OCR Scanning vs Image-Only Scanning: What’s the Real Difference?
The difference is simple to state and expensive to get wrong: an image-only scan is a photograph of a page, while an OCR scan is a photograph of a page with a machine-readable text layer sitting invisibly behind it. Both produce a PDF that looks identical on screen. Only one of them lets you type “Henderson” into a search box and find the file in two seconds. If you are digitising an archive to make it usable rather than just to clear shelf space, that distinction decides whether the project pays for itself or quietly becomes a folder of digital paper nobody opens.
What image-only scanning actually produces
An image-only scan captures each page as a bitmap — TIFF or a PDF with an embedded image — and stops there. The file contains pixels arranged in the shape of letters, but the software has no idea those pixels spell anything. To a computer it is no different from a photo of a brick wall.
That still has real uses. Image-only is faster to produce, cheaper per page, and entirely sufficient when:
- The documents are already indexed by a reference number captured at scan time (barcode, job number, client ID) and nobody will ever need to search inside them
- You are scanning for evidential preservation — signed deeds, wills, engineering drawings — where the visual record is the point
- The source material is handwritten, heavily annotated, or in poor condition, where OCR would produce more noise than signal
- The volume is enormous and the retrieval rate is near zero — deep archive that exists to satisfy a retention rule, not to be read
The trap is choosing image-only by default because it was the cheaper line on the quote, then discovering eighteen months later that the only way to find anything is to open files one at a time.
What OCR adds — and what it doesn’t
Optical Character Recognition analyses the scanned image, identifies character shapes, and writes out the text it recognises. In a searchable PDF that text sits as an invisible layer aligned to the image, so the page still looks exactly like the original — you just gain the ability to search it, copy from it, and feed it into other systems.
The practical gains
- Full-text search — find a supplier name, a policy number, or a date across 400 boxes’ worth of files in seconds rather than days
- Automated indexing — zonal OCR can read an invoice number or NI number from a fixed position on the page and populate metadata without anyone typing it
- Subject access requests — under the UK GDPR and Data Protection Act 2018 you have one month to respond to a DSAR. Searching a text layer for an individual’s name is the difference between a compliant response and a scramble
- Redaction — you cannot reliably redact what you cannot locate. Text-layer search makes finding every instance of a name tractable
- Downstream integration — document management systems, workflow tools, and analytics all need text, not pixels
The honest limits
OCR is not transcription and it is not perfect. On clean, machine-printed A4 at 300 dpi a good engine will typically reach the high 90s in character accuracy. On a faxed copy of a photocopy, on a dot-matrix printout, on a form with text overlapping ruled lines, or on cursive handwriting, accuracy drops sharply — and handwriting in particular is a job for ICR, with results that need human verification before you trust them.
The maths matters more than the headline percentage. A page holding roughly 2,000 characters at 99% character accuracy still contains about 20 errors. That is usually invisible in body text — but if one of those errors lands in the one surname you are searching for, that document is functionally missing from your archive. This is why serious providers quote accuracy at field level for indexed data, not just character level across the page.
A worked comparison
Take a mid-sized UK firm digitising 300 archive boxes — call it 750,000 pages. Assume a retrieval need of 40 file lookups a month across the archive.
Image-only route. Lower cost per page and a faster turnaround. But every lookup means identifying the right box from a spreadsheet, opening candidate PDFs, and eyeballing pages. Fifteen minutes per lookup is optimistic. That is 10 hours a month, 120 hours a year — around three working weeks of someone’s time, indefinitely, plus the failures where the file simply isn’t found.
OCR route. Higher cost per page and a longer project. But that same lookup is a search box and 30 seconds. The 120 hours a year collapses to under four. On a £35,000 fully-loaded salary, roughly £2,000 a year of recovered time — before you count the DSAR you answered inside the deadline, or the audit request you satisfied the same afternoon.
The premium for OCR is paid once. The cost of not having it is paid every month, forever. That is the calculation that actually decides it — not the per-page rate on the quote.
How to decide, box by box
The most common mistake is treating this as one decision for the whole archive. It rarely is. Most archives split cleanly:
- OCR it — anything containing personal data (HR files, client records, correspondence), anything you search by content rather than reference, anything with an active compliance or litigation dimension, and anything typed cleanly enough to give the engine a fair chance
- Image-only is fine — plans and drawings, signed originals kept for evidential value, bulk historic material retained purely to satisfy a retention schedule, and anything already reliably indexed by a captured reference
- OCR plus manual indexing — handwritten or degraded material that you genuinely need to find. Accept that the text layer will be imperfect and pay for keyed index fields on top
Ask any prospective provider three questions: what resolution and colour depth do you scan at before OCR runs; do you quote accuracy at character level or field level; and what does your QC process do when a page falls below threshold. A provider who cannot answer the third question is not doing QC — they are running OCR and hoping. Our document scanning service handles both routes, and a pilot on two or three representative boxes will tell you more about your own paper than any amount of theorising.
The bottom line
Image-only scanning solves a space problem. OCR scanning solves an access problem. If your files are leaving the office and never coming back, image-only may be all you need — and combining it with off-site document storage for the originals is a perfectly sound strategy. But if the point of digitising was to make information findable, an image-only archive has simply moved the haystack from a basement to a server. Decide by asking one question of each batch: will anyone ever need to search inside this? If the answer is yes, or even probably, OCR is not the upsell. It is the whole job.
More guidance on planning a digitisation project is available in our resources library.








