How to Evaluate OCR Accuracy Before Choosing a Scanning Partner

The only reliable way to evaluate OCR accuracy is to test a provider on your own paperwork, measure the result at character and field level, and compare it against a defined pass mark written into the contract. Headline claims of “99% accuracy” are close to meaningless on their own — 99% character accuracy on a typical A4 page of 2,500 characters still leaves around 25 wrong characters per page, which is enough to break a name, a National Insurance number, or an invoice total. What matters is not the number a provider quotes, but how it was measured, on what documents, and what happens when it is missed.

Understand what “accuracy” is actually measuring

Before you can compare quotes, you need to know which of three very different metrics a provider is quoting. They are routinely used interchangeably in sales material, and the gap between them is enormous.

  • Character accuracy — the percentage of individual characters recognised correctly. The most flattering number, and the one most often quoted. 99.5% character accuracy on a 2,500-character page means roughly 12 errors per page.
  • Word accuracy — the percentage of whole words correct. Always lower than character accuracy, because a single wrong character fails the whole word. 99.5% character accuracy typically lands around 97–98% word accuracy.
  • Field accuracy — the percentage of captured index fields (invoice number, date, surname, policy reference) that are correct. This is the number that actually determines whether your team can find a document, and it is the one you should be contracting on.

A provider who only quotes character accuracy is either not measuring field-level capture or does not want to be held to it. Ask for all three, in writing, and ask which one the service level agreement is enforceable against.

Run a paid pilot on your real documents

Vendor demos use clean, laser-printed, single-sided originals. Your archive does not look like that. A meaningful evaluation uses a sample pulled from the messiest end of your own holdings.

Build a representative sample

  • 200–500 pages is enough to produce statistically useful results without a large outlay.
  • Deliberately include your worst material: faded thermal receipts, carbon copies, fax printouts, dot-matrix payroll listings, documents with handwritten annotations, and anything with a coloured or shaded background.
  • Include the document types you retrieve most often — those are where errors cost you.
  • Include a few known-difficult characters: the 0/O, 1/l/I, 5/S and 8/B confusions account for a large share of numeric field errors.
  • Cover the full range of page sizes and bindings you hold, including A3 plans, stapled bundles and double-sided sheets.

Score it yourself

Key a “ground truth” version of 50–100 pages by hand before you send the sample. When the results come back, compare the OCR output against your ground truth rather than accepting the provider’s own scorecard. Grading your own supplier’s homework is the entire point of a pilot. Count errors in three buckets: characters wrong, whole fields wrong, and pages where the error made the document effectively unfindable. That third figure is the one to take to your board.

Ask the questions that expose the process behind the number

Accuracy is an output of a process, not a property of the software. Two providers running the same OCR engine can deliver wildly different results depending on what happens either side of it.

  • What scan resolution do you use? 300 dpi is the practical minimum for reliable OCR on standard business text. 200 dpi will visibly degrade small print and is a common cause of quietly poor results.
  • What image pre-processing is applied? Deskewing, despeckling, background removal and adaptive thresholding routinely move accuracy by several percentage points on aged paper.
  • Is there a human verification stage? Ask specifically whether key index fields are double-keyed or blind-verified by a second operator, and whether that is included in the price or billed as an extra.
  • How are low-confidence results handled? A good workflow flags characters the engine is unsure of and routes them to an operator. A bad one silently writes the best guess into your index.
  • What is your QC sampling rate? Look for a defined percentage of output inspected against a documented standard, not “we check everything” hand-waving.
  • Which OCR engine, and is it tuned per document type? Providers who train recognition profiles for your specific forms will beat generic processing every time.

Set realistic targets by document type

Demanding a single blanket accuracy figure across a mixed archive is the fastest way to end up with a provider who agrees to it and then misses it. Segment your expectations instead. Clean modern laser-printed documents should return character accuracy comfortably above 99%. Older photocopies, faxes and carbon copies will sit meaningfully lower. Handwriting recognition remains the weak point across the whole industry — free-form cursive on a decades-old form should be treated as needing manual keying, not OCR, and any provider promising high automated accuracy on it is overselling.

The pragmatic approach is to require high, verified accuracy on the handful of index fields you actually search by — surname, date, reference number — and accept best-effort full-text OCR on the body content. That keeps the verification cost proportionate while protecting retrieval. It is the same logic that underpins good document scanning project design generally: spend your quality budget where retrieval depends on it.

Why OCR accuracy is a compliance issue, not just a convenience one

Under the UK GDPR and the Data Protection Act 2018, the accuracy principle requires that personal data is accurate and, where necessary, kept up to date. If a digitisation project introduces errors into personal data — a transposed date of birth, a corrupted NHS or National Insurance number — that is a data quality failure you have created, and the ICO can take enforcement action for serious breaches, with penalties reaching up to £17.5 million or 4% of global annual turnover for the most severe cases.

The individual right of access compounds the risk. A subject access request must normally be answered within one calendar month. If poor OCR means your search returns nothing for a data subject who is in fact on file, you have failed the request without ever knowing it. The same exposure applies to legal discovery and to regulatory inspections in financial services and healthcare, where “we couldn’t find it” and “we don’t hold it” are treated very differently.

Ask any prospective partner how they evidence accuracy after the event. BS 10008, the British Standard for the evidential weight and legal admissibility of electronic information, sets out the audit trail expected if scanned records may later be relied on in a dispute. A provider working to it will be able to show you their QC records; one who isn’t will describe their process verbally and move on.

Get the standard into the contract

An accuracy figure that appears only in a sales deck is not a commitment. Before you sign, make sure the agreement states the metric being measured, the target percentage, the sampling method used to verify it, and the remedy if it is missed — normally free re-scanning and re-indexing of the affected batch at the provider’s cost, within a defined window. Confirm who owns the rework decision and how disputes over sampling are resolved.

It is also worth agreeing what happens to your paper originals while quality is still being confirmed. Destroying source documents before the digital set has been signed off removes your only means of correcting an error. Keeping the originals in secure document storage through the verification period, then moving to certified shredding once accuracy has been accepted, costs very little and eliminates the worst-case outcome entirely.

Evaluate on your own documents, measure at field level, insist on a documented QC process, and write the standard into the contract with a remedy attached. Do those four things and OCR accuracy stops being a marketing claim and becomes something you can actually hold a supplier to. For more practical guidance on planning a digitisation project, browse the rest of our resources library.

    See how affordable we are:

    I am happy to receive newsletters and offers from Evastore