Guide
Document extraction QA: how to check OCR and AI output
Deepen AI · Published 2026-10-02
OCR and AI extraction tools return a value for almost every field, including the fields they got wrong. A confident wrong total is worse than an empty one, because it flows into your accounts payable system, claims queue or database without anyone looking. Document extraction QA is the set of checks that decides which values to trust, which to send to a person, and how you know the people are right too.
Step 1: write the field spec
For each document type, list every field with:
- [ ] Name and data type: text, date, money or code
- [ ] Format, such as YYYY-MM-DD dates or ISO currency codes
- [ ] Whether it is required
- [ ] Where it appears, and what to do if it appears twice
- [ ] Validation rules (Step 2)
- [ ] Whether it is critical, meaning an error costs money or breaks compliance
Critical fields get stricter treatment at every later step. On an invoice, the total, currency, supplier and invoice number are usually critical. A free-text memo is not.
Step 2: add rules a machine can check
Run these before a person looks at anything. They cost little to run and catch many errors.
- Format: dates parse, amounts are numbers, codes match their pattern.
- Arithmetic: line items add up to the subtotal, and subtotal plus tax equals the total.
- Cross-reference: the supplier matches your supplier master, and the PO number exists in your system.
- Plausibility: dates fall in a sensible range, and amounts sit within the usual band for that supplier.
- Duplicates: the same supplier and invoice number have not been processed before.
Step 3: route the right items to people
Send three groups to human verification:
- Failed rules. Any document that failed a check above.
- Low confidence. Fields below a confidence cut-off you set from pilot data. Confidence scores are a routing signal, not proof. Test them against known answers before you trust them.
- A random sample of passed documents. Without this, you only ever measure the documents you already suspected, and you never learn what passes the rules wrongly.
The human task is narrow: compare each routed field with the source image, correct it, and tag the error type. If the source is unreadable, flag it. Do not guess.
Step 4: use a fixed error taxonomy
Tag every correction with one type. A short, fixed list makes the weekly pattern obvious.
| Error type | Example | |---|---| | Character confusion | O read as 0, l as 1, S as 5 | | Date format | 03/09 read as March 9 when the supplier writes the day first | | Default value | Currency set to USD because the field was blank | | Wrong field | Tax amount placed in the total | | Missed item | A line item on page 2 skipped | | Split or merge | Two invoices in one PDF treated as one | | Entity match | Supplier name not matched to the master record | | Unreadable | Source too poor to read, so the item was flagged |
Step 5: check the checkers
People make errors too. A QA reviewer re-checks a sample of human-verified documents, field by field, against the source. Sample across every person, every document type and every major supplier. Count errors per field as well as per document, because a document-level pass count hides which fields cause the trouble.
Build a gold set of 50 to 200 documents with known-correct fields, including the awkward ones: multi-page invoices, handwritten amounts, poor scans and foreign date formats. Use it to train new people and to retest whenever the extraction model or template changes.
Step 6: report and fix upstream
A useful weekly report shows documents processed, documents and fields routed to people, the QA sample size and what it found, errors by field and type, and open questions. The goal is fewer errors next week. If "default value" keeps appearing for currency, fix the default in the extraction setup instead of correcting it by hand forever.
Worked example
Illustrative example. Invented data, not from a customer.
Source invoice: Acme Parts Ltd · Invoice no. INV-20417 · Date 03/09/2026 (this supplier writes the day first) · Total 1,284.50 EUR · VAT 214.08
What the extraction returned, and what review did:
- invoice_number: INV-2O417 → INV-20417 (character confusion)
- invoice_date: 2026-03-09 → 2026-09-03 (date format)
- supplier: Acme Parts → Acme Parts Ltd (entity match to the supplier master)
- total: 1284.50 (checked, no change)
- currency: USD → EUR (default value)
- vat: 214.08 (checked, no change)
The arithmetic check passed, because the total and VAT were right. A pattern check on invoice numbers would have caught the letter O, and the supplier master check would have caught the name. The swapped date and the default currency looked valid to every rule and needed a person. Both became supplier-specific rules in the extraction setup the same week.
If you use Amazon Textract with A2I
Per AWS documentation, A2I's built-in Textract task type can start a human review when specific form keys are missing or detected with low confidence, and can also send a random percentage of documents for review. As of 2026-10-02, AWS states that A2I is no longer open to new customers, so this route applies to existing users. The A2I private workforce set-up guide covers the workforce side.
Pitfalls
- Only reviewing low-confidence fields. You will never see the confident errors.
- Trusting confidence scores untested. Check them against known answers, by field and by document type.
- Silent defaults. Log any default value as a default, not as extracted data.
- Template drift. A supplier redesigns its invoice and errors jump. Track errors by supplier.
- Correcting data, not causes. Repeated error types belong upstream, in rules or extraction settings.
- Reviewers fixing the source instead of flagging it. If a document is wrong at the source, flag it for your team to decide.
Checklist
- [ ] Field spec with types, formats and critical fields
- [ ] Format, arithmetic, cross-reference, plausibility and duplicate rules
- [ ] Routing for failed rules, low confidence and a random passed sample
- [ ] Fixed error taxonomy on every correction
- [ ] QA sample of human work across people, document types and suppliers
- [ ] Gold set of 50 to 200 documents, rerun after any model or template change
- [ ] Weekly report with errors by field and type
If you want a trained team to verify extracted fields in your own tools, see document data extraction, and how we check quality for the sampling process behind it.
FAQ
What is document extraction QA?
The checks that decide which OCR or AI-extracted values to trust, which to send to a person, and how you confirm that the human corrections are right too.
Which extracted documents should a person review?
Documents that fail validation rules, fields below your confidence cut-off, and a random sample of documents that passed, so you also catch confident errors.
How large should a gold set for document extraction be?
Start with 50 to 200 documents with known-correct fields, including multi-page, handwritten and poor-quality examples, and retest against it whenever the extraction model or template changes.
Should extraction errors be measured per document or per field?
Both. Field-level counts show which fields cause errors. Document-level counts show the effect on your process.