Services / Document data extraction and verification
Document data extraction and verification
We pull fields from documents into your systems, or check what your OCR or AI extraction already pulled. Each field is checked against the source document and your field rules before it moves on. Each week you also get a list of the errors your extraction makes most, so you can fix them upstream.
Supported tasks
- ·Key invoice fields into your accounts payable tool
- ·Verify and correct OCR or AI-extracted fields from invoices, claims and forms
- ·Review Amazon Textract results in your A2I portal
- ·Enter bills of lading and shipping documents into your logistics system
- ·Index, split and rename scanned files by customer, date and document type
- ·Flag missing or unreadable documents for your team
Sample input and output
Illustrative example. Invented data, not from a customer.
Sample input (the source invoice)
Acme Parts Ltd · Invoice no. INV-20417 · Date 03/09/2026 (this supplier writes day/month/year) · Total 1,284.50 EUR · VAT 214.08
Sample input (what the OCR extracted)
- invoice_number: INV-2O417
- invoice_date: 2026-03-09
- supplier: Acme Parts
- total: 1284.50
- currency: USD
- vat: 214.08
Sample output (after review)
- invoice_number: INV-20417 (fixed: the letter O was read in place of zero)
- invoice_date: 2026-09-03 (fixed: day and month were swapped)
- supplier: Acme Parts Ltd (fixed: matched to your supplier list)
- total: 1284.50 (checked, no change)
- currency: EUR (fixed: USD was a default, the invoice says EUR)
- vat: 214.08 (checked, no change)
- Error types logged: character mix-up (1) · date format (1) · supplier match (1) · default currency (1)
How it works
- 1
You share documents and field rules. Send sample documents, your field rules and access to the tool.
- 2
We train and build a gold set. The team trains on your samples. We build a set of documents with known-correct fields and confirm it with you.
- 3
The team keys or verifies each document. They work in your tool. Unreadable or unclear documents go to the team lead.
- 4
You get a weekly report. Documents done, QA sample results, error types by field, and open questions.
Quality checks
The team lead checks: unreadable and unclear documents, changes to field rules, and each person's QA results.
QA reviewers sample: a share of finished documents, re-checked field by field against the source. The report shows which fields cause most errors, so you know where to look.
What this means: every document is handled by one trained person. QA re-checks a sample, not every document. The weekly report shows the sample size and the acceptance rate.
Pricing scope
- $5 per person-hour (standard tasks): keying and verifying fields from English documents, to your field rules.
- Quoted after our ops team reviews the task: documents that need specialist knowledge (for example medical coding or legal review), languages other than English, or setup of a new extraction tool.
FAQ
Can you check our OCR or AI extraction instead of keying from scratch?
Yes. That is often the faster setup. We correct what the machine got wrong and log each error type, so you can see where the extraction needs work.
Can you work in our Textract review loop?
Yes. We join your Amazon A2I portal as a private workforce. See the Ground Truth and A2I page for the steps.
Our documents contain personal data. Is that a problem?
No, if you tell us first. We agree in writing who can see it, which fields they need and what gets masked. See the security page.
Is every document checked twice?
No. One trained person handles each document. QA reviewers re-check a sample, and the weekly report shows how many.
Try it on your own documents for two weeks.
Send your field rules and a few sample documents. We email a pilot plan with a quote.
Request a pilot plan →