Document understanding

Intelligent Document Processing (IDP) Services

Extraction is the easy half. What decides whether a document pipeline is safe to run is the threshold: which fields are confident enough to post automatically, and which ones a person still checks. We make that a number you set, not a number we hide.

ISO 27001 certified Human review by design Under 3 minutes to first reply

One document, two destinations

Extracted fields are split by confidence: those above the threshold post automatically, and those below it route to a human review queue rather than entering the system unchecked. Document in Extract fields, score each one layout + language model, not plain OCR Above your threshold? Posts automatically straight into your system Human review queue correct, do not retype

What intelligent document processing is

Intelligent document processing extracts structured data from invoices, contracts, KYC packs and forms using OCR and language models. Every field carries a confidence score, and anything below the agreed threshold routes to a human review queue instead of entering your systems unchecked.

153+Projects Delivered
4.9/5Across 153 Reviews
24–48hSOW To Discovery Sprint
100%Labelled Data Assigned To You

Why document projects disappoint after the demo

The demo used five clean invoices. Production has eleven supplier formats, a fax and something photographed at an angle.

A single accuracy number was quoted

Ninety-something per cent, on a sample chosen by whoever was selling. Accuracy varies enormously by document type, field and scan quality, and an average conceals exactly the cases you need to plan for. Ask instead what the review rate is at a threshold you would accept.

Reviewers ended up retyping everything

The review interface shows a document and an empty form, so the person reads and keys exactly as before, with an extra login. A correct review queue shows the page image and the extracted values side by side, with low-confidence fields highlighted, so the work is correcting rather than transcribing.

Extraction worked, posting did not

Fields come out cleanly and then sit in a CSV because the accounting system has no interface. Getting verified data into the system of record is frequently the larger half of the project, and it needs integration work or a robot to key it.

The corrections went nowhere

Reviewers fix the same field on the same supplier's invoice every week and nothing improves, because corrections were never captured as labelled data. That feedback loop is the difference between a review queue that shrinks and one that is permanent staffing.

The real decision

Move the threshold and watch the trade-off

Twelve fields from one supplier invoice, with the confidence the extraction returned. Three of them are actually wrong. Lower the threshold to automate more and you will start approving them.

Fields scoring below this go to a person.

0.90
Invoice number0.98auto
Supplier name0.96auto
Invoice date0.94auto
Invoice total0.93auto
Purchase order reference0.89review
Line 1 quantity0.86review
Tax amount · misread0.84review
Line 2 description0.81review
Currency0.79review
Line 3 quantity · misread0.76review
Payment terms0.72review
Handwritten note · misread0.61review

Posted automatically

4

of 12 fields

To a person

8

review queue

Wrong values approved without review

0

Nothing wrong is getting through

Every misread field is caught. The cost is a large review queue, which is the correct starting position for anything touching money or identity. Lower it deliberately once you have data.

Illustrative sample, not a benchmark. On a real engagement these numbers come from a labelled set of your own documents, which is why step two of the build is labelling rather than modelling.

Run this on your documents
Capabilities

What a document automation build includes

The extraction engine is a component you can swap. These are the parts that make it operable.

A review queue built for correcting, not retyping

The page image and the extracted values side by side, low-confidence fields highlighted and focused first, keyboard navigation between them, and the source region on the page highlighted as each field is selected. A reviewer should be able to clear an item in seconds rather than minutes.

Every correction is captured as labelled data, which is what makes the queue shrink over time rather than becoming permanent headcount.

Classification before extraction

A mailbox receives invoices, credit notes, statements and things that are none of those. Classifying first means each document type gets the extraction schema and threshold that suits it, instead of one model attempting everything.

Validation against what you already know

Does the supplier exist, does the total match the lines, does the PO reference resolve? Cross-checks catch errors that confidence scores miss, and they are cheap to add.

A labelled ground truth set

Built once, hand-checked, and used to measure every change afterwards. It is the only thing that makes an accuracy claim falsifiable, and it belongs to you at the end.

Documents are sensitive by default

Invoices and identity packs carry personal data. Where residency rules bite, extraction runs entirely inside your network. Reviewed under AI governance and by our security team.

Build sequence

How a document pipeline gets built

Five steps, and the first two are the ones that get skipped. Same delivery shape as the rest of the AI and automation practice.

  1. Collect a realistic document sample

    Two to three hundred documents, deliberately including the bad scans, the unusual suppliers and the ones somebody photographed on a phone. A sample of clean examples produces an accuracy figure that does not survive its first week in production, and everyone knows it except the person who chose the sample.

  2. Label a ground truth set

    A subset labelled by hand, field by field. This is tedious and it is the foundation of everything after it: without it, accuracy is an assertion and the threshold conversation has nothing to stand on. It also becomes an asset you keep.

  3. Benchmark engines per document type

    AWS Textract, Azure Document Intelligence, Tesseract and layout-aware language models all win on different material. We measure rather than standardise, and the winner for invoices is frequently not the winner for identity documents.

  4. Set the threshold with the business

    We model the trade-off explicitly, per field, using the labelled set: what auto-approves, what queues, and how many errors slip through at each level. Finance signs the number, not engineering, and it is recorded as a decision with a date.

  5. Build the review queue and the feedback loop

    Side-by-side correction, captured labels, and monitoring on review rate and correction rate so a drop in quality is visible before somebody complains. Then handover, with the labelled data and the configuration assigned to you.

Four document types, four different problems

Each needs a different schema, a different threshold and a different definition of a bad day.

Invoices and credit notes

The most common starting point and the best understood. Header fields extract reliably; line items are where the difficulty lives, because tables wrap across pages, merge cells and use supplier-specific column orders.

The strongest control here is not the model, it is the three-way match against the purchase order and the goods receipt, which catches errors no confidence score would flag. Usually orchestrated as a workflow.

Hard part

Line item tables across pages

Best control

Three-way match, not confidence alone

The pipeline, stage by stage

Six stages between a file arriving and a record existing. Select one to see what it does and how it fails.

Ingest and normalise

Files arrive from a mailbox, an SFTP drop, a portal or a phone. Everything is converted to a consistent image and page format, deskewed, and given a stable identifier so the same document arriving twice is recognised as a duplicate.

Fails when

Duplicate detection is missing and the same invoice enters twice, which is a finance problem long before it is a data problem.

One invoice, two working days

The same document handled manually and through a pipeline. Notice that a person still appears in the second version, and that this is the design rather than a shortfall.

Around 20 minutes of work, spread across two days

  1. Invoice sits in a shared mailbox

    Until someone on AP duty opens it.

  2. Header and line items keyed by hand

    Reading from a PDF on one screen and typing into another.

  3. Purchase order looked up separately

    A second system, a second search, from memory of the reference.

  4. Quantity mismatch noticed, or not

    Depending on how busy the day is. This is the expensive branch.

  5. Approver decides

    The one step that genuinely required judgement.

What disappeared was the transcription and the two-day wait. What did not disappear was the judgement, and any proposal that claims otherwise is describing a different process. The wider version of this argument sits under business process automation.

Where document volume concentrates

Four sectors where extraction pays back fastest. Profiles are anonymised.

Onboarding and KYC packs

Identity documents, address proofs and application forms arriving in one bundle. Extraction plus cross-document consistency checking removes the assembly work, and the analyst spends their time on the judgement rather than the collation.

Banking and insurance

Supplier invoice intake

Dozens of supplier formats into one accounting system, with three-way matching against purchase orders and goods receipts. The highest-volume and best-understood application in this list, and usually the first one built.

Manufacturing engagements

Proof of delivery and consignment notes

Photographed on a phone at a loading bay, often with a signature and a handwritten annotation. Capture quality dominates results here, and improving how drivers take the photograph beats any model change.

Logistics engagements

Referrals and claim forms

Arriving by fax, email and portal in equal measure. Extraction plus routing gets a referral to the right team the same day rather than whenever somebody opens the shared inbox. Personal data throughout, so residency is settled first.

Healthcare engagements

OCR, template extraction, or IDP

Three approaches frequently sold under the same name. The row that decides it is usually the third one.

Comparison of plain OCR, template-based extraction and intelligent document processing across six criteria.
CriterionPlain OCRTemplate extractionIntelligent document processing
OutputRaw textFields, per known layoutNamed fields with confidence per field
New supplier layoutWorks, but output is unusableNeeds a new template builtUsually works without configuration
Tells you when it is unsureNoRarely, and only per documentYes, per field, which is what makes review possible
Handles tablesNo structure preservedIf the template covers itYes, including continuation across pages
Ongoing maintenanceNoneA template per layout, foreverThreshold tuning and periodic re-evaluation
Best whenYou only need searchable textFew layouts, never changingMany layouts, and errors have consequences

If you genuinely have three supplier layouts that never change, template extraction is cheaper and more predictable, and we will say so. The case for IDP is variety plus consequence, and the confidence score is what buys you the second half of that.

Clear Answers

Document processing questions

The ten that come up in almost every first conversation.

Send us fifty documents, including the bad ones

Especially the bad ones. We will come back with a measured review rate at a threshold you would accept, rather than an accuracy percentage. Reply in under 3 minutes.

CYBER WARRIOR ZERO TRUST SECURITY CUSTOM WEB ENGINEERING VAPT AUDITING AWS CLOUD ARCHITECTURE ENTERPRISE AUTOMATION CYBER WARRIOR ZERO TRUST SECURITY CUSTOM WEB ENGINEERING VAPT AUDITING AWS CLOUD ARCHITECTURE ENTERPRISE AUTOMATION