Blog / Questions / FIG. 151
What Is Intelligent Document Processing? The Honest Explainer
Intelligent document processing turns invoices, forms, and mail into data and routed decisions. How IDP works, how it fails, and how to roll it out.
Intelligent document processing (IDP) is software that turns incoming documents into structured data and routed decisions, with humans reviewing only the cases the system is unsure about. Invoices, intake forms, contracts, scanned mail: IDP reads them, decides what they are, pulls out what matters, and sends each one somewhere useful.
It's the successor to template OCR. Old systems needed a box drawn around every field on every layout. Modern IDP reads documents it has never seen before, which is the whole point, and also the source of its new failure modes.
The five stages of an IDP pipeline
Every IDP system, bought or built, runs the same skeleton:
- Capture. Scans, PDFs, photos, and attachments arrive. OCR turns pixels into text and layout.
- Classify. What kind of document is this? Invoice, receipt, contract, junk.
- Extract. Pull the fields: vendor, amount, dates, line items.
- Validate. Do the values make sense? Totals add up, required fields exist, this isn't a duplicate.
- Route. Post to the system of record, or send to a human queue when confidence is low.
Vendors differ in how well they do each stage. They don't differ in the stages. If a pitch skips validation, that's the part to ask about.
How LLM-era IDP differs from template OCR
Template OCR broke every time a vendor redesigned its invoice. The ML generation that followed trained extractors per document type, which meant labeled data for every new type.
Generalist models changed the economics. They read unseen layouts, and the work moved from training to instructions and calibration. The catch is honest and important: a generative extractor can produce a confident, well-formatted total that isn't on the page. Validation stopped being a nice-to-have and became the stage that keeps accounting clean.
A worked example: one invoice, five stages
A PDF lands in the accounts inbox. The text layer is empty because it's a phone photo of a paper invoice, so capture runs OCR and keeps the page layout alongside the text.
Classify asks a closed question: invoice, receipt, statement, or other? Invoice comes back with high probability, so it moves on. A coin-flip answer here would stop the document before any field gets extracted, because extracting invoice fields from a statement just produces tidy garbage.
Extract pulls vendor, invoice number, date, line items, and total. Validate then checks them against each other and against your records: line items sum to the total, the vendor exists in your system, this invoice number hasn't been seen before, the date isn't in the future.
Say the line items sum to 1,240 and the extracted total says 1,420. That's a transposed digit or a misread, and it's exactly the case validation exists for. Route sends it to a person with both numbers highlighted. Clean invoices go to the approval queue as drafts. Nothing gets paid by the pipeline itself.
The full invoice version of this flow, with the accounting specifics, lives in AI invoice processing.
Where decision models fit, and where they don't
Jev, the decision model from TypeSafe AI, generates no text, so it doesn't extract fields. That's a feature here, not a gap. It judges, and IDP is full of judgments:
- Classification as a closed choice: which document type, with a probability per option.
- Validation as yes/no questions: does the extracted total match the line items? Is the scan legible enough to trust?
- Routing: auto-post or human review, decided by a confidence band.
The split is clean: an extractor proposes values, verdicts check them. That's the verifier shape from five cascade architectures, and the full buildable loop lives in the OCR-plus-verdicts intake recipe. For the theory of the classification step itself, document classification goes deeper.
Builders keep reaching for this shape. One sorted about 900 images in 40 seconds by running OCR first and letting Jev pick each category, as reported in the OCR image classifier build. Another classified a 26-sheet construction plan set in 2.9 seconds for $0.0052, as reported in the plan-set classifier build. More patterns by document type are collected in IDP examples.
Privacy gets easier too. The doc-router build asks Jev which PDF pages actually need OCR, extracts the text pages locally, and sends only the rest to an OCR provider, as reported by its author. Fewer pages leave the machine, and the OCR bill shrinks with them.
One rule holds in every IDP design on this site: payments, rejections, and anything else irreversible never happen on a lone verdict. A human or a second independent check signs off.
What IDP costs now
Packaged IDP platforms typically price per page or per document under annual contracts. Assembled pipelines price by their parts: OCR (often local and nearly free), extraction calls, verdict calls, and human review minutes.
The verdict line is usually the smallest. One builder classified 500 emails for 3.5 cents, as reported, and document batteries scale from there with question count (napkin method here). The real cost centers are extraction on long documents and the people working the review queue. Price those first.
Rolling out IDP without regretting it
Volume times variety decides whether you need it. A handful of documents a week is cheaper to handle by hand. Thousands a month across dozens of layouts is where IDP pays for itself. If you're past that line, the next question is whether to buy a platform or assemble one, which gets its own page.
Once you're in, the rollout that tends to work:
- Collect a real sample. A hundred or more documents from actual traffic, including the ugly scans, with human labels for type and key fields.
- Measure per stage. Classification agreement, field-level extraction accuracy, and validation catch rate, separately. One blended accuracy number hides which stage is failing.
- Run in shadow. The pipeline processes everything while humans keep doing the work. Compare for a few weeks.
- Set the bands. Pick the confidence cutoffs for auto-route versus review from your shadow data, using confidence thresholds as the method.
- Design the review seat. Show the reviewer the page, the extracted value, and why it was flagged. Human-in-the-loop design covers the queue mechanics.
The failure modes to watch are consistent. Silent extraction errors that pass validation because nobody wrote the check. Classification drift when a new document type starts arriving and gets forced into the nearest existing bucket, which is why an "other" option matters. Review queues that grow until people rubber-stamp them. And for contracts, scope creep: flagging whether a clause is present is a document task, while deciding what that clause means for you is a lawyer's.
Frequently asked questions
Is IDP the same as OCR?
No. OCR is the capture step, pixels to text. IDP is the whole pipeline that turns that text into classified, validated, routed decisions.
Does intelligent document processing need training data?
Not to start, in the LLM era. You still need about a hundred real documents with human labels to measure agreement before trusting it, per the calibration ritual.
What's the most common IDP failure?
Trusting extracted values without validation. A misread or invented total that flows straight into accounting is the classic incident.
Can IDP run without sending documents to the cloud?
Partly. OCR can run locally, and a verdict layer only needs the text it judges, as the screenshot-free build shows. Anything that does leave the device needs your own vendor review.
Where does a human stay in the loop?
On low-confidence classifications, failed validations, and every irreversible action. The uncertainty band is the human's seat, by design, and human-in-the-loop design covers how to build it.
Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.