shipwithjev

Blog / Recipes / FIG. 88

OCR + Jev: A Document Intake Pipeline

Document intake automation in three moves: scan, classify, route. OCR reads the page, Jev judges the text, humans handle anything touching money.

Every intake inbox is a pile of PDFs pretending to be a workflow. Invoices, contracts, forms, scans of forms, photos of scans of forms. Document intake automation usually stalls on the same fact: the model you'd like to use can't see, and the model that can see costs too much to run on every page.

The split that works: OCR reads, Jev judges. Jev is the decision model from TypeSafe AI (primer), and it's text-only, per ecosystem documentation. Why that's a pattern rather than a limitation is covered in can Jev judge images, and the theory of classifying documents belongs to the document classification page. For the wider category this pipeline sits in, see what intelligent document processing is. This recipe is the pipeline: scan, classify, route.

The pipeline at a glance:

  1. Land: files arrive by email, upload, or a watched folder.
  2. Extract: text-bearing pages yield text directly; image pages go to OCR.
  3. Classify: Jev answers closed questions about each page or document.
  4. Route: code sends each document to a queue, a folder, or a human.

Every stage is replaceable. Swap OCR engines, change the queues, reword questions, without touching the others.

Step 1: extract text, cheaply first

Don't OCR what already has text. Many PDFs carry a text layer, and pulling it locally is free and exact. doc-router, a public build in the directory, is built on precisely this idea: a Rust tool that asks Jev which PDF pages actually need OCR, extracts text pages locally, and sends only the rest to the OCR provider.

For scanned pages and photos, any OCR engine works. The screenshot-free pattern from the macOS loop build, local Apple Vision OCR feeding Jev's decisions, is the same trick applied to a screen: perception stays local, judgment goes to the cheap text model, and pixels never travel.

Step 2: classify with closed questions

Per page or per document, ask what intake needs to know:

  • What kind of document is this: invoice, receipt, contract, ID, form, correspondence, other? (choice)
  • Is the text readable enough to judge, or is the OCR garbage? (yes/no)
  • Does it reference an amount due or a payment deadline? (yes/no)
  • Does it appear to be incomplete, such as a missing page or a cut-off signature block? (yes/no)

The "is the OCR garbage" question is the one people forget. A bad scan should route to a human, not to a confident wrong category.

# pseudocode, not real API syntax
for doc in new_documents:
    text = has_text_layer(doc) ? extract_local(doc) : ocr(doc)
    v = judge(text, INTAKE_QUESTIONS)
    route(doc, v)

Per ecosystem documentation, Jev is reached through the Vercel AI Gateway as typesafe-ai/jev and returns probabilities with choice answers; check docs.typesafe.ai for real syntax. Long documents may exceed what fits in one call; judge the first pages or per-page and roll up.

Step 3: route by confidence and consequence

Routing is plain code:

  • Confident and low-stakes (correspondence, receipts for filing): straight to the right folder.
  • Confident but money-touching (invoices, payment demands): to the finance queue, pre-labeled, for a person to approve.
  • Unreadable, incomplete, or low confidence: to a manual review queue with the reason attached.

Nothing gets paid, signed, or submitted because a classifier said so. Classification tells a person which pile to look at; it doesn't replace the look. For tax and legal documents, this is triage, not advice. The public tax-doc-classifier build says exactly that in its own description, with a confidence gate and a note that scanned pages need OCR.

Document intake automation receipts

Two builds show the economics, both as reported by their authors. Fayaz Ahmed's image classifier, OCR first and Jev second, sorted about 900 images in 40 seconds (build). Trinay Hari reports Jev classified a 26-sheet construction plan set in 2.9 seconds for $0.0052, matching GPT-4.1 and GPT-6 Astra on 100% of sheet-level classifications in his comparison (build). Those are single builders on their own data; there are no official Jev benchmarks, so test on yours.

Where it breaks

  • Handwriting and bad scans: OCR quality caps everything downstream.
  • Layout-heavy meaning: a table where position carries the meaning can lose it in extraction.
  • Truly visual questions: "is this signature genuine?" is not a text question. Use a vision model or a person.

Frequently asked questions

What is document intake automation?

A pipeline that extracts text from incoming files, classifies each document, and routes it to the right queue. Humans handle low-confidence and high-stakes items.

Can Jev read scanned documents?

Not directly, because it judges text. Run OCR first, then send the text; the image workaround page explains the pattern.

How do I scan, classify, and route invoices safely?

Classify and pre-label them automatically, then send them to a person who approves payment. Payments should never trigger on a lone model verdict.

Is this reliable for contracts and compliance?

It's reliable enough to sort and flag, not to decide. The document classification page covers the review posture; nothing here is legal advice.

Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.