shipwithjev

Blog / Builds & people / FIG. 92

Jev for Data Teams: Judgment as Infrastructure

Putting an LLM in the data pipeline without the chaos: typed verdicts for AI data quality checks, enrichment, dedup, and labels, all versioned.

Data teams have always had a judgment gap. The pipeline can check that a field is not null, that a date parses, that a foreign key exists. It can't check that the "company name" field contains a company, that the free-text "other" category is actually one of the real categories, or that two customer records are the same business. Those checks went to humans, which in practice meant they went nowhere.

Putting an LLM in the data pipeline fills that gap, but only if the model behaves like infrastructure: typed outputs, stable questions, versioned results, measurable error. That's a narrow description of what Jev, the decision model from TypeSafe AI, does (primer). It returns closed answers with probabilities and never writes free text, which is exactly the property you want in a column.

This page owns the pipeline seat: where judgment sits in batch jobs and what discipline it needs. Semantic filters inside queries are the natural language database queries page's territory.

Four jobs for an LLM in the data pipeline

1. AI data quality checks. Semantic validation next to your schema tests: Is this "job title" a job title? Is this address plausibly real and complete? Does this product description match its category? Failures route to a quarantine table, not the trash.

2. Enrichment. New columns derived from text: ticket category, "mentions a competitor," review sentiment by aspect. The materialized verdict columns recipe is the build for this: judge once, store with the question version, query forever.

3. Deduplication. Pairwise "same entity?" verdicts on candidates from a blocking step. Entity resolution owns the method, including why blocking is mandatory unless you enjoy paying for the square of your row count.

4. Labels for training data. Model-first labeling with a human-checked sample, feeding a classifier you train yourself. The data labeling page covers the workflow and its honesty section.

The infrastructure discipline

A model in a pipeline is a dependency, and dependencies get the same treatment as any other.

  • Version the questions. A question's wording is code. Store the version alongside every answer, and treat a rewording as a migration with a backfill.
  • Store probabilities. Downstream consumers can filter by confidence, and you can move thresholds without re-judging.
  • Make it incremental. Judge new and changed rows only. Full-table re-runs are for version bumps.
  • Quarantine, don't drop. Low-confidence or failing rows go to a table a human reviews. Deleting data on a lone verdict is irreversible; quarantining isn't.
  • Monitor drift. Track the answer distribution per question over time. A sudden jump in "other" means either your data changed or your upstream did; either way, someone should look. When that someone is on call, AI incident triage covers how verdicts can sort the alert and what stays human.
  • Keep an answer key. A few hundred human-labeled rows per question, re-scored whenever the question or the model changes.

The pseudocode shape of a quality check, for orientation only (per ecosystem documentation, Jev is reached through the Vercel AI Gateway as typesafe-ai/jev; real syntax lives at docs.typesafe.ai):

# pseudocode, not real API syntax
for row in new_rows(contacts):
    v = judge(row.company_name, ["Is this a real company or organization name?"])
    if v.answer == no and v.probability over 0.9:  quarantine(row, reason = v)
    else:                                            pass_through(row)

What the receipts show

The building blocks already exist in the directory. pg-jev is a PostgreSQL extension for semantic questions over table rows, and JevQL offers semantic WHERE clauses for vanilla Postgres. Hamilton Ulmer's DuckDB extension classifies rows in CSV, Parquet, or DuckDB tables at about ten seconds per thousand rows, as reported, and he called it more ergonomic than a classifier (build).

On cost at volume, Ian Nuttall's archive run is the useful reference: 3,282 posts with eight questions each, roughly 26,000 verdicts, for $0.1282, as reported (build). Your mileage depends on text length and question count, so do the arithmetic on your own tables before promising anyone a number.

No official Jev benchmarks exist. For a data team that should feel familiar: the only accuracy number worth trusting is the one you measured against your own answer key.

Where it doesn't belong

  • Anything a rule can do. Regex, lookups, and constraints are free and deterministic. Use judgment for what rules can't express.
  • Generating text. Descriptions, summaries, and translations need a generative model; Jev only decides.
  • Unreviewed high-stakes fields. If a column drives pricing, credit, or anything regulated, the verdict is an input to a human process, with an audit log, not the decision itself.

Frequently asked questions

How do I put an LLM in a data pipeline safely?

Treat it like any dependency: typed outputs, versioned questions, stored probabilities, incremental runs, and a human-labeled answer key. Low-confidence rows go to quarantine, not straight to production tables.

What are AI data quality checks?

Semantic validations that rules can't express, like whether a field contains a real company name or a plausible address. They run beside schema tests and route failures for review.

Should verdicts be computed at query time or stored?

Store stable, frequently read answers as columns and keep query-time judging for exploration. The verdict columns recipe covers the storage side.

How do I handle duplicates across large tables?

Use cheap blocking rules to find candidate pairs, then ask a pairwise "same entity?" question. The entity resolution page explains why blocking is essential.

Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.