Blog / Recipes / FIG. 73
Classify a CSV With Jev in 20 Minutes
Classify a CSV with AI in about 20 minutes: pick columns, write one question, batch rows through Jev, and audit a sample. Real costs as reported.
You have a spreadsheet with a few thousand rows and a column that should exist but doesn't: "category", "is this a complaint", "priority". This recipe shows how to classify a CSV with AI using Jev, the decision model from TypeSafe AI, in roughly the time it takes to argue about which intern should do it by hand.
Whether a trained classifier would serve you better long term is a real question, and it lives in LLM vs traditional ML. This page assumes you've decided to just label the spreadsheet today.
What this costs before you start
The best public receipt for batch classification is Ian Nuttall's X archive run: 3,282 posts, eight questions each (roughly 26,000 verdicts), for $0.1282 in 8 minutes 34 seconds, as reported (build). On the tooling side, a DuckDB extension that classifies rows in CSV or Parquet files reported about 10 seconds per 1,000 rows (build).
Both numbers are builder-reported, not benchmarks, because no official Jev benchmarks exist. They're still enough for napkin math: rows times questions times a fraction of a cent. Most spreadsheets land in coffee-money territory.
The recipe
-
Pick the one column that matters (2 minutes). Jev judges text you hand it. Choose the column (or two) a human would read to make the call, like "ticket_body" or "comment". Drop IDs and timestamps from the payload; they're noise that invites weird verdicts.
-
Write one closed question (5 minutes). "What category is this?" is a wish. "Which of these five categories best fits this customer comment: billing, bug, feature request, praise, other?" is a question. Always include an escape hatch like "other" or "unclear". The craft here is its own discipline, and how to write judge questions covers it properly.
-
Label 30 rows yourself (5 minutes). Before any API call, open the file and label a random 30 by hand. This is your tiny gold set. Skipping it is how people end up trusting 5,000 verdicts they never checked.
-
Batch the rows (5 minutes of setup, then wait). Per ecosystem documentation, the model is addressed as typesafe-ai/jev through the Vercel AI Gateway, and the native shape is the Choice API: a question plus candidate answers in, a probability per choice out. Exact syntax lives at docs.typesafe.ai, not here. The loop looks like this:
# pseudocode, not real API syntax
for row in read_csv("comments.csv"):
result = jev.choose(
question = CATEGORY_QUESTION,
text = row["comment"],
choices = ["billing", "bug", "feature_request", "praise", "other"],
)
row["category"] = result.top_choice
row["confidence"] = result.top_probability
write_csv("comments_labeled.csv")
Keep the confidence column. It's the most useful thing in the output and the first thing people throw away.
- Score against your 30 (3 minutes). Compare Jev's labels to yours. If agreement is poor, the question is usually the problem, not the model: tighten category definitions and rerun. Rerunning costs cents, so iterate freely.
Use the confidence column, not just the label
Sort by confidence ascending. The bottom slice is where disagreements cluster, and it's the only part you need to eyeball. Pick a threshold (say, anything under 0.7 goes to a "review" tab), then hand-check that tab. This is the same cascade logic the bigger pipelines use, scaled down to a spreadsheet.
The honest objection: "If I still have to review some rows, what did I save?" You review a slice instead of every row, and you know which slice. That's the whole trade.
Batch classification tips from builds that worked
Split multi-part questions. The X archive run asked eight separate questions per post rather than one compound one, as reported. Eight simple verdicts beat one question that tries to do everything.
Stay under the choice cap. Per ecosystem documentation, choice sets cap around 255 options. If your taxonomy is bigger, stage it: broad category first, subcategory second.
Don't ask for prose. Jev doesn't generate text. If you also want a one-line summary per row, that's a chat model's job in a separate pass. Decision model for the closed column, generative model for the open one.
Once you've got labels, the next moves (training a model on them, building a real labeling workflow) belong to the AI data labeling playbook.
Frequently asked questions
How big a CSV can I classify with Jev?
Builders report runs from hundreds to tens of thousands of verdicts in minutes, such as roughly 26,000 in under nine minutes as reported. Rate limits and batching rules are set by TypeSafe AI, so check docs.typesafe.ai before a large run.
Do I need to write code to label a spreadsheet?
A short script is the simplest route today, though the DuckDB extension build lets you classify rows from a query instead. If you prefer no-code tools, the webhook pattern in Jev automations works row by row.
How accurate will the labels be?
That depends on your question and your data, and nobody can give you a universal number because no official benchmarks exist. Score it yourself against a hand-labeled sample, as the getting-started guide recommends.
Can Jev assign more than one label per row?
Ask separate yes/no questions per label instead of one multi-select question. Each gets its own confidence score, which makes review much easier.
Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.