Blog / Evals & judging / FIG. 80
Turn Survey Open-Ends Into Data
Analyze open-ended survey responses with Jev: build a codebook, code every answer as verdicts, check a hand-coded sample, then chart it.
Every survey has a "tell us more" box, and every survey team has a folder of those answers nobody coded because it would take a week. This recipe shows how to analyze open-ended survey responses with Jev, TypeSafe AI's decision model: turn a codebook into closed questions, run every response through them, and end up with columns you can chart.
The broader workflow of model-first labeling, including how to audit labels and reuse them, belongs to AI data labeling. This page is the survey-specific recipe.
The receipts
The closest large-scale run is Ian Nuttall's: 3,282 posts, eight questions each, roughly 26,000 verdicts for $0.1282 in 8 minutes 34 seconds, as reported (build). That's structurally identical to coding survey responses: many short texts, several fixed questions each.
More directly on topic, one builder rebuilt a survey-research workflow on Jev with 200 samples and 12 questions per survey, and reports it ran 2x faster and cost 85% less than Gemini 3.5 Flash-Lite (build). Both are builder-reported. There are no official Jev benchmarks.
The recipe
-
Draft a codebook from a sample (30 minutes). Read 50 random responses. Write down the themes you actually see, not the ones you hoped for. Aim for 6 to 12 codes, plus "other" and "no substantive answer".
-
Turn each code into a yes/no question. Survey responses often mention several themes, so multi-select beats single-choice. One question per code:
- "Does this response mention price or cost as a concern? YES / NO / UNCLEAR"
- "Does this response describe a missing feature? YES / NO / UNCLEAR"
- "Does this response mention customer support, positively or negatively? YES / NO / UNCLEAR"
Each code gets its own confidence score, which you'll want later. Question wording is where accuracy is won or lost, so be specific about what counts.
-
Hand-code 50 responses (40 minutes). A different 50 from step 1. This is your check set. If two people on the team can code them independently, even better: their disagreement rate tells you how fuzzy your codes are before any model gets involved.
-
Run every response.
# pseudocode, not real API syntax
for resp in survey_responses:
for code in CODEBOOK:
v = jev.choose(code.question, resp.text, ["YES", "NO", "UNCLEAR"])
resp[code.name] = v.top_choice
resp[code.name + "_conf"] = v.top_probability
export(survey_responses, "coded_responses.csv")
Per ecosystem documentation, the model is typesafe-ai/jev via the Vercel AI Gateway with per-choice probabilities. Check docs.typesafe.ai for current syntax and limits.
-
Compare against your hand-coded 50. Look per code, not overall. Usually most codes agree well and one or two don't. The weak ones are almost always ambiguous definitions; rewrite them and rerun only those codes.
-
Chart and cross-tab. Now your open-ends are columns. Cross them with the closed survey questions: which themes show up among detractors, which among new customers, which by plan tier. This is where coded open-ends earn their keep.
Sentiment is one column, not the whole analysis
The tempting shortcut is to run sentiment over everything and call it done. Resist it. "Negative" tells you nothing about what to fix. Code themes first, then add a sentiment or intent question if you need it; the approach for that lives in sentiment analysis with LLMs.
What Jev doesn't do here
Jev doesn't write the summary. It won't produce "Customers mostly want better onboarding" or pick representative quotes with commentary. That's a generative job for a chat model or, better, the analyst. The clean split: Jev codes every response into closed fields, then a human (optionally helped by a chat model) writes the story from the counts.
It also doesn't discover your codebook. Step 1 is human work. If you want a machine assist, a chat model can propose candidate themes from a sample, and you then turn the good ones into questions.
Code survey responses with AI, honestly
Report method with the results. "Responses were coded by a decision model against a 10-code codebook, validated on a 50-response hand-coded sample" is a sentence your readers deserve.
Watch low-confidence rows. Review anything under your threshold before it lands in a chart.
High-stakes surveys need human sign-off. Employee surveys, patient feedback, and anything feeding pay, hiring, or grading decisions: use the coding to speed humans up, keep a person accountable for conclusions, and keep an audit trail of codebook versions.
Frequently asked questions
How accurate is AI coding of open-ended survey responses?
It depends on how clearly your codes are defined, and no official benchmark exists to quote. Measure it yourself against a hand-coded sample, per code, before trusting the counts.
Can Jev summarize survey responses?
No, Jev returns decisions rather than text. Use it to code every response into themes, then write the summary from the counts, with a chat model helping if you like.
How much does it cost to code a survey?
Builder-reported runs put tens of thousands of verdicts at well under a dollar, such as about 26,000 for $0.1282 as reported. Multiply responses by codes for your own estimate, and see AI data labeling for the wider cost picture.
Should I use single-choice or multi-select codes?
Multi-select, as one yes/no question per code, since real responses often mention several themes. Single-choice forces the model to drop information.
Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.