Blog / Builds & people / FIG. 91
Jev for PMs: Feedback Triage at Last
AI product feedback analysis that reads every ticket, review, and call note: feature request triage by area, pain, and segment for pennies.
Every product manager has a feedback graveyard. A Slack channel called #feedback, a spreadsheet someone started in Q2, a tag in the support tool nobody applies consistently, and a board of feature requests where the loudest customer wins. The feedback exists. It just isn't read, because reading all of it is a full-time job nobody has.
AI product feedback analysis changes that math, and not because a model "summarizes themes." Summaries are where nuance goes to die. The better move is to ask every piece of feedback the same closed questions and count the answers. Jev, the decision model from TypeSafe AI, does exactly that job and nothing else (primer). This page owns the PM feedback loop: from raw text to a triaged, countable backlog.
Why AI product feedback analysis should ask, not summarize
Theme extraction gives you "users want better reporting" with no denominator. Questions give you a table, something like: 212 items this month, 41 about reporting, 29 of those describe a workaround, 18 come from accounts on your top plan. That's a roadmap argument, not a vibe.
It's also the same technique behind the sturdiest sentiment work: replace a single positive/negative gauge with specific questions about what the text says. For PMs, that means pain, area, and segment instead of mood.
The feature request triage battery
Run every item, from tickets, app reviews, sales call notes, community posts, and NPS comments, through a battery like this:
- Is this a feature request, a bug report, a question, or praise? (choice)
- Which product area does it concern? (choice from your own area list, plus "other")
- Does the user describe a workaround they currently use? (yes/no)
- Does the user say this blocks their work or purchase? (yes/no)
- Is a competitor mentioned as having this? (yes/no)
Then join the verdicts to customer data you already have: plan, company size, revenue. The model never needs to know who's paying; your join does.
Two pieces of craft matter. First, your area list should mirror how your team is organized, so every count has an owner. Second, "workaround described" is gold: users who built a workaround have proven the need with their own time.
Treat it like labeling, because it is
Feedback triage is data labeling with a product hat on, and the labeling page's rules apply: check a human-labeled sample before trusting the counts, version your questions, and store probabilities so low-confidence items can go to a person.
Duplicates are the other trap. "Export to CSV," "download as spreadsheet," and "I need this in Excel" are the same request. Group them with a pairwise question on candidates that share an area ("do these two items request the same capability?") before you count, or your top request will be split into five mid-sized ones.
The weekly loop
- Ingest the week's feedback from every source into one table.
- Judge each item with the battery; store answers and probabilities.
- Review a sample of 20 items by hand, especially low-confidence ones.
- Count by area, pain, and segment; compare with last week.
- Read the actual quotes behind the top three movers. The counts tell you where to look. The quotes tell you what to build.
Step 5 is the one that keeps this honest. A PM who ships from counts alone is shipping from a model's reading of the product. A PM who reads 30 quotes selected by counts is shipping from customers.
What the receipts show
The best evidence that the economics work comes from outside product teams. Ian Nuttall ran 3,282 of his X posts through Jev with eight questions each, about 26,000 verdicts, for $0.1282 in 8 minutes 34 seconds, as reported (build). Swap posts for feedback items and the question set for your battery, and a year of feedback costs less than the coffee during the roadmap meeting.
A closer analog is the Hacker News comment verdicts build, which gives each comment a support, critical, or neutral stance plus substance and quotability scores. "Substance" and "quotability" are exactly the filters a PM wants before pasting customer voice into a spec.
There are no official Jev benchmarks for feedback triage; the numbers above are builders' own, on their own data.
What it won't do: Jev doesn't write your synthesis doc, your PRD, or the "what we heard" summary. For prose, hand the counted, quoted results to a chat model or write it yourself. And it doesn't decide the roadmap. Frequency is one input; strategy, cost, and conviction are others.
Frequently asked questions
What is AI product feedback analysis?
Asking every piece of customer feedback the same closed questions, such as type, product area, and whether it's blocking, then counting the answers. It produces a countable backlog instead of a theme summary.
How does feature request triage handle duplicates?
Run a pairwise "same capability?" question on requests that share an area before counting. Otherwise one popular request gets split across many phrasings.
Can I trust the counts?
Check a human-labeled sample each week and send low-confidence items to a person. The data labeling page covers how to validate model labels.
Does this replace user interviews?
No. It tells you which topics deserve interviews and which quotes to read first. Talking to users is still the job.
Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.