shipwithjev

Blog / Recipes / FIG. 83

Build a Churn-Flag Pipeline From Support Text

Churn detection from tickets as a working pipeline: cancellation language alerts, account rollups, and a human on every save play.

Churn detection from tickets usually fails in a boring way: nobody reads the tickets in time. The customer wrote "we're evaluating alternatives" in a Tuesday reply, the agent fixed the actual bug, and the sentence went into the archive with everything else. This recipe builds the pipe that catches it, using Jev, the decision model from TypeSafe AI: every inbound message gets a few closed questions, flags roll up per account, and a human gets a short daily list.

Which signals matter is the churn prediction page's job; this page is plumbing. If you need the one-paragraph version of what Jev is: it answers typed questions about text, with probabilities, and it does not write prose.

The pipeline has four stages, each dumb on purpose:

  1. Ingest: new tickets, replies, and chat transcripts land in a queue. If you already run support triage, you have this; churn flags ride the same stream as extra questions.
  2. Judge: each message gets three to five closed questions.
  3. Roll up: flags aggregate per account over a rolling window.
  4. Alert: accounts crossing a threshold land in a CSM queue with the triggering sentences quoted.

No stage takes an action on the account. The whole recipe is built around that rule.

Step 1: write the churn detection questions

Keep the battery small and operational. For a first pipeline:

  • Does the message mention canceling, downgrading, or not renewing? (yes/no)
  • Is a competitor or an evaluation of alternatives mentioned? (yes/no)
  • Is a problem described as recurring or still unresolved? (yes/no)
  • Which fits best: routine, frustrated, at-risk, leaving? (choice)

Resist asking "will this customer churn?" It asks the model to forecast from one message, which is a coin flip with extra steps. Ask what the text says, and let the rollup do the predicting. The judge-questions guide covers the craft.

Step 2: call the judge, keep the probabilities

Per ecosystem documentation, Jev is reached through the Vercel AI Gateway as typesafe-ai/jev, and choice answers come back with probabilities. Check docs.typesafe.ai for real client syntax; the block below is pseudocode.

# pseudocode, not real API syntax
for message in new_messages:
    verdict = judge(message.body, CHURN_QUESTIONS)
    store(message.id, message.account_id, verdict.answers, verdict.probabilities)

Store the probabilities, not just the winning answer. A "leaving" at 0.51 and a "leaving" at 0.97 are different events, and you'll want that distinction when you tune in week two.

Cost is the least interesting part. Riley Brown classified 500 emails for 3.5 cents, as reported (build). Do the math on your own volume, but expect it to be unexciting.

Step 3: roll up to accounts

A single angry ticket is weather. Three signals across two contacts in a fortnight is climate. The rollup is SQL or a spreadsheet, not a model:

# pseudocode
account_risk = sum over the last 14 days of:
    2 x cancel_mention      (probability at least 0.8)
  + 2 x competitor_mention  (probability at least 0.8)
  + 1 x recurring_problem
  + 1 x tone is at-risk or leaving
flag the account when account_risk crosses your threshold

The weights should be embarrassing at first. Start with a threshold that flags too many accounts, have CSMs mark each flag "real" or "noise" for two weeks, then adjust. That marked list becomes your gold set, the only honest way to know whether the pipeline works.

Step 4: alert a human, not a workflow

The alert is a queue item: account, score, and the exact sentences behind each flag. The quote matters more than the score. A CSM who reads "we've been talking to another vendor since the outage" acts faster than one staring at "risk: 6."

There's a live receipt for this shape: one builder has Jev watch more than 25 customer WhatsApp groups and decide whether something needs his attention, such as an urgent problem or an open order; only then does an LLM write him the message, as reported (build). Decision model decides whether to wake someone, a generative model writes the note, a person does the work.

What the pipeline must never do on a lone verdict: apply a retention discount, change a plan, or email the customer. Those are real actions with real costs, and a classifier reading one message doesn't get to trigger them. Flags open conversations; people decide what to say.

Week two: tune, then widen

With two weeks of CSM labels, three adjustments usually matter:

  • Threshold: move the probability cutoff until the noise rate is tolerable for the people reading the queue.
  • Wording: if "recurring problem" fires on every ticket containing "again," tighten the question.
  • Sources: add call transcripts and NPS comments, same questions, same rollup.

Then check flags against last quarter's actual churn. If flagged accounts don't leave more often than unflagged ones, your questions measure mood, not intent. Back to Step 1.

Frequently asked questions

How does churn detection from tickets work?

A decision model answers closed questions about each ticket, such as whether it mentions canceling or a competitor, and a simple rollup scores accounts over time. Humans review flagged accounts; the churn prediction page covers which signals matter.

Can I set up cancellation language alerts without a data team?

Yes, if you can move tickets into a script or no-code workflow and write results to a table. Judging is one call per message, and the rollup is arithmetic.

Should the pipeline trigger retention offers automatically?

No. Offers, plan changes, and customer emails are consequential actions and should never fire on a single model verdict. The pipeline flags; a CSM decides.

How do I know the flags are accurate?

Have CSMs label every flag as real or noise, then compare flagged accounts against actual churn. There are no official Jev benchmarks, so your own labeled data is the evidence; the one-day eval suite recipe shows how to formalize it.

Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.