Blog / Comparisons / FIG. 99
Jev vs Kimi K3: Partners, Not Rivals
Jev vs Kimi: a decision model and a general model do different jobs. In the Kimi K3 cascade fraud build, the pair scored 96/100 for ~$0.07.
Search "Jev vs Kimi" and you'll expect a winner. There isn't one, because the two don't compete for the same job. Jev is TypeSafe AI's decision model: it answers closed questions (yes, no, which category) with a probability attached, and it doesn't generate text. Kimi K3 is a general-purpose model from Moonshot AI's Kimi line: it reads, reasons, and writes. Asking which is better is like asking whether a triage nurse is better than a surgeon.
The reason this comparison exists at all is one build, and it's the most-cited receipt in the directory: a fraud pipeline where the two work as a team.
The receipt: the Kimi K3 cascade
Hassan gave Jev 100 emails to check for fraud. Jev judged all of them in 1.42 seconds, and only the cases it was unsure about went to Kimi K3. The full pipeline got 96 of 100 right for about $0.07, and he notes the video isn't sped up (build, all as reported).
One run, 100 emails, one builder: not a benchmark, and there are no official Jev benchmarks to compare against. What makes it instructive isn't the exact score; it's the shape. The expensive model never saw the easy cases, and easy cases are most cases.
The general theory of cascades (why they save money, how to tune thresholds, when they're the wrong choice) lives in the LLM routing guide. This page is about why this particular pairing works.
Jev vs Kimi: why the pairing works
Jev's probabilities are the handoff signal. Per ecosystem documentation, Jev returns a probability for each answer choice. A 0.97 "fraud" or a 0.96 "legitimate" is a confident call; a 0.55 is Jev saying "I'm not sure." That uncertainty is exactly the routing key. A generative model's confidence is harder to read from its output, so it makes a worse first tier.
Kimi K3 is good at what the uncertain slice needs. The hard emails are hard for reasons: a legitimate vendor with a changed bank account, a scam that copies a real invoice. Those benefit from a model that can read the whole thread, weigh details, and explain itself in prose a reviewer can check. That's generative reasoning, which Jev doesn't do.
The costs stack in the right order. Jev's per-verdict cost is a fraction of a cent in builder-reported workloads, so running it on everything is cheap. The general model's cost only applies to the slice that earned it.
Here's the shape as pseudocode (not real API syntax; see docs.typesafe.ai for the actual interface):
pseudocode:
for each email:
p = jev.ask("Is this email attempting fraud?", choices = [YES, NO, UNCLEAR])
if p.YES is high: route to fraud review queue
else if p.NO is high: route to normal inbox
else: send to Kimi K3 with full thread
Kimi's answer + reasoning go to the review queue
Notice what the pseudocode doesn't do: it doesn't block a payment, freeze an account, or delete an email. Both tiers produce labels for a queue. In money-adjacent domains, a person makes the irreversible call, with both models' outputs logged as evidence.
When you'd use one without the other
Jev alone when the decisions are closed-set, high-volume, and cheap to get wrong: tagging, sorting, filtering, pre-screening. If a wrong verdict costs a human a second look, the cascade's second tier might not earn its keep.
Kimi K3 alone when every item needs reasoning, a written explanation, or generation, and volume is low enough that per-item cost doesn't matter. Drafting replies, summarizing threads, and open-ended research are its territory, not Jev's.
Both when volume is high and a meaningful slice is genuinely hard. That describes fraud, moderation, support triage, and most evals. The broader lineup of alternatives (small chat models, trained classifiers, rules) is compared in Jev alternatives.
Tuning your own version
Don't copy the fraud build's thresholds; copy its method. Run a few hundred of your own labeled cases through Jev, look at where confident verdicts go wrong, and set the "unclear" band wide enough to catch those. Then measure how often the second tier fixes them. The confidence thresholds guide walks through it. Widen the band and you pay more for accuracy; narrow it and you save money and take more risk. That dial is yours.
Frequently asked questions
Is Jev better than Kimi K3?
They do different jobs. Jev answers closed questions fast and cheaply; Kimi K3 reads, reasons, and writes. The strongest reported setup uses both in a cascade.
What is a Kimi K3 cascade?
Jev judges every item first, and only the items it's unsure about go to Kimi K3. In the fraud build, that pipeline got 96 of 100 right for about $0.07, as reported.
Can I swap Kimi K3 for another model?
The pattern doesn't depend on Kimi specifically; any capable general model can be the second tier. Measure on your own data, because the fraud build's numbers only describe that run.
Should a cascade block fraudulent payments automatically?
No. Both tiers should produce labels and evidence for a review queue; blocking payments or freezing accounts is an irreversible, money-touching action that needs a human decision.
Where do I learn to build one?
Start with the routing guide for the theory and five cascade architectures for variants.
Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.