Blog / Questions / FIG. 128
Can I Fine-Tune Jev?
Can you fine-tune Jev? No documented fine-tuning path exists. Your questions are the tuning surface: boundary clauses, optimizers, and layers on top.
Not that we've seen documented. Jev, TypeSafe AI's decision model, is hosted-only with unpublished weights, and no fine-tuning or custom-model offering appears in the ecosystem documentation at the time of writing. If that changes, docs.typesafe.ai will say so before anyone else does.
Here's the part that matters more: for a decision model, fine-tuning is rarely the lever you actually need. Your questions are the tuning surface.
Why questions do the job fine-tuning does elsewhere
With a chat model, fine-tuning teaches behavior that's hard to describe in a prompt: tone, format, house style. A decision model has no tone or format to teach. Its output is already one of your allowed answers. What's left to "tune" is where the boundaries fall, and boundaries can be written down.
That's the boundary-clause point. Most disagreements between Jev and your reviewers live at an edge the question never defined, and a single parenthetical fixes them: "Does this review mention a product defect? (Cosmetic damage counts; disliking the color does not.)" The fine-tuning vs prompting guide puts it bluntly: teams that fine-tune first often buy a training pipeline to discover what a boundary clause would have fixed. The method for writing those clauses lives in the judge-questions guide.
Three ways to get a "custom Jev" anyway
Optimize the question automatically. The jev-align library takes your labeled examples, lets an optimizer (GEPA) rewrite the question, and returns a Jev function with a known error rate, per its author. It's the closest thing to fine-tuning, applied to text you can read and diff.
Fit a layer on top. One post-scoring build asks 61 questions per draft for $0.0004, as reported, and its author says the scorer on top was fitted on 9,481 posts from 207 creators (build). We haven't verified its accuracy claims; the architecture is the point: Jev's answers become features, and your model learns the weights.
Graduate to a trained model. Log verdicts, audit them, and train a classic classifier on the result when volume justifies it, as covered in the LLM vs traditional ML guide.
Frequently asked questions
Does TypeSafe AI offer fine-tuning for Jev?
No fine-tuning offering is documented in the sources we track at the time of writing; check docs.typesafe.ai for current options.
How do I customize Jev for my domain?
Write operational questions with boundary clauses for the edge cases your reviewers disagree on, then calibrate against 50 to 100 labeled cases. The judge-questions guide is the method.
Can I train a model on Jev's outputs?
Yes. Using logged, human-audited verdicts as training labels is a common pattern; see whether Jev learns from feedback for how that loop works.
When would I need a fine-tunable model instead?
When your task is generative (style, format, voice) or your data can't leave your infrastructure; that's open-model territory, and TypeLLM is the ecosystem's pattern for it.
Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.