Blog / Monitoring / FIG. 159
LLM Monitoring: The Working Guide
LLM monitoring tracks errors, latency, cost, and output quality in production, and alerts when they shift. The signals, thresholds, and weekly routine.
LLM monitoring is the set of production checks that tell you when a model-backed feature is getting slower, costlier, flakier, or worse, ideally before your users do. It's the alerting half of LLM observability: fewer signals, watched continuously, each with a threshold and an owner.
The four signal families
Reliability. Error rate, timeouts, rate-limit responses, and how often fallbacks activate. A fallback that fires silently for a week is an outage you didn't notice (uptime and fallbacks).
Latency. Median and 95th percentile per call and per pipeline run. The 95th percentile is where users feel it.
Cost. Spend per day, per feature, and per thousand items, plus the retry ratio. Retries inflate real cost over quoted cost without anyone deciding to spend more. Cost metering has the full spec.
Quality. Sampled evaluator scores, label distribution shifts, escalation and unclear rates, and agreement on a canary set. This is the family teams skip, and the one that catches the expensive problems.
Setting thresholds that don't cry wolf
Baseline for two weeks before alerting on anything. Then alert on relative change rather than absolute noise: "escalation rate up 50% against its 14-day average" beats "escalation rate above 8%."
For quality, alert on distribution changes. If a label's share moves outside its normal band, something changed upstream: new input types, a broken integration, or model drift.
Every alert gets an owner and one line of runbook. An alert nobody owns trains everyone to ignore alerts.
Monitoring judge pipelines specifically
When the pipeline is a judge, like a Jev verdict layer, three signals matter most:
- Escalation rate is the canary. A rise usually means inputs drifted or a question broke.
- Unclear rate rising means new kinds of input the questions weren't written for.
- Canary agreement. Re-judge a fixed set on a schedule; a drop means investigate now, per drift monitoring.
Log the question version with every verdict, so step changes line up with the edits that caused them (question versioning).
The weekly routine
- Read a sample of escalations and misroutes.
- Check the cost-per-thousand trend.
- Re-run the canary set.
- Add any new edge cases to the gold set.
- Fix or version any question that caused repeated misroutes, per the human-in-the-loop rules.
Thirty minutes a week, and it's the cheapest insurance in the stack.
Frequently asked questions
What's the difference between LLM monitoring and observability?
Monitoring alerts on known signals. Observability captures enough detail to explain problems you didn't anticipate.
What should I alert on first?
Error rate, spend spikes, and shifts in escalation rate. Those three catch most production problems early.
How do I monitor output quality without reading everything?
Sample outputs, judge them with closed questions, and watch the distributions over time.
How often should canary checks run?
Daily for high-volume pipelines, weekly for low-volume ones, and after every question or model change.
Do I need a dedicated LLM monitoring tool?
Your existing metrics stack plus structured logs covers most needs. The tools comparison helps when you outgrow it, and what Jev is covers the judge model.
Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.