shipwithjev

Blog / Monitoring / FIG. 152

LLM Observability: What to Watch and Why

LLM observability means seeing what your model calls actually did: traces, inputs, outputs, cost, latency, and quality. What to capture, and the traps.

LLM observability is the ability to see what your language-model calls actually did in production: which inputs went in, what came out, how long it took, what it cost, and whether the output was any good. Traditional observability watches servers. LLM observability has to watch behavior, because a model can return a fast, cheap, perfectly formatted answer that's wrong.

That last part is why the usual dashboards don't save you. Every light can be green while the product quietly gets worse.

Observability vs monitoring

Monitoring watches known signals and alerts on them: latency, error rate, spend. It answers "is something wrong?" The working guide is LLM monitoring.

Observability captures enough detail to answer questions you didn't plan to ask: why did this user get that answer on Tuesday? It answers "what went wrong, and why?"

You need both. Monitoring pages you, and observability tells you what to do about it.

What to capture for LLM observability

  • Traces and spans for every step of a chain or agent: retrieval, tool calls, model calls.
  • Inputs and outputs, redacted where they contain personal data.
  • Versions: model, prompt, and for judge pipelines, question version.
  • Tokens, latency, and cost per call and per pipeline run (cost metering spec).
  • Quality signals: user feedback, evaluator scores, label distributions.

Structured logs with a request ID tying all of it together get you most of the way before you buy any tool.

The details that save you later are the boring ones. Log the exact model identifier, not "the fast one," because providers ship silent revisions. Store the rendered prompt as sent, after templating, since the template in git isn't what the model saw. For retrieval steps, keep the document IDs and their order, not just the final answer. And record which branch the code took after the model replied: retried, fell back, escalated, or shipped.

One field people skip: the user-visible outcome. Did the answer get shown, edited, thumbs-downed, or abandoned? Without that, you can explain what the model did but not whether anyone cared.

The quality layer is the hard part

Latency and cost are just numbers. Quality needs judgment, and nobody reads every output.

Sampled evaluation. Run closed-question judges over a sample of production outputs daily, the pattern LLM-as-a-judge covers. Closed questions ("does the answer cite a retrieved document, yes or no?") are cheap to run and easy to trend. Open-ended "rate this 1 to 10" scores drift and argue with each other.

Canary sets. Re-judge a fixed set of known items on a schedule. When agreement drops, something upstream changed, per drift monitoring.

Distribution watching. If one label's share doubles overnight, investigate before users complain.

The useful trick is reusing the same checks in two places. The jevals build turns agent checks into typed questions so one definition runs on saved traces offline and as a runtime guardrail, per its author. Your pre-release tests and your production quality signal then measure the same thing, which is the whole point of prompt testing before a change ships.

A worked example: tracing one bad answer

A hypothetical, but a common one: a support bot tells a customer their plan includes a feature it doesn't. Monitoring saw nothing: latency normal, no errors, cost flat.

With observability, you pull the request ID from the complaint and open the trace. Retrieval returned three docs, and the top one is last year's pricing page. The prompt version is current, the model version is current. So the model did its job on bad context.

Next question: how many other answers used that stale doc? Because you logged document IDs per call, it's one query. You find 140 affected conversations this week, fix the index, and add a canary item ("does the answer match current pricing?") so the next stale doc gets caught by a check instead of a customer. Retrieval-specific checks like this one are what RAG evaluation is about.

Without the document IDs, that same incident is a week of guessing.

Observability for judge pipelines

When the thing in production is itself a judge, a Jev pipeline for instance, watch the verdicts: probability distributions, escalation rate, unclear-answer rate, and canary agreement. A rising escalation rate usually means inputs drifted or a question broke. Log each verdict with its evidence and question version, as the audit trail spec describes, and production posture covers the rest.

Pin question versions so a log line always maps to one exact question. The workbench build publishes immutable judgment versions and pins callers to them, per its author, which is the same discipline question versioning argues for.

Decision models also show up on the other side of the pipe, reading the observability data itself. Tocsin grouped 22.8 million log lines into repeating patterns and asked Jev one question per pattern, finishing in six minutes for $0.64, as reported. That's triage, not diagnosis: it narrows what a human reads, and paging decisions still follow a policy you wrote.

Rollout checklist and common mistakes

  1. Day one: request IDs, model and prompt versions, tokens, latency, cost, redacted inputs and outputs.
  2. First multi-step chain: spans for every retrieval, tool call, and model call.
  3. First month: a daily sampled evaluation with closed questions, plus a small canary set.
  4. First incident: write the query you wished you had, and make sure the data for it exists.
  5. When logs outgrow grep: reach for a platform. The tools comparison sorts the options by type.

The common mistakes are predictable. Logging prompts but not versions, so you can't tell which change caused a regression. Keeping raw user inputs forever with no retention policy. Sampling so little that a 5% failure mode never shows up. And building beautiful dashboards nobody owns, which is monitoring cosplay. Every signal needs a person who acts when it moves.

Frequently asked questions

What is LLM observability?

Capturing traces, inputs, outputs, versions, cost, latency, and quality signals so you can explain model behavior in production.

How is LLM observability different from APM?

APM watches system health. LLM observability adds behavior and output quality, which APM can't see.

Do I need a tool for LLM observability?

Not to start. Structured logs with IDs and versions go a long way; tools add tracing views and evaluation workflows, compared in the observability tools guide.

What's the most overlooked observability signal?

Output distribution shifts. A label's share jumping overnight often catches drift before any user report.

Does observability create privacy risk?

Yes, since logs hold user inputs. Redact, minimize, and set retention limits. The judge model itself is covered in what Jev is.

Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.