Blog / Recipes / FIG. 120
Cost Monitoring for Verdict Pipelines
LLM cost monitoring for verdict pipelines: what to meter per call, the ratio that catches runaway loops, and API spend alerts worth setting.
Expensive models train you to watch the bill. Cheap ones train you to stop looking. That's the trap with verdict pipelines: when a single decision costs a fraction of a cent, nobody instruments spend, and the first cost signal anyone sees is an invoice.
LLM cost monitoring for decision pipelines isn't about the price of a verdict. It's about noticing, within hours, when the number of verdicts stops matching the work. The reduce-costs playbook owns the levers that shrink a bill: cascades, caching, prompt diets, batching. The production guide puts "meter cost per pipeline per day" on its ops list. This page owns the metering itself: what to log, what to compute, and which API spend alerts actually catch problems.
Why cheap pipelines need more metering, not less
Consider an illustrative scenario, a composite of a common failure shape rather than a reported incident. A nightly job classifies new support tickets with eight questions each. One night the ticket API starts returning errors for a subset of records. The job's retry logic re-enqueues failures, and each retry re-asks all eight questions. The failures never clear, so the queue grows every night. Nothing crashes. Verdicts still land. On day one the extra spend is invisible; by day thirty the pipeline is making dozens of calls per ticket, and the bill has multiplied while the ticket volume stayed flat.
At frontier prices, someone would have caught that in a week because the spend line would have jumped. At decision-model prices, the multiplied bill can still look small enough to ignore. That's the point: cheap calls don't make runaway loops less likely, they make them less visible. The loop costs little per call and plenty in total, and it's usually a symptom of a real bug (bad retries, a stuck agent, a cron firing twice) that is also hurting something other than your budget.
What to log on every call
You can't alert on what you don't record. Per verdict call, capture:
- Pipeline and question ID, plus question version, so spend can be sliced by the thing that caused it.
- Item ID, so you can count calls per item.
- Token counts and cost, if the response or gateway reports them. Builders do publish this level of detail: one run of 3,282 posts at eight questions each reported 4,252,330 tokens, $0.1282, and 8 minutes 34 seconds. If your builders can post it, your dashboard can show it.
- Cache hit or miss, so savings are measurable and a cache outage shows up as spend.
- Attempt number, so retries are countable.
- Environment (production, staging, backfill), so a test script can't hide in production spend.
Per-call cost figures and pricing should come from your gateway's billing data and docs.typesafe.ai, not from blog posts; the cost-per-request page collects what builders report, as reported.
The ratio that catches runaway loops
The most useful single number is calls per item. You know what it should be: a pipeline asking eight questions per ticket should average eight, a little less with caching, a little more with occasional retries. When it climbs to twelve, something is re-asking. When it hits forty, something is looping.
This works better than absolute spend alerts, because it's independent of volume. A busy day doubles your spend and leaves calls per item flat, which is fine. A retry loop leaves item volume flat and doubles calls per item, which isn't. Pseudocode for the check, not any real API:
// pseudocode: daily cost checks per pipeline
expected = questions_per_item for this pipeline
actual = verdict_calls_today / distinct_items_today
if actual is well above expected: alert "possible loop or retry storm"
if retry_rate is above its baseline: alert "upstream failures"
if cache_hit_rate falls sharply: alert "cache broken, spend rising"
if spend_today is far above trailing avg: alert "volume or loop, check ratio"
if spend_today hits hard ceiling: stop non-critical jobs, page owner
Spend alerts worth setting
A per-pipeline daily budget with a hard ceiling. Not just a warning email. For non-critical batch jobs, the ceiling should stop the job; a paused backfill is cheaper than a month of loop.
Rate-of-change on spend, compared against a trailing week, to catch volume changes you didn't plan.
Per-hour cost for continuous loops. Some workloads never finish. The launch-week Doom demo picked moves about ten times a second at a reported ~$7 an hour. For game loops, agents, and live monitors, cost per hour is the natural unit, and an agent that forgets to stop is the classic surprise.
Environment split. Backfills and eval runs get their own budget lines so a one-off experiment can't be mistaken for production drift, or vice versa.
Pair cost alerts with verdict alerts. A spend spike on the same day as a verdict distribution shift is usually one bug with two symptoms, and seeing them together cuts diagnosis time in half.
Unit economics, not just totals
Once metering exists, report cost per business outcome: per ticket routed, per document classified, per thousand posts scored. That's the number a finance conversation needs, and it's how you'll know whether a question rewrite or a cascade change was worth it. For Jev pipelines, the per-outcome figure is usually tiny; the reason to track it is that tiny numbers that suddenly aren't are the earliest sign something broke.
Frequently asked questions
What should I monitor to control LLM costs?
Spend per pipeline per day, calls per item, retry rate, cache hit rate, and cost per business outcome, all tagged by question version and environment. Calls per item is the best single early warning.
How do I detect a runaway loop in an AI pipeline?
Compare calls per item against the number of questions the pipeline should ask. A ratio well above expected means something is re-asking, whatever total spend looks like.
Do I need spend alerts if each verdict costs a fraction of a cent?
Yes, more than ever: cheap calls hide loops and retry storms until they've run for weeks. The production guide lists daily cost metering as a core signal for that reason.
Where should cost numbers come from?
Your gateway's billing data and the official docs for pricing. Builder-reported figures are useful for napkin math, as reported, but your own metering is the only number you should alert on.
Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.