shipwithjev

Blog / Cost / FIG. 142

How Builders Actually Measure Cost

How to measure LLM cost the way Jev builders do: run totals, unit rates, hourly burn and same-clock comparisons, with each receipt linked.

Every cost receipt in this directory was measured by somebody, and they didn't all measure the same thing. If you want to measure LLM cost for your own project, the builds made with Jev, the decision model from TypeSafe AI, are a useful field guide: five distinct methods show up again and again, and each answers a different question. This page sorts them, with the receipts that use each one.

All figures are as reported by the builders and linked. Setting up alerts and dashboards for live pipelines is covered in cost monitoring pipelines; the quick estimate before you build is the napkin method in what Jev builds actually cost. This page is the part in between: what to measure once something runs.

Method 1: the whole-run total

Run the job once, read the bill. The simplest and most common receipt.

The best examples report three numbers together. Ian Nuttall's X archive run gives 3,282 posts, 4,252,330 tokens, $0.1282 and 8 minutes 34 seconds, as reported. With tokens, cost and time side by side, anyone can sanity-check the other two. Compare 500 emails for 3.5 cents: a great headline, but on its own you can't tell what drove the cost.

If you only do one thing from this page: when you report a total, report token count and item count with it. That's the core of token cost tracking.

Method 2: the unit rate

Totals don't scale; rates do. Builders running ongoing tools report per unit:

That last word is the detail to copy. Kyle McLaren separates cached from uncached traffic, which is the difference between a worst-case rate and an average one. If your pipeline caches verdicts, report both.

Method 3: burn per hour or per day

For anything that runs continuously, per-item cost is the wrong lens. The Doom demo is reported at about $7 an hour; Bouncer, a policy check on every Claude Code tool call, at about $0.04 a day. These are the numbers that tell you what a month costs.

The TBC leveling agent adds a twist worth borrowing: it reuses decisions it has already made, so its cost per hour falls as the run goes on. If your workload repeats, track cost per hour over time, not just once.

Method 4: the same-clock comparison

The most persuasive receipts compare against an alternative on the same data.

  • 100,000 viral posts: $0.67 for the full corpus, against 214 posts for $0.98 from a frontier model on the same corpus and clock, which the builder puts at roughly 680 times the cost per post.
  • Tocsin: $0.64 for 22.8 million log lines, against an estimated $1,120 for running an LLM over every line.
  • Survey research: 85% cheaper and twice as fast as the model it replaced, per the author.

Note the difference between "measured both" (the viral posts run) and "estimated the alternative" (Tocsin). Both are useful; label which one you did.

Method 5: the whole-pipeline cost

A cascade has more than one model, and the honest total includes all of them. The fraud detection build reports about $0.07 for the full pipeline, including the larger model that handled the unsure cases. Reporting only the cheap model's share would flatter the result.

The same goes for estimated versus billed. The X archive dashboard is careful to call its ~$0.086 an estimated model cost. Estimates are fine. Presenting them as invoices isn't.

A record worth keeping

Pulling the five methods together, this is the shape of a cost record that makes every one of them possible later. Pseudocode, not an API:

// pseudocode: one row per job, not a real schema
job: "label-posts-2026-09-30"
items: 3282
questions_per_item: 8
tokens: 4252330
cost_billed_or_estimated: "estimated"
cost_total: 0.1282
cached_share: 0.0
wall_time_seconds: 514
other_models_in_pipeline: none

What to log per call, and which alerts to set on top of it, is the monitoring guide's territory.

Frequently asked questions

What is the simplest way to measure LLM cost?

Run the job once and record the total alongside item count and token count. The three together let you compute a unit rate and compare against later runs.

Should I track tokens or dollars?

Both: dollars answer "what did it cost," tokens explain why, and prices can change while token counts stay comparable. Builders who report both, like the X archive run, make the most checkable receipts.

How do I compare Jev's cost against another model fairly?

Run both on the same data, in the same window, and report whether the alternative was measured or estimated. Same-clock comparisons are the strongest receipts in the directory.

What costs do people forget to count?

The escalation model in a cascade, retries, and uncached traffic. For catching runaway spend in production, see cost monitoring pipelines.

Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.