shipwithjev

Blog / Calibration / FIG. 162

LLM Evaluation Metrics: The Catalog

LLM evaluation metrics grouped by task: precision and recall, BLEU and ROUGE, faithfulness, pass@k, judge agreement, and calibration.

LLM evaluation metrics are the numbers you use to decide whether a model, prompt, or pipeline is good enough to ship. No single metric fits every task. The right one depends on whether your output is a label, a piece of code, an answer grounded in documents, or free-form text.

This catalog groups the common metrics by task and says what each actually measures. For building the suite that uses them, see the LLM evals guide.

Metrics for classification and closed answers

Accuracy is the share of correct answers. It misleads on imbalanced data, where labeling everything as the majority class scores well.

Precision, recall, and F1 per class. Precision: of the items labeled X, how many were X? Recall: of the real X items, how many were caught? F1 balances the two.

Exact match scores short answers as right or wrong, character for character after normalization.

Calibration asks whether confidence means anything: are answers at 80% confidence right about 80% of the time? Expected calibration error (ECE) summarizes the gap. It matters most when thresholds decide what a human sees, per probability outputs explained.

Metrics for generated text

BLEU and ROUGE measure word overlap with a reference text. They're cheap and reproducible, but weakly tied to quality for open-ended writing, since a good answer phrased differently scores poorly.

Embedding-based similarity (BERTScore and relatives) compares meaning rather than exact words. Better, still bound to a reference.

Perplexity measures how surprised a model is by text. It's useful for comparing language models, not for scoring your app's outputs.

Metrics for grounded answers (RAG)

  • Faithfulness: is every claim supported by the retrieved context?
  • Answer relevance: does the answer address the question?
  • Context precision and recall: did retrieval fetch the right passages and skip the wrong ones?

These are usually scored by splitting an answer into claims and asking a closed question per claim. RAG evaluation covers the method.

Metrics for code and agents

pass@k is the probability that at least one of k generated solutions passes the unit tests. It's the standard metric for code generation.

For agents: task success rate verified against the end state, steps to completion, and cost and latency per task. AI agent evaluation covers the full playbook.

Metrics for judges themselves

When a model is the grader, grade the grader:

  • Agreement with human labels, and Cohen's kappa, which corrects agreement for chance.
  • Escalation and unclear rates as health signals.
  • Bias checks, especially self-preference when a model judges its own family's output (self-preference bias).

LLM-as-a-judge covers the setup these metrics audit.

Frequently asked questions

Which LLM evaluation metric is best?

The one that matches your output type. Most production suites pair a task metric with a human-agreement check.

Are BLEU and ROUGE still useful?

As cheap baselines for translation and summarization, yes. For open-ended quality, pair them with judged metrics.

What's a good agreement rate for an LLM judge?

Compare it with how often your human labelers agree with each other. A judge that matches that rate is as good as your labels allow.

How do I measure faithfulness?

Split the output into individual claims and ask, for each one, whether the source text supports it.

Which metric catches drift?

Agreement on a fixed canary set, re-run on a schedule. The judge model is covered in what Jev is.

Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.