Blog / Monitoring / FIG. 109
Verdict Logging and Audit Trails
Build an AI decision audit trail: what each verdict log record needs, what to leave out, and how to answer who decided and why. Not legal advice.
Six months after launch, someone will ask why a specific item was routed, rejected, or flagged on a specific day. It might be a customer, an auditor, a regulator, or your own team chasing a metric that moved. An AI decision audit trail is what lets you answer with a record instead of a shrug. This page owns the logging itself: what each verdict record needs, what it must not contain, and how to structure it so "who decided, and why?" has an answer. Examples use Jev, TypeSafe AI's decision model, but the discipline applies to any judge.
Scope: the broader operational playbook (caching, monitoring, fallbacks) is Jev in production, and data-minimization architecture is Jev security and privacy. This page covers only the record.
Why verdict logs are different from app logs
Most LLM logging best practices were written for chat: prompts, completions, token counts, latency. That's the core of LLM observability, and you still need it. Application logs answer "what did the system do?" Verdict logs have to answer "what did the system decide, on what evidence, under which rules, and who acted on it?" Two properties make that harder than it sounds.
First, re-asking is not replaying. Decision-model verdicts can occasionally differ on identical inputs, and questions, thresholds, and models change over time. If your audit story is "we'll re-run it," you'll get today's answer to last spring's question. The log is the record.
Second, the verdict is rarely the whole decision. A probability becomes an action through a threshold, a routing policy, and sometimes a human. An audit trail that stores only the model's answer can't explain the outcome.
What every verdict record needs
One record per question asked, linked into one decision per item. In pseudocode (a schema shape, not any vendor's format):
# PSEUDOCODE: a verdict log record, illustrative only
verdict_record = {
decision_id, item_ref, # reference or hash, not the raw payload
question_id, question_version, # the exact wording that was asked
choice_set_version,
model_id_as_reported, # whatever the API reports, stored verbatim
probabilities, # per choice, as returned
threshold_version, policy_version,
outcome, # auto_accept | escalated | rejected
acted_by, # system | reviewer id
human_override, override_reason,
timestamp
}
The fields that get skipped and later missed: question version (without it, you can't tell whether a metric moved because the world changed or because someone reworded the question; see question versioning), threshold and policy versions (the same probability can mean different actions in different months), and acted_by plus overrides (the line between "the model flagged it" and "a person decided it" is exactly what an auditor will ask about).
What the record should leave out
Verdict logs are both your audit trail and a second copy of whatever you logged. Store an item reference, a hash, or a redacted excerpt rather than the raw input wherever the source system already keeps the original. Apply your normal retention classes to verdict logs explicitly; they are data. Gate access like the source data. The minimization patterns that keep payloads small in the first place are the security page's territory.
A useful precedent from the directory: jevcache keys its local decision cache on model, schema, and state, with redaction and canonicalisation before hashing, for cheaper repeats and deterministic replay in CI, per its entry. The same idea applies to logs: canonicalize and redact first, then hash, so identical inputs are traceable without the log becoming a data lake of sensitive text.
Designing the AI decision audit trail for "who decided?"
For anything consequential, design the log so the answer is unambiguous. Two cataloged builds show the posture. jev-logtriage has Jev score collapsed log batches for noise, severity, and whether an operator should act, then plain code maps the answers to suppress, watch, review, notify, or page; per its entry, nothing is executed. The verdict informs; code applies policy; a person acts. The Alan AI review console screened 3,518 internship applications in 4 minutes 54 seconds and surfaced 20 for human review, as reported. In a hiring context, the log that matters records that a human made each decision, which items the model surfaced, and which question version did the surfacing.
Compliance framing (not legal advice)
This section is not legal advice, and we mean it. If you work in a regulated domain (hiring, lending, insurance, healthcare, education), your counsel decides what records you must keep and for how long. What we can say is architectural: regulators and auditors tend to ask whether a decision procedure was documented, applied consistently, reviewable by a human, and correctable. Versioned questions serve as your documented procedure. Logged probabilities, thresholds, and policies show consistent application. Human decisions and overrides, recorded with who and why, show review. Keep model verdicts advisory wherever someone could reasonably ask "which person decided this?", and never let an irreversible action execute on a lone verdict.
TypeSafe AI's own retention, residency, and certification commitments are theirs to state at docs.typesafe.ai; your vendor review should cite them directly.
Frequently asked questions
What should an AI decision audit trail include?
Per verdict: an item reference, the exact question version, the returned probabilities, the threshold and policy versions, the resulting action, and whether a system or a named person acted, with any override and reason. That set lets you reconstruct why an outcome happened.
Should I log the raw input text?
Only where you must; prefer references, hashes, or redacted excerpts, and apply retention rules to verdict logs like any other sensitive data. The security and privacy page covers minimization.
Can I re-run old inputs instead of storing verdicts?
No, not for audit purposes. Verdicts can vary and questions, thresholds, and models change, so replayability comes from the log, not from re-asking.
Is logging enough for compliance?
It's necessary, not sufficient, and this isn't legal advice. Human review of consequential decisions and a documented, versioned procedure matter as much as the records.
Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.