shipwithjev

Blog / Recipes / FIG. 119

Human-in-the-Loop That Humans Don't Hate

Human in the loop design for verdict pipelines: evidence-attached queues, sane volume, and review queue UX that keeps reviewers sharp instead of numb.

Every responsible AI architecture diagram has a box labeled "human review". It's usually the smallest box on the page and the least designed part of the system. Then it ships as a spreadsheet of IDs, a reviewer opens it on Monday, sees 1,400 rows with no context, and starts clicking "approve" at a pace that would alarm anyone who watched.

That's the failure this page is about. Human in the loop design isn't whether a human is involved; it's whether the human can actually exercise judgment in the time you've given them. The escalation design patterns guide owns which cases reach a person. This page owns what happens when they get there.

Why reviewers end up hating the queue

Talk to anyone who has worked a review queue and the complaints rhyme:

  • No context. A ticket ID and a label, so every item means opening three other tools.
  • No idea what's being asked. Is the reviewer confirming the label, choosing a new one, or deciding an action?
  • Too much volume. When the queue is bigger than the day, review turns into rubber-stamping.
  • No visible effect. Reviewers correct the same mistake every day and nothing upstream ever changes.
  • Automation bias. When the machine's answer is shown prominently and usually right, people stop checking. The loop still has a human in it; it no longer has judgment in it.

Each of these has a design fix.

Attach the evidence, not just the verdict

The biggest single improvement is an evidence-attached queue: every card arrives with what the pipeline knew and why it's uncertain. For a verdict pipeline built on a decision model like Jev, that means:

  • The input itself, rendered readably, with the relevant section on top.
  • Each question asked and its verdict, with the per-choice probabilities. "Refund requested: YES 0.55, UNCLEAR 0.40" tells a reviewer exactly where the doubt is.
  • Why it's here: low confidence, a tie between two labels, a policy that forces review, or a random audit sample.
  • Pointers to the evidence. Jev returns verdicts, not highlighted spans, but if your pipeline asks per-section questions, you can point the reviewer to the section that fired.

A reviewer who can see "every question passed except the one about disclosure" decides in seconds what used to take minutes. The call QA guide shows the same principle in coaching: a review session that opens with the agent's actual flagged moments, not a random call.

One decision per card, from a closed set

Give the human the same discipline you gave the model. Each card asks one question with a small set of answers, including an honest "can't tell" and a way to flag a bad question. Free-text boxes are optional extras, not the decision.

This pays twice. Reviewers move faster when the decision is framed. And every human answer arrives already structured, so it can flow straight into your gold set and eval suite without anyone transcribing notes.

Size the queue before you set the threshold

Most teams set a confidence threshold, then discover how big the resulting queue is. Reverse it. Decide how many items your reviewers can genuinely examine per day, then set the escalation threshold to fill that capacity with the least confident cases.

The Alan AI review console shows the shape, as reported: 3,518 internship applications checked in 4 minutes 54 seconds, with 20 surfaced for human review. The model narrows; people decide. Because that example is hiring, the rules get stricter, and we mean them: in hiring, grading, medical, or money decisions, a human makes the final call, rejected candidates shouldn't be silently filtered by a verdict nobody checked, and every decision needs a logged record of the verdicts shown and the human's ruling. Sample the "not surfaced" pile too, because that's where a biased filter hides. None of this is legal advice; regulated domains need your counsel's review.

The same posture shows up well outside hiring. The jev-auto-approve action asks Jev whether a pull request is the kind of change that can go in unattended, and everything else waits for a person, as reported. A deal-review demo in the directory has Jev map which teams need to look while, in the builder's words, people still make the call. Irreversible actions never fire on a lone verdict.

Fight automation bias on purpose

Showing the model's verdict speeds up triage queues. It also anchors people. Split the two jobs:

  • Triage queues (the model is uncertain, a human decides): show the verdicts and probabilities, since the evidence is the point.
  • Audit samples (checking whether confident verdicts are right): hide the model's answer until the reviewer commits. Otherwise you're measuring agreement with the machine, not accuracy.

Seed audit queues with a few known gold cases the reviewer doesn't know about. Agreement on those tells you whether the reviewer is still sharp, without anyone feeling watched by a spreadsheet.

Close the loop so reviewers see it matter

A queue feels pointless when corrections vanish. Make the effect visible: every disagreement becomes a candidate eval case, repeated corrections on the same pattern trigger a question rewrite, and a weekly note tells the team which boundary clause changed because of their calls. Log the verdicts shown, the human ruling, and the reviewer for every item; the audit trail guide covers the record. Rotate reviewers, cap session lengths, and treat reviewer fatigue as a system metric, not a personal failing.

Frequently asked questions

What is human in the loop design?

It's the practice of designing the human review step itself: what cases reach a person, what evidence they see, what decision they make, and how their answers feed back into the system. The goal is real judgment, not a rubber stamp.

What is an evidence-attached review queue?

A queue where every item arrives with the input, each question the model was asked, its verdicts and probabilities, and the reason the item was escalated. Reviewers decide from the card instead of hunting through other tools.

How many items should go to human review?

Set the number from reviewer capacity first, then tune the confidence threshold to fill it with the least certain cases. Routing logic itself is covered in escalation design patterns.

Should reviewers see the model's answer?

In triage queues, yes, because the evidence speeds decisions. In audit samples, no: hide it until they commit, or the audit only measures agreement with the model.

Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.