Guide
Human review of LLM outputs: a practical playbook
Deepen AI · Published 2026-10-02
Automated evaluation catches a lot. It does not tell you whether an answer followed your refund policy, whether a summary dropped the one sentence that mattered, or whether an agent took a step a customer would object to. That is the job of human review. It works when you design it like any other measurement: a clear question, a fixed scale, known answers, and checks on the checkers.
This playbook is for teams starting or repairing a review program for chatbots, agents, summaries or extraction.
1. Decide what the review is for
Write one sentence before anything else. The three common purposes need different designs:
- Gate: decide whether an output can ship. Needs fast decisions and a clear escalation path.
- Measure: score a sample to track quality over time or compare two models. Needs a stable rubric and a fixed sample design.
- Correct: fix outputs to use as training data or to send to customers. Needs a correction standard, not just a score.
Mixing purposes without saying so is a common reason review data ends up unusable.
2. Write the rubric
A rubric is a short document reviewers can apply the same way every time. Include:
- The scale. Three levels is often enough: Correct, Partly correct, Wrong. Define each in terms a reviewer can check, such as "accurate and follows policy," not "good."
- Failure types. A fixed list, such as wrong fact, missed required step, policy break, unsafe content, formatting, or refused when it should have answered. Ask reviewers to tag one primary type.
- Tie-breakers. What wins when an answer is accurate but rude, or complete but too long.
- Sources of truth. Which policy document, which version, and what to do when it is silent.
- Examples. At least one per score level, including a hard one.
3. Build a gold set
A gold set is a group of items with known-correct scores. Take 50 to 200 real outputs, include the hard cases you already know about, and have two people who own the policy score them independently. Where they disagree, discuss, decide and write the reason down. That reason often becomes a rubric rule.
Once live work starts, mix gold items into the queue without marking them. Their value is that reviewers do not know which items are gold.
4. Calibrate reviewers before live work
Run a practice round on gold items. Compare each reviewer with the key, then hold a short session on the disagreements. They fall into three groups: the reviewer misread the rubric (train), the rubric is ambiguous (rewrite it), or the key is wrong (fix the key). Repeat until most disagreements are in the first group.
5. Measure agreement, not only accuracy
Accuracy against gold tells you about each reviewer. Agreement between reviewers tells you about the rubric.
- Percent agreement is how often two reviewers gave the same score. Easy to read, but it rewards defaulting to the most common score.
- Cohen's kappa corrects for the agreement two reviewers would reach by chance. Krippendorff's alpha handles more reviewers, missing scores and ordered scales.
Double-score a small slice of live items each week and track the trend. A drop usually means new kinds of outputs, a model change, or a rubric that no longer covers reality. Set your own threshold from pilot data rather than borrowing one, because what is acceptable depends on the task.
6. Sample for QA
You rarely need to re-review everything. A QA reviewer re-checks a sample of finished items against the rubric and gold set.
- Size. The standard sample-size formula for a proportion says a worst-case estimate within plus or minus 5 points, at 95 percent confidence, needs about 385 items. Within 3 points, about 1,070. Smaller samples still catch patterns; they give wider error bars.
- Stratify. Sample every reviewer every week, and over-sample new failure types and new model versions.
- Record. Items sampled, errors found, error types, and what changed as a result.
7. Escalate instead of guessing
Give reviewers a third option besides pass and fail: flag. Flagged items go to a lead who either decides from the rubric and adds the case to it, or asks the policy owner. Keep a dated decision log. A reviewer guessing silently on an unclear item does more damage than a slower answer.
8. Report so someone can act
A weekly report should show items reviewed, items sampled, errors found by type, open rubric questions, and patterns worth sending upstream. Reviewers often notice that a prompt quotes an outdated policy days before anyone else does.
Worked example
Illustrative example. Invented data, not from a customer.
A support assistant answers return questions. Policy: full refund within 30 days, store credit from day 31 to day 60.
- Question: Can I return a jacket I bought 40 days ago?
- Model answer: Yes. You can return any item within 60 days for a full refund.
- Score: Wrong. Failure type: policy break (refund window).
- Correction: You can return it for store credit. Full refunds apply within 30 days of purchase. From day 31 to day 60, we offer store credit.
- Flag to the prompt owner: three other answers this week gave 60 days for full refunds. The system prompt may quote an old policy.
In that week's QA sample of 300 items, the QA reviewer found 11 errors. Seven were reviewers scoring "Correct" when a required step was missing, so the rubric gained a new example and the team retrained on it. Four were refund-window errors the reviewers had missed, all on one topic, so the lead had every refund answer from that week re-checked.
Pitfalls
- Position bias in pairwise review. When comparing two answers, randomize which appears first and hide which model wrote which.
- Length bias. Longer answers look more thorough. Add a rubric line about unnecessary length.
- Rubric drift. Decisions made in chat never reach the document. Date every change.
- Reviewers using AI to review. A 2023 study of crowd workers on a summarization task (Veselovsky et al., arXiv:2306.07899) found evidence that some used LLMs to do the work. Ask any provider how they prevent and detect this.
- No owner for the rubric. Someone on your side must answer reviewer questions within an agreed time.
Checklist
- [ ] One-sentence purpose: gate, measure or correct
- [ ] Rubric with scale, failure types, tie-breakers and examples
- [ ] Gold set of 50 to 200 items with written reasons
- [ ] Calibration round before live work
- [ ] Weekly double-scored slice and an agreement trend
- [ ] QA sample that covers every reviewer
- [ ] Flag path and a dated decision log
- [ ] Weekly report sent to the people who can change prompts and models
If you want a trained team to run this process in your tools, see AI output review and how we check quality. For the rubric itself, how to write annotation guidelines applies directly.
FAQ
How many items should a gold set for LLM review contain?
Start with 50 to 200 real outputs, including the hard cases you already know about, scored independently by two people who own the policy. Add items as new failure types appear.
Does every LLM output need to be reviewed twice?
No. Most programs have one trained reviewer per item and a QA reviewer who re-checks a sample. Double-score a small slice each week to measure agreement.
What is the difference between reviewer accuracy and agreement?
Accuracy compares a reviewer with known-correct answers. Agreement compares reviewers with each other and shows whether the rubric is clear enough to apply the same way.
Should reviewers see which model produced an answer?
Not in comparisons. Hide the model name and randomize which answer appears first, so the score reflects the answer and not its source or position.