Guide
How to write annotation guidelines people can follow
Deepen AI · Published 2026-10-02
Most labeling problems are instruction problems. Two careful people read the same guideline, look at the same item and choose different labels, because the guideline never said what to do in that case. Good annotation guidelines remove those choices one at a time. This guide covers a structure that works, a process for testing it, and how to keep it current.
The structure
Keep the guideline as one document with these sections, in this order:
- Purpose. Two or three sentences on what the labels are for. Annotators make better calls on unclear items when they know whether the data trains a model, routes tickets or feeds a report.
- Label set. Every label with a one-sentence definition written in observable terms.
- Decision rules. The order in which to apply labels, and what wins when two seem to apply.
- Examples. For each label, at least two clear positives and one near miss that belongs elsewhere, each with its reason.
- Edge cases. A growing list of resolved hard cases, each with the decision and date.
- When you cannot decide. The flag path: who to ask, and what to do with the item meanwhile.
- Format rules. Span boundaries, capitalization, and how to handle empty or corrupted items.
- Version history. What changed, when and why.
Step 1: start from real data
Pull 50 to 100 real items before writing anything. Sort them by hand into the labels you think you need. You will find labels that never occur, items that fit none, and pairs that are hard to separate. Write the guideline around what you found, not around the label list you started with.
Step 2: write rules people can check
- Use observable criteria. "The customer mentions a charge, refund or invoice" beats "the ticket is about billing."
- One decision per rule. Split compound rules into steps.
- State precedence. For example: "If a ticket raises a billing issue and a shipping issue, label the one the customer asks us to act on. If they ask about both, label Billing."
- Remove vague words. Obviously, clearly, generally and usually hide the hard cases.
- Define boundaries. For spans, say whether to include punctuation, titles and possessives. For numbers, state units and rounding.
- Show, do not only tell. An example with its reason often does more than a paragraph.
Step 3: pilot with two or three people
Give the draft to two or three annotators who did not write it. Each labels the same 50 to 100 items independently. Do not let them discuss until they finish.
Step 4: measure agreement and read the disagreements
- Percent agreement: the share of items where annotators chose the same label. Easy to read, but inflated when one label dominates.
- Cohen's kappa: agreement between two annotators, adjusted for chance.
- Fleiss' kappa: the same idea for more than two annotators.
- Krippendorff's alpha: also handles missing labels and ordered scales.
The number tells you whether to look. The disagreements tell you what to fix. Put every disagreement in a table with the item, each label and each annotator's reason. Most fall into a handful of patterns, and each pattern becomes a rule, an example or a changed definition.
Set the agreement level you need from how the data will be used and from your pilot results. Subjective tasks, such as sentiment, rarely reach the agreement of clear-cut ones, such as language identification. Forcing the number up with arbitrary rules can make the labels less useful.
Step 5: revise, version and freeze
Make the changes, give the guideline a version number and date, and run a second, smaller pilot on new items. When results settle, freeze that version for live work. Record the guideline version on every labeled item, so you know later which rules each label followed.
Step 6: keep it alive
Once live work starts, edge cases arrive daily. Route them through one person who decides, adds the case to the guideline with a date, and tells the whole team at once. Batch small changes weekly. For a change that alters existing labels, decide whether to relabel older items, and record that decision too.
Worked example
Illustrative example. Invented data, not from a customer.
A support team labels tickets by topic: Billing, Shipping, Account or Product.
- Pilot result: two annotators disagreed on 14 of 80 tickets. Nine of those were between Billing and Account.
- Pattern: tickets about a failed card update. One annotator read "card" as payment (Billing), the other as profile settings (Account).
- Rule added: "Payment method changes, including adding, removing or updating a card, are Billing. Password, email and profile changes are Account."
- Example added: "I can't update my card, it keeps saying error." → Billing, because the customer is changing a payment method.
- Second pilot: Billing and Account disagreements fell to one ticket in 60.
A template you can copy
``` Guideline: <task name> Version: 1.0 Date: <YYYY-MM-DD> Owner: <name>
- Purpose
- Labels <Label>: <one-sentence observable definition>
- Decision rules (apply in order)
- Examples <Label> / positive / <item> / <reason> <Label> / near miss / <item> / <correct label and reason>
- Edge cases (dated)
- If you cannot decide: flag to <name or channel>; leave the item <unlabeled | in queue>
- Format rules
- Version history ```
Pitfalls
- Writing for yourself. The author knows what they meant. Test with people who do not.
- Too many labels. If two labels cannot be separated reliably, merge them or add a rule that separates them.
- Rules in chat, not in the document. Decisions made in messages get lost.
- No "cannot decide" path. People guess, and guesses look like labels.
- Changing rules without versioning. Old and new labels mix without anyone knowing.
- Text-only rules for visual data. Image, video and multi-sensor labeling also needs geometry rules: how tight a box must be, how to treat occluded or truncated objects, and how to keep an object's ID consistent across frames. That specialized work is the core of the Deepen AI enterprise team; see Deepen AI enterprise.
Checklist
- [ ] Purpose written in two or three sentences
- [ ] Labels defined in observable terms
- [ ] Precedence rules for overlapping labels
- [ ] Positive and near-miss examples, with reasons, for each label
- [ ] Flag path for items nobody can decide
- [ ] Pilot on 50 to 100 items with two or three annotators
- [ ] Agreement measured and every disagreement read
- [ ] Version number on the guideline and on every label
The same method works for review rubrics; the LLM output review playbook applies it to scoring model answers. If your labels are text, the text annotation page shows how a trained team works from guidelines like these.
FAQ
What should annotation guidelines include?
A purpose statement, label definitions, decision and precedence rules, examples with reasons, resolved edge cases, a path for items nobody can decide, format rules and a version history.
How do I test annotation guidelines?
Have two or three people who did not write them label the same 50 to 100 items independently, measure agreement, and turn each pattern of disagreement into a rule or an example.
What is inter-annotator agreement?
A measure of how often annotators choose the same label for the same item. Cohen's kappa, Fleiss' kappa and Krippendorff's alpha adjust for agreement expected by chance.
How often should annotation guidelines change?
Collect edge cases continuously, route them through one decision owner, and release dated versions, usually in weekly batches. Record the guideline version on every labeled item.