Guide
Glossary of human data and annotation terms
Deepen AI · Published 2026-10-02
Plain definitions of the terms that come up when you plan, buy or run human data work. Each term has its own anchor, so you can link to it directly.
A
Acceptance rate. The share of items in a QA sample that pass review against the guidelines and gold set. It describes the sample, so always read it next to the sample size.
Activation condition. In Amazon A2I, the rule that decides when a prediction goes to a person, such as a confidence score below a threshold. AWS documents activation conditions for its Textract and Rekognition built-in task types; for custom task types, your code decides when to start a review.
AI output review. People checking what a model or agent produced, scoring it against a rubric, correcting it and flagging problems. See AI output review.
Amazon Augmented AI (A2I). An AWS service that sends model predictions to people for review. As of 2026-10-02, AWS documentation says A2I is no longer open to new customers; existing customers can continue to use it. See Ground Truth and A2I private workforce.
Amazon SageMaker Ground Truth. An AWS service for running data labeling jobs with a choice of workforce. As of 2026-10-02, AWS documentation says it is no longer open to new customers; existing customers can continue to use it.
Annotation guidelines. The written instructions that tell annotators which label to apply, in what order, with examples and edge cases. See how to write annotation guidelines.
ASAM OpenLABEL. An open standard from ASAM for the format of labels on multi-sensor data and for scenario tagging. Deepen AI is one of its authors.
B
Bounding box. A rectangle (2D) or cuboid (3D) drawn around an object to mark its position and extent.
C
Calibration (sensor). Measuring a sensor's internal properties (intrinsic calibration) and its position relative to other sensors (extrinsic calibration), so data from cameras, lidar and radar lines up. See Deepen AI enterprise.
Cohen's kappa. A measure of agreement between two annotators that adjusts for the agreement expected by chance.
Content moderation. Reviewing user content against a policy and deciding whether it stays, goes or needs escalation.
D
Data annotation (data labeling). Adding labels, tags or structured information to raw data so a model or a process can use it. See text annotation.
Data entry. Moving information from one place to another, such as from forms into a database, following field rules. See data entry.
Document data extraction. Pulling specific fields, such as totals, dates and names, from documents into structured data, by hand or by checking OCR and AI output. See document data extraction.
E
Edge case. An item the guidelines do not clearly cover. Good programs resolve each one, record the decision and add it to the guidelines.
Escalation. Sending an unclear item to a team lead or the customer for a decision instead of guessing.
F
Flow definition. The AWS API name for an Amazon A2I human review workflow. It sets the workforce, task template, activation conditions and output location.
G
Gold set. A set of items with known-correct answers, used to train people, test them and measure quality over time. See how we check quality.
Ground truth. The reference answer for an item, against which model output or human work is measured. Also the name of the AWS labeling service above.
H
HIT (Human Intelligence Task). A single unit of work on Amazon Mechanical Turk.
Human-in-the-loop (HITL). A workflow in which people review, correct or decide on some of a system's outputs, usually the uncertain or high-risk ones, before they are used.
Human loop. In Amazon A2I, one review of one item by the assigned workforce.
I
Inter-annotator agreement. How often annotators give the same label to the same item. It shows whether the guidelines are clear enough to apply consistently.
K
Krippendorff's alpha. An agreement measure that works with any number of annotators, missing labels and ordered scales.
M
Managed team. A vendor-supplied group of trained people with a lead and a quality process, who work on your tasks over time. See managed team vs freelancers.
Mechanical Turk (MTurk). Amazon's crowdsourcing marketplace. It stopped accepting new requesters on 2026-07-30 and closed permanently on 2026-09-30. See the MTurk shutdown FAQ.
O
OCR (optical character recognition). Software that turns images of text into machine-readable text. Its output often needs checking before use.
P
Pairwise comparison. A review task in which a person sees two outputs for the same input and picks the better one, usually with a reason.
PII (personally identifiable information). Data that can identify a person, such as a name, email or ID number. Plan which reviewers may see it and what gets masked before work starts.
Private workforce. In Ground Truth and A2I, a group of people you choose who sign in to a labeling portal for your AWS account. See the A2I private workforce set-up guide.
Product data enrichment. Filling and fixing product attributes, titles, categories and descriptions so listings are complete and consistent. See product data enrichment.
Q
QA sampling. A QA reviewer re-checking a sample of finished items, rather than every item, against the guidelines and gold set. See how we check quality.
R
Red teaming. Deliberately trying to make a model produce harmful, wrong or policy-breaking output, to find weaknesses before users do.
RLHF (reinforcement learning from human feedback). Training a model using human judgments, such as preferences between two answers, as the signal for what good output looks like.
Rubric. The scoring guide for a review task: the scale, what each level means, failure types and examples. See the LLM output review playbook.
T
Taxonomy. The set of labels or categories, and how they relate, that annotators apply.
V
Vendor workforce. In Ground Truth and A2I, a labeling company you subscribe to through AWS Marketplace.
W
Work team. In Ground Truth and A2I, a group within your private workforce that receives a specific job or review workflow.
Definitions checked 2026-10-02. AWS service facts come from the AWS SageMaker documentation; MTurk dates from mturk.com.
FAQ
What is the difference between data annotation and data labeling?
In practice, none. Both mean adding labels, tags or structured information to raw data so a model or a process can use it.
What is a gold set?
A set of items with known-correct answers, used to train people, test them and measure quality over time.
What does human-in-the-loop mean?
A workflow in which people review, correct or decide on some of a system's outputs, usually the uncertain or high-risk ones, before they are used.