Amazon Mechanical Turk closed on 2026-09-30. Here is how to keep your work running →

Quality

How we check quality

We use the same approach as our annotation work. We train people on your rules, measure them against known-correct answers, check a sample of live work, and report what we find every week. This page explains each step, and what is and is not checked twice.

What a person does, and what QA checks

StepWhoWhich items
Doing the workOne trained personEvery item
Team lead reviewYour named team leadFlagged and unclear items, and each person's error pattern
QA checkA QA reviewerA sample of finished items, not every item
Gold checkScored against the known-correct answerGold items mixed into live work

We do not check every item twice by default.

1. Training before live work

Each person trains on your guidelines and finished examples before they touch live work. When your guidelines change, the team retrains, and the weekly report notes the date.

2. Reference answers (gold sets)

We build a gold set from your guidelines and examples: items with known-correct answers that you approve. It lets us measure accuracy instead of guessing at it. We add gold items when new edge cases appear. You can keep the gold set and use it with any team.

3. QA sampling

QA reviewers re-check a sample of finished items against your guidelines and the gold set. The weekly report shows how many items were done, how many were sampled, and the acceptance rate on the sample.

4. Escalation

Unclear items are flagged, not guessed. The person sends them to the team lead. If the guidelines answer the question, the team lead decides and adds the case to the guidelines. If not, the team lead asks your contact, and the question appears in the weekly report under open questions. Hard cases go up a review tier.

5. Corrections

You can send back any item you think is wrong. Your team lead replies with a fix or a reason.

6. What you see each week

The weekly report shows hours worked, items completed, items sampled, the acceptance rate on the sample, error categories, rework done, open questions and what changes next week.

See an example report →

What we do not claim

We do not promise zero errors. We show you the error rate we find, and what we changed because of it.

See it on your own work.

Two-week paid pilot. Weekly reports from week one.

Request a pilot plan →