Amazon Mechanical Turk closed on 2026-09-30. Here is how to keep your work running →

Services / AI output review

AI output review

Your models and agents produce answers, summaries and extracted data. Our reviewers score each one against your rubric, correct what is wrong, and flag what needs your decision. You get a record of what was checked, what changed and why. The work happens in your tools, including your SageMaker Ground Truth or Amazon A2I portal.

Supported tasks

  • ·Score chatbot and agent answers against your rubric
  • ·Compare two model answers and pick the better one, with a reason
  • ·Correct AI-written text, such as product descriptions or support replies, before it goes live
  • ·Check AI-extracted fields against the source document
  • ·Review low-confidence items sent by Amazon A2I or Ground Truth
  • ·Tag failure types, such as wrong fact, missed step or policy break
  • ·Build small evaluation sets from your real data

Sample input and output

Illustrative example. Invented data, not from a customer.

Sample input

  • Your rubric: Correct = accurate and follows policy. Partly correct = accurate but misses a required step. Wrong = inaccurate or against policy.
  • Your policy: Full refund within 30 days. Store credit from day 31 to day 60.
  • Customer question: Can I return a jacket I bought 40 days ago?
  • AI answer: Yes. You can return any item within 60 days for a full refund.

Sample output

  • Score: Wrong
  • Failure type: Policy mismatch (refund window)
  • Corrected answer: You can return it for store credit. Full refunds apply within 30 days of purchase. From day 31 to day 60, we offer store credit.
  • Note to your team: Two other answers this week gave 60 days for full refunds. The prompt may quote an old policy.

How it works

  1. 1

    You share the rubric. Send your rubric, policy documents and 10 to 20 scored examples.

  2. 2

    We train and build a gold set. Reviewers train on your examples. We build a set of items with known-correct scores and confirm it with you.

  3. 3

    Reviewers score live outputs. They work in your tool. Unclear items go to the team lead, not a guess.

  4. 4

    You get a weekly report. Items reviewed, QA sample results, failure types and rubric questions.

Quality checks

The team lead checks: rubric questions from reviewers, every flagged item, and each reviewer's QA results. Open rubric questions go to your contact and into the weekly report.

QA reviewers sample: a share of finished items, re-scored against your rubric and the gold set. Disagreements go to the team lead and into the next training round.

What this means: every item is reviewed by one trained person. QA re-checks a sample, not every item. The weekly report shows the sample size and the acceptance rate.

How we check quality →

Pricing scope

  • $5 per person-hour (standard tasks): rubric scoring, pairwise comparison, correction of general English text, and extraction checks.
  • Quoted after our ops team reviews the task: review that needs a domain expert (for example medical, legal or financial content), languages other than English, or custom tool setup.
  • Not taken: graphic or violent content.

FAQ

Can you work inside our Amazon A2I or Ground Truth portal?

Yes. Add our people to a work team in your private workforce. Your task template stays the same. See the Ground Truth and A2I page for the steps.

Do you write the rubric?

You own the rubric. If you only have scored examples, we draft a rubric from them for you to approve before live work starts.

Is every item checked twice?

No. One trained reviewer handles each item. QA reviewers re-check a sample, and the weekly report shows how many.

Is there content you do not review?

Yes. We do not review graphic or violent content. Tell us in the pilot request if your outputs might contain it.

Try it on your own outputs for two weeks.

Send your rubric and a sample. We email a pilot plan with a quote.

Request a pilot plan →