Outsource
AI output review service for chatbot, agent and RAG outputs
An AI output review service: trained reviewers score, correct and sign off on chatbot, agent and RAG outputs against your rubric. $5 per person-hour.
$5 per person-hour for standard tasks.
An AI output review service puts trained people between your model and your users, or between your model and the decision to ship it. Deepen AI reviewers score outputs against your rubric, correct what is wrong, and record why, in your tools.
What review means, and who buys it
Buyers are teams shipping chatbots, agents, AI search and AI-written content, teams with a human review step in production, and teams that need evaluation data before a launch or a model change. "Review" covers three different jobs, and the pilot plan names which one you are buying:
- Gate. Check outputs before they reach a user or a system of record, and correct or block the bad ones.
- Measure. Score a fixed set of outputs so you can compare one prompt, model or retrieval setup with another.
- Teach. Produce preference judgments and rubric scores that your team uses to tune a model.
All three need the same foundation: a written rubric, a gold set of items with agreed answers, and a sampling plan. They differ in volume, in who sees the result, and in what happens next.
What our team does
- Score chatbot and agent answers against your rubric
- Check retrieval-augmented (RAG) answers claim by claim against the retrieved passages
- Review agent traces: the right tool, in the right order, stopping at the right point
- Compare two answers and record which is better, and why
- Write reference answers for evaluation sets
- Re-score a fixed evaluation set after each prompt or model change
- Tag failure types, such as unsupported claim, wrong tool, missed step, policy break or exposed personal data
- Review low-confidence items in your Amazon A2I or Ground Truth private workforce portal
For the service page, with a refund-policy example, see AI output review.
How it runs
- You share the rubric and examples. Send your rubric, the policy or source documents it refers to, and scored examples, including a few your own team disagreed on.
- We agree scope and train. Our ops team sets the people and hours in your pilot plan. Reviewers train on your scored examples, the team lead goes through each disagreement with them, and we build a gold set for you to confirm.
- We deliver, with a named team lead and QA sampling. Reviewers work in your tool. The team lead handles rubric questions and sends the ones only you can answer to your contact.
- You get a weekly report. Items reviewed, the score distribution, failure types, QA sample results and open rubric questions. For measurement work, scores are shown per version.
Illustrative example: a RAG answer checked claim by claim
Illustrative only. The product, plans and answer are invented.
- Input, question: "Does the Pro plan include single sign-on?"
- Input, retrieved passage: "SAML single sign-on is available on the Enterprise plan. The Pro plan includes two-factor authentication and audit logs."
- Input, AI answer: "Yes. The Pro plan includes single sign-on and two-factor authentication."
- Output, claim 1, Pro includes single sign-on: not supported. The passage puts it on Enterprise.
- Output, claim 2, Pro includes two-factor authentication: supported
- Output, score: Fail, one unsupported claim · failure type: unsupported claim
- Output, corrected answer: "No. Single sign-on is on the Enterprise plan. Pro includes two-factor authentication and audit logs."
- Who did what: one reviewer checked and corrected the answer. A QA reviewer re-scored it as part of the weekly sample and agreed.
How we keep review consistent
Review quality depends on the rubric first. If two careful people would score an item differently, the rubric needs a new line, and the team lead brings you that question rather than letting reviewers drift apart. Gold items with agreed scores are mixed into live work, so each reviewer is measured on real conditions. QA reviewers re-score a sample of finished items. They do not re-score every item. Disagreements go to the team lead and into the next training round. The weekly report shows how many items were sampled and the acceptance rate on that sample. If you also run an automated judge, the same gold set gives you a way to check it. See how we check quality.
Pricing scope
Rubric scoring, claim checks, pairwise comparison and correction of general English text are standard tasks at $5 per person-hour, with a named team lead and QA sampling included. Work starts with a two-week prepaid pilot: people × hours per week × 2 weeks × $5, with at least 20 hours per person per week. After the pilot, we invoice weekly in arrears. Review that needs a domain expert, such as medicine, law, finance or code, and review in other languages are quoted after our ops team reviews the task. We do not review graphic or violent content. See the pricing page.
When to use the enterprise team instead
If the outputs come from a Physical AI model, such as a driving or robot policy, reviewing its decisions against sensor recordings is a different job. The Deepen AI enterprise team runs human review and evaluation for Physical AI models, scoped and priced per project. See enterprise AI data services.
Send your rubric and a sample of outputs. We email you a pilot plan with a quote.
FAQ
What is a human-in-the-loop service?
A team of people who sit at a defined point in an AI workflow: they check, correct or approve what a model produced before it is used, or they score outputs so you can measure the model. We provide that team, trained on your rubric, working in your tools.
Do you offer RLHF outsourcing?
We produce the human feedback part: pairwise preference judgments, rubric scores and corrected answers, in English, on general topics. Your team trains the model. Feedback that needs domain experts, such as code, medicine, law or finance, is quoted after our ops team reviews the task.
How is human review different from an automated LLM judge?
An automated judge scores fast but needs checking itself. Human reviewers build the gold set you use to test the judge, review a sample of what it scores, and handle the cases it gets wrong. Many teams use both.
What if our rubric is still changing?
That is normal. The team lead logs every rubric question, sends it to your contact, and retrains the team when the rubric changes. The weekly report notes the date of each change, so you know which items were scored under which version.
Can you review outputs before they reach customers?
Yes, as a review step in your tool or your Amazon A2I portal. The hours your queue needs covered are agreed in your pilot plan. Items the rubric does not settle go to the team lead or your team, not to a guess.