Amazon Mechanical Turk closed on 2026-09-30. Here is how to keep your work running →

Outsource

Data annotation outsourcing for text and AI evaluation data

Outsource data annotation for text, search relevance and LLM evaluation data to a trained team in your labeling tool. $5 per person-hour, QA included.

$5 per person-hour for standard tasks.

Data annotation outsourcing gives the labeling behind your models to a trained, managed team. Deepen AI labels text, search results and model outputs in your labeling tool, to your guidelines. If your data is images, video, lidar or other sensor data, that is our enterprise team's core work, covered in the last section.

What the work is, and who buys it

Buyers are machine learning and product teams that need training or evaluation data labeled the same way every time, data teams relabeling after a taxonomy change, and operations teams whose classifications feed dashboards. The usual choice is between labeling in-house, which spends engineers' time, an open crowd, which starts quickly but is hard to keep consistent, or a managed team. A managed team suits work where the guidelines will keep changing and consistency matters more than raw volume on day one.

Annotation is Deepen AI's core business. We have built annotation tools, guidelines and review tiers since 2017. The text and AI-data work on this page uses the same method: written guidelines, gold examples, sampled QA and a weekly report.

What our team does

  • Classify text such as tickets, reviews, messages or documents to your taxonomy
  • Mark entities and spans, such as product names, companies, dates and amounts
  • Label intent and slots in search queries and chat turns
  • Judge search relevance by scoring how well each result answers a query, on your scale
  • Compare two model answers and record which is better, and why
  • Build and label small evaluation sets from your real data
  • Audit an existing labeled dataset and relabel it after a taxonomy change

The team works in your labeling tool, such as SageMaker Ground Truth, an open-source tool, an internal tool or a spreadsheet. For more on text labeling, with a support-ticket example, see text annotation. For scoring and correcting what your models produce, see AI output review outsourcing.

How it runs

  1. You share guidelines and examples. Send your label set, guidelines and labeled examples. If you have examples but no written guidelines, we draft guidelines from them for your approval.
  2. We agree scope and train. Our ops team sets the people and hours in your pilot plan. Annotators train on your examples, and we build a gold set of items with known-correct labels for you to confirm.
  3. We deliver, with a named team lead and QA sampling. Annotators label live items. The team lead keeps an edge-case log, and guideline updates go to the whole team at once, with the date noted.
  4. You get a weekly report. Items labeled, QA sample results, the label pairs most often confused, and open guideline questions.

Illustrative example: search relevance judgments

Illustrative only. The query, products and scale are invented.

  • Input, query: "waterproof hiking boots women wide"
  • Input, your scale: 3 exact match · 2 one minor attribute off · 1 same product type, wrong core attribute · 0 different product type
  • Output, "Women's TrailGuard Waterproof Hiking Boot, Wide": 3
  • Output, "Women's TrailGuard Waterproof Hiking Boot, Regular Width": 2, width is the minor attribute
  • Output, "Men's TrailGuard Waterproof Hiking Boot": 1, wrong gender
  • Output, "Women's Merino Hiking Socks": 0 under the current scale
  • Flag for your contact: should accessories for the same activity score 1? Sent as a guideline question.
  • Who did what: one annotator judged all four results. A QA reviewer re-judged the third result as part of the weekly sample and agreed.

How we measure label quality

Label quality depends on guidelines that answer the hard cases, so the team lead turns every new edge case into a guideline question for you. Gold items with known-correct labels measure each annotator in training and in live work. QA reviewers relabel a sample of finished items and compare the result with the annotator's label and the gold set. They do not relabel every item. The weekly report shows the acceptance rate on the sample and which label pairs get confused most, which usually points to the guideline line that needs rewriting. We do not quote an accuracy target before we have seen your data. See how we check quality.

Pricing scope

English text labeling, relevance judgments and evaluation data are standard tasks at $5 per person-hour, with a named team lead and QA sampling included. Work starts with a two-week prepaid pilot: people × hours per week × 2 weeks × $5, with at least 20 hours per person per week. After the pilot, we invoice weekly in arrears. Text that needs domain expertise, such as clinical notes, legal contracts or code, other languages, and annotation in a tool we must set up for you are quoted after our ops team reviews the task. See the pricing page.

When to use the enterprise team instead

Image, video, lidar and multi-sensor annotation is Deepen AI's core work, run by the enterprise team with our own annotation tools, trained annotators and review tiers. Deepen AI has delivered annotation and data services for Daimler Trucks, BMW, Bosch, Mercedes, Hexagon, Aptiv and Ford Otosan. The enterprise team scopes and prices each project. See enterprise AI data services.

Send your label set and a sample of your data. We email you a pilot plan with a quote.

Request a pilot plan

FAQ

How is data labeling priced?

English text labeling, relevance judgments and evaluation data are $5 per person-hour, with a named team lead and QA sampling included. We do not price per label. The weekly report shows items completed per hour, so after the two-week pilot you can work out your real cost per label.

Do you label images, video or lidar?

Yes, through the Deepen AI enterprise team. Image, video, lidar and multi-sensor annotation is Deepen AI's core work, done with our own annotation tools since 2017. The enterprise team scopes and prices each project. See the enterprise page.

What should we ask a data labeling company before we start?

Ask how annotators are trained on your guidelines, how quality is measured, how much finished work is re-checked, how guideline changes reach the whole team, and where your data goes. Ask to see the weekly report format. Then run a short paid pilot on your own data.

Can you label in our own tool?

Yes, if it runs in a browser and you can create accounts for us. That includes SageMaker Ground Truth, open-source labeling tools, internal tools and spreadsheets.

Can you help with AI data labeling for LLM evaluation?

Yes. We compare model answers, score them against a rubric, write reference answers and build evaluation sets from your real data, in English. Work that needs domain experts, such as code, medicine or law, is quoted after our ops team reviews the task.