Amazon Mechanical Turk closed on 2026-09-30. Here is how to keep your work running →

Enterprise / Human review and evaluation

Deepen AI enterprise

Human review and evaluation for Physical AI models

A model can pass its metrics and still make the wrong call in a real scene. Deepen reviewers watch the recording, judge the model’s decision against your rubric, correct it, and say why.

1. What it solves

What it solves

Vision-language-action and end-to-end models turn camera, audio and sensor input into decisions and actions. Automated metrics can say whether an output matches a label. They cannot say whether stopping, yielding or proceeding was right for that scene. That takes a person who watches the recording, reads the telemetry and applies a written rubric, the same way each time.

Who uses it

TeamWhat they use it for
Robotics and embodied AI teams training VLA modelsReasoning-trace data: the situation, the goal, and the right decision and action for each recorded scenario
Autonomous driving teams evaluating end-to-end or VLA modelsScenario-by-scenario review of model decisions against a written rubric
Perception teamsHuman review of model predictions where they disagree with verified labels
Safety and validation teamsScenario, behavior and intent labels for evaluation sets

A good fit when

  • · You have recorded scenarios (multi-camera video, audio, telemetry) or model predictions on sensor data.
  • · You can describe what a correct decision looks like, or want help writing that rubric.
  • · You need a trained, managed team that stays on your program.

Not the right fit when

  • · The outputs are text-only chatbot or agent answers. See AI output review, a standard managed service.
  • · You need dataset-level coverage scoring and a certification package. See Validate as a Service ↗.

2. Inputs

What we work from

InputFormat or requirement
Scenario recordings for VLA reviewOne .zip per scenario. Inside one scenario folder: metadata.json, one MP4 per camera feed named by position (for example FRONT_CENTER.mp4), and an audio WAV whose file name matches metadata.json. metadata.json carries the scenario and run IDs, start and stop timestamps, duration, a transcription and timestamped telemetry records.
Model predictions to compareTwo label sets, such as model predictions and verified labels. For 3D point-cloud datasets in Validate, data and labels are uploaded as JSON.
Sensor data for scenario and behavior labelsImages, video, LiDAR and radar in the formats Deepen Annotate reads.
VLA data upload format ↗

What you provide

ItemWhy it matters
Rubric or guidelinesDefines a correct decision, action and post action. We can draft it with you from examples.
Attribute schemaThe fields reviewers fill. In VLA, Situation, Goal and the response attributes are configured per dataset.
ExamplesA few scenarios with answers you agree with, including hard ones.
AccessAccounts in your tool, or an agreed way to share data into Deepen's platform.
RightsConfirmation that you can share the data for this work.
A contactOne person who answers rubric questions during the pilot.

3. Work performed

What reviewers do, and what the software does

TaskWhat the reviewer doesWhat the software does
Situation and goalWatches the selected camera, audio and telemetry feeds. Writes the observable situation and the goal.Plays the selected feeds with run metadata.
Response reviewChecks the model’s Decision, Action and Post Action against the scene and the rubric. Corrects what is wrong, records why, and saves.VLA generates the responses from the inputs, situation and goal.
Prediction reviewReviews frames where model output and verified labels disagree, riskiest first. Confirms or corrects.Validate scores matching rate, category agreement and IoU, and ranks frames by risk.
Scenario and behavior labelsTags scenarios, events, intent and behavior.Annotate provides AI-assisted tools where they apply.
EscalationFlags unclear items instead of guessing. The team lead resolves them from the rubric, or asks your contact, and the rubric is updated.Issues on a record notify the reviewer and hold the discussion in comments.

Illustrative VLA review example

Scene
Warehouse shelf. Target: blue tote, top shelf. Two stacked totes immediately to its left. Task: "Pick the blue tote from the top shelf and place it on the conveyor without disturbing the stack beside it."
AI draft action plan
1. Approach shelf.
2. Grasp blue tote from the side.
3. Lift and rotate 90°.
4. Place on conveyor.Side grasp clips the adjacent stack by ~2 cm at this shelf spacing
Reviewed plan
1. Approach shelf.
2. Grasp blue tote from the top-center handle.
3. Lift and rotate 90°.
4. Place on conveyor.
Reviewer note
Side-grasp approach doesn’t clear the adjacent stack at this shelf spacing. Corrected the grasp point to top-center and flagged the failure mode for the next training batch.

Illustrative VLA review example. Not a real robot deployment or customer record.

4. Outputs

What you receive

OutputDetail
Reviewed responsesFor each scenario: the situation, the goal, and the reviewed Decision, Action and Post Action, saved in your VLA dataset.
Comparison reportsMatching rates, IoU distribution and downloadable mismatch breakdowns, tied to the frames they came from.
LabelsScenario and behavior labels, exported to ASAM OpenLABEL or another format your pipeline reads.
Issue logIssues raised during review, with comments and how each was resolved.
Batch reportsContent and cadence agreed in the statement of work.
DeliveryExport format and destination for review outputs agreed in the statement of work.

Traceability: Each VLA record is identified by the scenario ID and run ID from its metadata.json. Validate records each check, score and correction against the frames it came from.

Illustrative review record. Not customer data.

Model response

Decision: Proceed · Action: Keep speed and lane · Post action: Continue straight

Reviewed response

Decision: Yield · Action: Slow and hold lane position until the cyclist completes the turn · Post action: Resume speed once the cyclist clears the ego lane

Scenario
Urban intersection, light rain · 20 s · scenario_id: demo-0142
Feeds
FRONT_LEFT · FRONT_CENTER · FRONT_RIGHT · audio · telemetry
Situation (written by reviewer)
Ego vehicle approaching a signalized intersection at about 30 km/h on a green light. A cyclist ahead in the right lane is signaling a left turn across the ego lane.
Reviewer note
Changed Proceed to Yield. The cyclist’s left-turn signal is visible in FRONT_CENTER and FRONT_RIGHT from 00:06. Rubric rule 4.2: yield to a signaled turn across the ego path.
QA
In the QA sample. QA agreed with the correction.

Field names follow Deepen’s VLA workflow (Situation, Goal, Decision, Action, Post Action). The reviewer note is a text attribute configured for the dataset. One reviewer did this record; QA re-checks a sample of records, not every one.

5. Quality

How quality is checked

PracticeHow it worksDocumented, or agreed per project
Rubric developmentWe turn your guidelines and examples into a written rubric with worked cases, and add cases as new ones appear.Agreed in the statement of work
Reviewer trainingReviewers train on your rubric and examples before live work.Documented: Deepen trains and QAs its in-house team directly. The training plan is agreed in the statement of work.
Reference examplesScenarios with answers you approve, used to check reviewers.Agreed in the statement of work
Review tiersDeepen's annotation programs run an annotator pass, peer review and QA lead sign-off.Documented for annotation. Tiers for review work are agreed in the statement of work.
QA samplingA QA reviewer re-checks a sample of reviewed records.Sample size agreed in the statement of work
DisagreementsTwo review passes can be compared and scored for agreement in Validate. Disagreements go to the team lead.Tool capability documented. Whether double review is used is agreed in the statement of work.
AcceptanceAcceptance criteria are agreed before production. In Validate, you run checks, track issues and accept each dataset in Customer Review.Tool documented. Criteria agreed in the statement of work.
Correction and reworkErrors go back to the reviewer as issues with comments and are corrected.Tool documented.

What we do not claim: we do not publish an accuracy figure or a turnaround time for review and evaluation work. Targets, if any, are agreed in the statement of work.

6. Workflow and integration

How it fits your workflow

QuestionAnswer
Where does the work happen?In Deepen's tools (VLA, Validate and Annotate on the Deepen platform), or in your own tool on accounts you create.
Is working in our tool an integration?No. Our team can work in your tool with the access you give it. That is not a native integration.
How does data get in?Upload in the documented format, or through the Deepen Annotate API for datasets and labels.
Where does the Deepen platform run?As SaaS; as hybrid on-premise, where source data stays behind your VPN and generated annotation data is stored on Deepen servers; or fully on-premise, deployed with Deepen’s engineering team. Which applies to your program is agreed in the statement of work.

Annotate API quickstart ↗SaaS and on-premise options ↗

Who does what

DeepenYou
Reviewers, training and a team leadData, and the right to share it for this work
A dedicated program managerRubric, guidelines and examples
QA, issue handling and reportingOne contact who answers questions
Tool setup in Deepen's platformAccounts, if the work happens in your tool
Acceptance decisions

7. Security and procurement

Security and procurement

Data handling
Before work starts, we agree in writing what data moves, where it is processed and stored, who can access it, and when it is deleted.
Certifications
Deepen AI’s security program includes a SOC 2 Type II report, ISO 27001 certification and a TISAX assessment, and Deepen AI handles personal data in line with GDPR.
The certification scope covers the Deepen AI services delivery team.
Documentation
Reports, policies and the sub-processor list are in the Deepen AI Trust Vault ↗. Data processing terms are in the Data Protection Addendum ↗. NDA on request.

Procurement steps

  1. 1.Send an inquiry with the capability, data and timeline.
  2. 2.We scope the work with you by email or on a call.
  3. 3.NDA, security review and data processing terms.
  4. 4.Statement of work for the pilot.
  5. 5.Pilot on a sample of your data.
  6. 6.Statement of work for production.

Invoicing. Projects of $50,000 or more can be invoiced against a purchase order, on net terms, paid by wire transfer, with contract-based billing.

Security questionnaires and documentation requests: info@deepen.ai.

8. Engagement

Pilot, then production

PilotProduction
PurposeConfirm the rubric, the team and quality on your dataRun the full program
ScopeA sample of your scenarios, agreed in the statement of workVolume and schedule agreed in the statement of work
TeamTrained reviewers, a team lead and a dedicated program managerThe same team, sized to the program
ReportingPer batch, as agreed in the statement of workPer batch, as agreed in the statement of work

Pricing approach. Review and evaluation work is scoped and priced per project by the Deepen AI enterprise team. It is not billed at the standard rate.

What we need for an estimate

  • · The model type and what it outputs
  • · Feeds per scenario (cameras, audio, telemetry) and the file formats
  • · Number of scenarios or clips, and typical length
  • · Your rubric or guidelines, or a few examples with agreed answers
  • · Where the work should happen, and your timeline
  • · Security and data residency requirements
Discuss an enterprise project →

Put reviewers on your model’s decisions.

Tell us the model, the data and the timeline. The enterprise team replies by email with next steps.