Enterprise / Human review and evaluation
Deepen AI enterprise
Human review and evaluation for Physical AI models
A model can pass its metrics and still make the wrong call in a real scene. Deepen reviewers watch the recording, judge the model’s decision against your rubric, correct it, and say why.
1. What it solves
What it solves
Vision-language-action and end-to-end models turn camera, audio and sensor input into decisions and actions. Automated metrics can say whether an output matches a label. They cannot say whether stopping, yielding or proceeding was right for that scene. That takes a person who watches the recording, reads the telemetry and applies a written rubric, the same way each time.
Who uses it
| Team | What they use it for |
|---|---|
| Robotics and embodied AI teams training VLA models | Reasoning-trace data: the situation, the goal, and the right decision and action for each recorded scenario |
| Autonomous driving teams evaluating end-to-end or VLA models | Scenario-by-scenario review of model decisions against a written rubric |
| Perception teams | Human review of model predictions where they disagree with verified labels |
| Safety and validation teams | Scenario, behavior and intent labels for evaluation sets |
A good fit when
- · You have recorded scenarios (multi-camera video, audio, telemetry) or model predictions on sensor data.
- · You can describe what a correct decision looks like, or want help writing that rubric.
- · You need a trained, managed team that stays on your program.
Not the right fit when
- · The outputs are text-only chatbot or agent answers. See AI output review, a standard managed service.
- · You need dataset-level coverage scoring and a certification package. See Validate as a Service ↗.
2. Inputs
What we work from
| Input | Format or requirement |
|---|---|
| Scenario recordings for VLA review | One .zip per scenario. Inside one scenario folder: metadata.json, one MP4 per camera feed named by position (for example FRONT_CENTER.mp4), and an audio WAV whose file name matches metadata.json. metadata.json carries the scenario and run IDs, start and stop timestamps, duration, a transcription and timestamped telemetry records. |
| Model predictions to compare | Two label sets, such as model predictions and verified labels. For 3D point-cloud datasets in Validate, data and labels are uploaded as JSON. |
| Sensor data for scenario and behavior labels | Images, video, LiDAR and radar in the formats Deepen Annotate reads. |
What you provide
| Item | Why it matters |
|---|---|
| Rubric or guidelines | Defines a correct decision, action and post action. We can draft it with you from examples. |
| Attribute schema | The fields reviewers fill. In VLA, Situation, Goal and the response attributes are configured per dataset. |
| Examples | A few scenarios with answers you agree with, including hard ones. |
| Access | Accounts in your tool, or an agreed way to share data into Deepen's platform. |
| Rights | Confirmation that you can share the data for this work. |
| A contact | One person who answers rubric questions during the pilot. |
3. Work performed
What reviewers do, and what the software does
| Task | What the reviewer does | What the software does |
|---|---|---|
| Situation and goal | Watches the selected camera, audio and telemetry feeds. Writes the observable situation and the goal. | Plays the selected feeds with run metadata. |
| Response review | Checks the model’s Decision, Action and Post Action against the scene and the rubric. Corrects what is wrong, records why, and saves. | VLA generates the responses from the inputs, situation and goal. |
| Prediction review | Reviews frames where model output and verified labels disagree, riskiest first. Confirms or corrects. | Validate scores matching rate, category agreement and IoU, and ranks frames by risk. |
| Scenario and behavior labels | Tags scenarios, events, intent and behavior. | Annotate provides AI-assisted tools where they apply. |
| Escalation | Flags unclear items instead of guessing. The team lead resolves them from the rubric, or asks your contact, and the rubric is updated. | Issues on a record notify the reviewer and hold the discussion in comments. |
Illustrative VLA review example
2. Grasp blue tote from the side.
3. Lift and rotate 90°.
4. Place on conveyor.Side grasp clips the adjacent stack by ~2 cm at this shelf spacing
2. Grasp blue tote from the top-center handle.
3. Lift and rotate 90°.
4. Place on conveyor.
Illustrative VLA review example. Not a real robot deployment or customer record.
4. Outputs
What you receive
| Output | Detail |
|---|---|
| Reviewed responses | For each scenario: the situation, the goal, and the reviewed Decision, Action and Post Action, saved in your VLA dataset. |
| Comparison reports | Matching rates, IoU distribution and downloadable mismatch breakdowns, tied to the frames they came from. |
| Labels | Scenario and behavior labels, exported to ASAM OpenLABEL or another format your pipeline reads. |
| Issue log | Issues raised during review, with comments and how each was resolved. |
| Batch reports | Content and cadence agreed in the statement of work. |
| Delivery | Export format and destination for review outputs agreed in the statement of work. |
Traceability: Each VLA record is identified by the scenario ID and run ID from its metadata.json. Validate records each check, score and correction against the frames it came from.
Illustrative review record. Not customer data.
Model response
Decision: Proceed · Action: Keep speed and lane · Post action: Continue straight
Reviewed response
Decision: Yield · Action: Slow and hold lane position until the cyclist completes the turn · Post action: Resume speed once the cyclist clears the ego lane
- Scenario
- Urban intersection, light rain · 20 s · scenario_id: demo-0142
- Feeds
- FRONT_LEFT · FRONT_CENTER · FRONT_RIGHT · audio · telemetry
- Situation (written by reviewer)
- Ego vehicle approaching a signalized intersection at about 30 km/h on a green light. A cyclist ahead in the right lane is signaling a left turn across the ego lane.
- Reviewer note
- Changed Proceed to Yield. The cyclist’s left-turn signal is visible in FRONT_CENTER and FRONT_RIGHT from 00:06. Rubric rule 4.2: yield to a signaled turn across the ego path.
- QA
- In the QA sample. QA agreed with the correction.
Field names follow Deepen’s VLA workflow (Situation, Goal, Decision, Action, Post Action). The reviewer note is a text attribute configured for the dataset. One reviewer did this record; QA re-checks a sample of records, not every one.
5. Quality
How quality is checked
| Practice | How it works | Documented, or agreed per project |
|---|---|---|
| Rubric development | We turn your guidelines and examples into a written rubric with worked cases, and add cases as new ones appear. | Agreed in the statement of work |
| Reviewer training | Reviewers train on your rubric and examples before live work. | Documented: Deepen trains and QAs its in-house team directly. The training plan is agreed in the statement of work. |
| Reference examples | Scenarios with answers you approve, used to check reviewers. | Agreed in the statement of work |
| Review tiers | Deepen's annotation programs run an annotator pass, peer review and QA lead sign-off. | Documented for annotation. Tiers for review work are agreed in the statement of work. |
| QA sampling | A QA reviewer re-checks a sample of reviewed records. | Sample size agreed in the statement of work |
| Disagreements | Two review passes can be compared and scored for agreement in Validate. Disagreements go to the team lead. | Tool capability documented. Whether double review is used is agreed in the statement of work. |
| Acceptance | Acceptance criteria are agreed before production. In Validate, you run checks, track issues and accept each dataset in Customer Review. | Tool documented. Criteria agreed in the statement of work. |
| Correction and rework | Errors go back to the reviewer as issues with comments and are corrected. | Tool documented. |
What we do not claim: we do not publish an accuracy figure or a turnaround time for review and evaluation work. Targets, if any, are agreed in the statement of work.
6. Workflow and integration
How it fits your workflow
| Question | Answer |
|---|---|
| Where does the work happen? | In Deepen's tools (VLA, Validate and Annotate on the Deepen platform), or in your own tool on accounts you create. |
| Is working in our tool an integration? | No. Our team can work in your tool with the access you give it. That is not a native integration. |
| How does data get in? | Upload in the documented format, or through the Deepen Annotate API for datasets and labels. |
| Where does the Deepen platform run? | As SaaS; as hybrid on-premise, where source data stays behind your VPN and generated annotation data is stored on Deepen servers; or fully on-premise, deployed with Deepen’s engineering team. Which applies to your program is agreed in the statement of work. |
Annotate API quickstart ↗SaaS and on-premise options ↗
Who does what
| Deepen | You |
|---|---|
| Reviewers, training and a team lead | Data, and the right to share it for this work |
| A dedicated program manager | Rubric, guidelines and examples |
| QA, issue handling and reporting | One contact who answers questions |
| Tool setup in Deepen's platform | Accounts, if the work happens in your tool |
| Acceptance decisions |
7. Security and procurement
Security and procurement
- Data handling
- Before work starts, we agree in writing what data moves, where it is processed and stored, who can access it, and when it is deleted.
- Certifications
- Deepen AI’s security program includes a SOC 2 Type II report, ISO 27001 certification and a TISAX assessment, and Deepen AI handles personal data in line with GDPR.
- The certification scope covers the Deepen AI services delivery team.
- Documentation
- Reports, policies and the sub-processor list are in the Deepen AI Trust Vault ↗. Data processing terms are in the Data Protection Addendum ↗. NDA on request.
Procurement steps
- 1.Send an inquiry with the capability, data and timeline.
- 2.We scope the work with you by email or on a call.
- 3.NDA, security review and data processing terms.
- 4.Statement of work for the pilot.
- 5.Pilot on a sample of your data.
- 6.Statement of work for production.
Invoicing. Projects of $50,000 or more can be invoiced against a purchase order, on net terms, paid by wire transfer, with contract-based billing.
Security questionnaires and documentation requests: info@deepen.ai.
8. Engagement
Pilot, then production
| Pilot | Production | |
|---|---|---|
| Purpose | Confirm the rubric, the team and quality on your data | Run the full program |
| Scope | A sample of your scenarios, agreed in the statement of work | Volume and schedule agreed in the statement of work |
| Team | Trained reviewers, a team lead and a dedicated program manager | The same team, sized to the program |
| Reporting | Per batch, as agreed in the statement of work | Per batch, as agreed in the statement of work |
Pricing approach. Review and evaluation work is scoped and priced per project by the Deepen AI enterprise team. It is not billed at the standard rate.
What we need for an estimate
- · The model type and what it outputs
- · Feeds per scenario (cameras, audio, telemetry) and the file formats
- · Number of scenarios or clips, and typical length
- · Your rubric or guidelines, or a few examples with agreed answers
- · Where the work should happen, and your timeline
- · Security and data residency requirements
Put reviewers on your model’s decisions.
Tell us the model, the data and the timeline. The enterprise team replies by email with next steps.