Outsource
Scenario extraction from recorded drives and logs
Scenario extraction from recorded drives, logs and incident reports: events tagged to your taxonomy, variations and coverage gaps. Quoted per project.
Quoted per project, above the standard rate.
A scenario extraction service turns hours of recorded driving and system logs into a library of tagged test scenarios that your validation team can search, simulate and count. Deepen AI's specialist team finds the events, tags actors, maneuvers and conditions to your taxonomy, and reports where your coverage is thin.
What the work is, and who buys it
Buyers are validation, simulation and safety teams in automated driving, driver assistance and mobile robotics programs. They have recorded far more data than anyone has watched. Triggers such as hard braking, disengagements and near misses flag candidates, but someone still has to decide what happened, describe it the same way every time, and connect it to the operational design domain (ODD) the system is meant to handle.
This work sits close to Deepen AI's roots. Deepen AI is one of the authors of the ASAM OpenLABEL standard for multi-sensor data labeling and scenario tagging, and has worked with automotive and robotics teams since 2017.
What our team does
- Review trigger-flagged events in your replay tool, and confirm or reject each candidate
- Mark scenario start and end times, and add the window to your scenario library
- Tag actors, maneuvers, road layout, weather and lighting to your taxonomy, or to OpenLABEL scenario tags if you use them
- Turn incident and disengagement reports into structured scenario descriptions
- Write scenario variations from a seed scenario and your parameter ranges, in your template
- Check scenario coverage against your ODD and report the gaps
- Find and merge duplicate scenarios in an existing library
The team works in your log replay or visualization tool, your scenario database and your tracker, on accounts you create.
How it runs
- You share the taxonomy and examples. Send your scenario taxonomy or tag set, your ODD description, 20 to 30 scenarios your team extracted well, and access to the replay tool and log fields.
- We agree scope and train. Our ops team reviews the work and sends the people, hours and rate in your pilot plan. Analysts train on your examples, and we build a gold set of events with known-correct windows and tags for you to confirm.
- We deliver, with a named team lead and QA sampling. Analysts work through the drives. The team lead settles borderline maneuvers, keeps a decision list, and sends taxonomy gaps to your contact.
- You get a weekly report. Events reviewed, scenarios extracted by type, rejected candidates and why, coverage gaps, QA sample results and taxonomy questions.
Illustrative example: a hard-braking trigger to a tagged scenario
Illustrative only. The drive, values and IDs are invented.
- Input: drive 2026-08-14_route7, trigger "hard brake" at 00:41:12, safety driver note "car cut in from right, rain"
- Output, scenario: SC-0412, window 00:41:05 to 00:41:20
- Output, ego vehicle: lane keeping at 62 km/h on a divided highway, braking from 00:41:11
- Output, actor 1: passenger car, cut-in from the right adjacent lane, gap at entry 9 m (from your log's object list)
- Output, conditions: rain · daylight · wet road
- Output, tags (your taxonomy): cut-in · right side · short gap · ego brake response
- Output, coverage note: third rainy cut-in this month. Your ODD list has no night-and-rain cut-in example yet.
- Who did what: one analyst confirmed, extracted and tagged the event. A QA reviewer re-checked the maneuver tags and the window as part of the weekly sample and agreed.
How we keep scenarios consistent
A scenario library is only useful if the same event is tagged the same way by every analyst, every month. The decision list records rulings on borderline cases, such as when a lane change counts as a cut-in, and every analyst follows it. Gold events measure each analyst in training and in live work. QA reviewers re-check a sample of extracted scenarios, comparing windows and tags with the recording. They do not re-check every scenario. The weekly report also lists rejected candidates and the reasons, which helps your team tune its triggers. See how we check quality.
Pricing scope
Scenario generation and extraction is quoted per project, above the standard rate. The rate depends on your taxonomy, your tools and how much judgment each event needs. Describe the work in the pilot form, and our ops team sends the rate with your pilot plan. Work starts with a two-week prepaid pilot at the quoted rate, with at least 20 hours per person per week. The team lead and QA sampling are included. After the pilot, we invoice weekly in arrears. See the pricing page.
When to use the enterprise team instead
Annotating the sensor data itself, such as 3D boxes in lidar point clouds or camera labels, sensor calibration, automated edge-case selection across raw logs, and evaluation programs for driving models are enterprise work. Deepen AI has delivered annotation and data services for Daimler Trucks, BMW, Bosch, Mercedes, Hexagon, Aptiv and Ford Otosan. The enterprise team scopes and prices each project. See enterprise AI data services and human review and evaluation for Physical AI models. For map features and place data, see mapping data services.
Send your taxonomy and a few extracted scenarios. We email you a pilot plan with a quote.
FAQ
What is scenario extraction?
Finding meaningful events in recorded driving or robot logs, marking where each one starts and ends, and describing it in a consistent structure: who was involved, what they did, and under what conditions. The result is a scenario library your validation team can search, count and turn into simulation tests.
Can you extract driving scenarios from our recorded drives?
Yes, in your replay or visualization tool, using the log fields you give us access to. We tag each scenario to your taxonomy and add it to your scenario database or a file in your format.
Do you do AV edge case mining?
We do the human part: reviewing candidate events flagged by your triggers, such as hard braking or disengagements, confirming or rejecting each one, and tagging the real ones. Automated selection of edge cases across raw sensor logs is an enterprise service, scoped per project.
Do you use OpenLABEL or our own taxonomy?
Yours. If your taxonomy is based on ASAM OpenLABEL scenario tags, we follow it as written. Deepen AI is one of the authors of that standard.
Can you write scenario variations for simulation?
Yes. From a seed scenario and your parameter ranges, we write variations in your template, such as different speeds, gaps, weather or actor types. Your team runs them in your simulator.
Why is scenario work quoted per project?
It needs domain judgment, your tools and your taxonomy, so the rate depends on the setup. Our ops team reviews the work and sends the rate with your pilot plan.