Outsource
Outsource transcription to a trained team
Outsource English transcription: interviews and calls typed or machine transcripts corrected to your style guide, with speaker labels. $5 per person-hour.
$5 per person-hour for standard tasks.
Outsourcing transcription gives your recorded interviews, calls and meetings to a trained team that turns them into clean, consistent text. Deepen AI transcribes English audio from scratch or corrects your machine transcripts, to your style guide, in your tools.
What the work is, and who buys it
Buyers are user research and market research teams with interview recordings, product and sales teams reviewing recorded calls, podcast and media producers, and speech AI teams that need corrected transcripts to measure a speech model. Most of them already run automatic speech recognition. The open question is what to do about its errors: names, product terms, numbers, crosstalk, and who said what.
Human transcription vs AI. Machine transcription is good enough for search and rough notes. People are worth paying for when the transcript will be quoted, published, used as evidence in research, or used as reference data to test a speech model. The setup we see most is both: the machine produces a first draft, and our team corrects it against the audio and your glossary. For poor audio or heavy crosstalk, typing from scratch can work better. The pilot shows which approach suits your recordings.
What our team does
- Transcribe interviews, focus groups and recorded calls in English
- Correct machine transcripts against the audio, using your glossary of names and terms
- Add timestamps and speaker labels in your format
- Apply clean verbatim or full verbatim, as your style guide says
- Mark inaudible or overlapping speech with a timestamp instead of guessing
- Mask personal details, such as phone numbers, where your rules require it
- Build corrected reference transcripts for testing a speech model, tagged by the error types the machine made
The work happens in your tools: a transcription editor, your research repository, a shared drive, or a spreadsheet of segments. If you also need AI-written call summaries checked against the transcript, see AI output review outsourcing.
How it runs
- You share recordings and the style guide. Send a few sample recordings, your style guide or a transcript you like, and a glossary of names and terms.
- We agree scope and train. Our ops team sets the people and hours in your pilot plan. Transcribers train on your samples, and we build a gold set of segments with approved transcripts for you to confirm.
- We deliver, with a named team lead and QA sampling. The team lead settles unclear terms, adds them to the glossary, and sends open questions to your contact, such as how a guest spells their name.
- You get a weekly report. Hours worked, audio minutes completed, QA sample results, the most common error types and open questions.
Illustrative example: correcting a machine transcript
Illustrative only. The recording and names are invented.
- Input, machine draft [00:12:40]: "SPEAKER_1: so we we moved the sink job to night lee because the cash was stale"
- Input, machine draft [00:12:51]: "SPEAKER_2: right and that fixed the dash board numbers"
- Input, your rules: clean verbatim · name speakers from the intro · product terms from the glossary
- Output [00:12:40]: Priya (engineering lead): So we moved the sync job to nightly because the cache was stale.
- Output [00:12:51]: Tom (product manager): Right, and that fixed the dashboard numbers.
- Changes logged: four glossary fixes (sync, nightly, cache, dashboard) · one repeated word removed · two speakers named
- Who did what: one transcriber corrected the segment against the audio. A QA reviewer re-checked it as part of the weekly sample and agreed.
How we check transcripts
Transcript quality comes from three things: a clear style guide, a glossary that grows each week, and checks against the audio. Gold segments measure each transcriber against approved transcripts. QA reviewers then listen to a sample of finished segments and compare them with the text. They do not re-check every minute. The weekly report shows how many minutes were sampled, the acceptance rate on that sample, and the error types found: terms, speakers, omissions and punctuation. We do not quote a word accuracy figure in advance, because it depends on your audio. The report shows what QA found on your recordings. See how we check quality.
Pricing scope
English transcription and transcript correction are standard tasks at $5 per person-hour, with a named team lead and QA sampling included. We price by the person-hour, not per audio minute. Work starts with a two-week prepaid pilot: people × hours per week × 2 weeks × $5, with at least 20 hours per person per week. After the pilot, we invoice weekly in arrears. Other languages, and medical or legal transcription that needs specialist knowledge, are quoted after our ops team reviews the task. See the pricing page.
When to use the enterprise team instead
If you are building speech or multimodal datasets for vehicles or robots, such as in-cabin voice commands recorded alongside camera and sensor streams, that is collection and annotation work for the Deepen AI enterprise team. It is scoped and priced per project. See enterprise AI data services.
Send a few recordings and your style guide. We email you a pilot plan with a quote.
FAQ
Is transcription priced per minute or per hour?
Per person-hour: $5 for English transcription and transcript correction, with a team lead and QA sampling included. How long an hour of audio takes depends on sound quality, the number of speakers and whether we start from a machine draft. The weekly report shows hours worked and audio minutes completed, so after the pilot you can work out your own cost per minute.
Human transcription vs AI: which do we need?
Machine transcription is fine for search and rough notes. Pay for people when the transcript will be quoted, published, used as research evidence, or used as reference data to test a speech model. Many teams use both: the machine drafts, and our team corrects the draft against the audio.
Can you correct our existing machine transcripts?
Yes. Send the drafts with the audio and your glossary of names and terms. We correct them in your editor or file format and log the error types the machine made most.
Do you add speaker labels and timestamps?
Yes, in your format. Speakers are named from the introduction or a list you provide. Where a speaker cannot be identified, we label them consistently and flag it rather than guess.
Are you a transcription services company or a managed team?
A managed team. The same trained people work on your recordings each week, under a named team lead, to your style guide. That suits steady volume and house rules. For a single file now and then, a per-file service may fit better.
Which languages do you transcribe?
English at the standard rate. Other languages, and medical or legal transcription that needs specialist knowledge, are quoted after our ops team reviews the task.