Study workflow
Enter the length of your recording and how it was captured, and get a realistic estimate of transcription hours, transcript word count, cost and turnaround — for typing it yourself, hiring a transcriptionist, or running AI transcription and correcting it.
Free · No sign-up · Runs entirely in your browser
That is 1 h of audio.
Your typing time
4.2–6.3 h
4.2–6.3 h of work per audio hour
Transcript
9,000 words
at 150 words per minute
Pages
36
double-spaced, ~250 words per page
Cost
Your own time
4.2–6.3 h of it
Turnaround
2–4 days
at 2–4 h of transcription a day
Treat this as a planning range. Published ratios disagree — 4:1 is the common benchmark for a professional, but sources put an untrained person anywhere from 4:1 to 12:1 — so these are bands, not quotes. Time one recording of your own material and you will have a better multiplier than any table, including this one.
Reference · for the curious
Almost every published answer to "how long does it take to transcribe an hour of audio?" is the same number: about four hours. That figure is a reasonable starting point and a poor planning tool, because the variance around it is enormous. The same hour of audio can take two hours or ten depending on how many people are talking, how the recording was made, and what you need the transcript to capture. This calculator makes those factors explicit instead of hiding them in an average.
The industry convention is a 4:1 ratio — four hours of work per hour of clear audio for a competent professional. Rev puts the average person at roughly four hours per audio hour and professionals at two to three. Ditto Transcripts describes four hours as the industry standard, with beginners at six to eight and non-specialists at eight to twelve. GMR Transcription quotes four to six. The Transcription Certification Institute, via GoTranscript, gives three to five.
These sources agree more than they disagree, and their consensus is worth stating plainly: a skilled transcriptionist on good audio needs three to five hours per audio hour, and everyone else needs more. If you are planning your own time and you have not done this before, budget six to ten hours for your first interview and expect to get faster.
Nothing else comes close. A clean recording made with a dedicated microphone in a quiet room transcribes at close to the theoretical minimum. A recording made on a laptop microphone across a table, in a room with air conditioning and corridor noise, can double or triple the time — not because the words are harder to type but because the transcriptionist rewinds constantly. Every unclear passage costs several times its own duration.
Ditto's published figures show the progression clearly: one speaker in clear dictation takes a professional two to three hours; a two-person interview with clear audio takes three to four; three people about four; four or more, four to five. Poor audio or heavy accents push it to five or six and beyond. Crosstalk is the specific problem — overlapping speech has to be untangled by ear, and speaker attribution becomes guesswork exactly when it matters most.
A clean-read transcript that removes false starts and filler is much faster to produce than true verbatim. If your analysis depends on hesitations, repetitions, self-corrections or pause lengths — as it does for anything touching disordered speech, therapy process coding, or conversation analysis — you are asking for verbatim, and verbatim is slower. Decide this before you commission the work; converting a clean-read transcript to verbatim afterwards means doing the job twice.
Conversational English runs about 130 to 160 words per minute; the National Center for Voice and Speech puts average US conversational speech near 150. An hour of interview audio therefore produces roughly 8,000 to 9,500 words — about 30 to 40 double-spaced pages. Two podcast hosts at a brisk 160 words per minute generate closer to 9,600 words an hour.
This matters for more than curiosity: word count drives your coding workload downstream. A study with twenty hour-long interviews is roughly 180,000 words to read, code and check — which is the real reason transcription time is worth estimating honestly. The transcript is the cheap part.
Machine transcription is effectively instant: one to three minutes of processing per hour of audio. The work does not disappear, it moves to review. Correcting a good automatic transcript of clear audio takes roughly half an hour to an hour and a half per audio hour. Poor recordings, strong accents, overlapping speakers and specialist vocabulary push that up substantially, and clinical vocabulary is exactly the case where automatic systems produce confident, plausible errors.
AI-plus-review is usually the fastest total route, and for most research it is the right default. But an uncorrected machine transcript is not analysis-ready, and treating it as though it were is how errors get laundered into findings. Budget the review honestly. Our guide to audio and speech annotation covers what to check first, and where automatic diarisation typically fails.
Professional transcription is generally priced per audio minute, not per hour of work, which is how the provider absorbs the variance you are trying to estimate. Published rates commonly sit in the region of $1.50 to $3.00 per audio minute for standard turnaround, with rush delivery and difficult audio charged higher — Ditto, for instance, publishes tiered rates that roughly double for its harder category and rise again for one-to-two-day turnaround.
Treat the figures this calculator produces as indicative market ranges for planning, not as a quote. Rates depend on your language, your turnaround, your verbatim requirements, whether you need timestamps and speaker labels, and whether the provider will sign a data processing agreement — which, for clinical interviews, is usually non-negotiable and narrows the field considerably.
Recordings of clinical or research interviews are among the most sensitive data a study holds, and sending them to a transcription service is a disclosure that needs a legal basis. Check what your ethics approval and participant consent actually permit, and what the provider commits to. Our guide to consent and licensing for annotation data covers the HIPAA and GDPR considerations.
If you need to share a transcript more widely than the recording, de-identify it first — the transcript de-identifier runs entirely in your browser, so the text is never uploaded anywhere.
Transcription is one line in a budget that also has to cover coding, double-coding for reliability, adjudication of disagreements, and analysis. Our annotation cost and timeline guide works through the rest, and if you intend to report inter-rater reliability, the sample-size planner will tell you how much double-coding you actually need — which is usually less than people fear, and more than they budget.
This calculator runs entirely in your browser. No inputs are sent anywhere and there is no account.
The widely cited industry benchmark is about four hours of work per hour of clear audio for a professional transcriptionist — a 4:1 ratio. Experienced specialists working on a single clear speaker can reach 2–3:1. People transcribing for the first time typically need 6–10 hours per audio hour. Multiple speakers, crosstalk, background noise and strong accents each push the ratio up; verbatim conventions that capture every false start push it up further.
Conversational English runs about 130–160 words per minute, so an hour of interview audio yields roughly 8,000–9,500 words of transcript. Interviews with two speakers and natural pausing sit near the lower end; fast single-speaker narration sits above it. At standard double spacing that is roughly 30–40 pages.
Machine transcription itself is effectively instant — typically 1–3 minutes of processing per hour of audio. The real cost moves to review. Correcting a good AI transcript of clear audio takes roughly 0.5–1.5 hours per audio hour, and considerably longer for poor recordings, heavy accents, overlapping speech or specialist vocabulary. AI-plus-review is usually the fastest route overall, but budget the review honestly: an uncorrected transcript is not analysis-ready, and for clinical or verbatim-sensitive work the corrections are the point.
Professional transcription is normally priced per audio minute rather than per hour of work, commonly in the region of $1.50 to $3.00 per audio minute for standard turnaround — roughly $90 to $180 for an hour of recording. Rush delivery and difficult audio (several speakers, poor recording quality, strong accents) push rates higher, sometimes close to double. Machine transcription costs a fraction of that per minute, but the human review it needs afterwards is the real expense. Treat all of these as indicative market ranges for planning, not a quote: your language, turnaround, verbatim requirements and whether the provider will sign a data processing agreement all move the price.
Audio quality first, then the number of speakers. A clean single-speaker recording made with a dedicated microphone transcribes quickly; a multi-party conversation captured on a laptop microphone in a room with background noise can double or triple the time, because the transcriptionist rewinds constantly. After that: verbatim requirements, speaker identification, timestamps, accents, and domain vocabulary the typist has to look up.
Written by Enrique Gutiérrez, PhD (Computer Science) — founder of Tagaroo and Associate Professor of Computer Science, working on inter-rater reliability, measurement and annotation methodology (ORCID).
Last verified: 29 July 2026. Formulas, thresholds and cited figures on this page were checked against their original sources on that date. Every calculation runs in your browser; nothing you enter is transmitted or stored.
Tagaroo transcribes your audio and video and hands you a speaker-separated transcript you can code immediately — with every annotation linked to the utterance that justifies it.