Study workflow

Transcription time & cost calculator

Enter the length of your recording and how it was captured, and get a realistic estimate of transcription hours, transcript word count, cost and turnaround — for typing it yourself, hiring a transcriptionist, or running AI transcription and correcting it.

Free · No sign-up · Runs entirely in your browser

minutes

That is 1 h of audio.

Your typing time

4.2–6.3 h

4.2–6.3 h of work per audio hour

Transcript

9,000 words

at 150 words per minute

Pages

36

double-spaced, ~250 words per page

Cost

Your own time

4.2–6.3 h of it

Turnaround

2–4 days

at 2–4 h of transcription a day

How this was calculated

  • 2.50 h/hBase typing ratio2–3 h of typing per audio hour for a professional on clear, single-speaker audio (Rev; Ditto Transcripts)
  • × 1.40Experience: some practiceBetween Rev's optimistic 4:1 for an average person and Ditto's 8-12:1 for an untrained one
  • × 1.502 speakersOverlapping speech and speaker labelling force repeated rewinding (Ditto Transcripts' speaker table)
  • × 1.00Clear audioDedicated microphone, quiet room, nothing to decipher
  • 4.2–6.3work-hours per audio hourThe band, not the midpoint, is the estimate.

Treat this as a planning range. Published ratios disagree — 4:1 is the common benchmark for a professional, but sources put an untrained person anywhere from 4:1 to 12:1 — so these are bands, not quotes. Time one recording of your own material and you will have a better multiplier than any table, including this one.

Reference · for the curious

How long transcription actually takes, and what drives the number

Almost every published answer to "how long does it take to transcribe an hour of audio?" is the same number: about four hours. That figure is a reasonable starting point and a poor planning tool, because the variance around it is enormous. The same hour of audio can take two hours or ten depending on how many people are talking, how the recording was made, and what you need the transcript to capture. This calculator makes those factors explicit instead of hiding them in an average.

The 4:1 rule and where it comes from

The industry convention is a 4:1 ratio — four hours of work per hour of clear audio for a competent professional. Rev puts the average person at roughly four hours per audio hour and professionals at two to three. Ditto Transcripts describes four hours as the industry standard, with beginners at six to eight and non-specialists at eight to twelve. GMR Transcription quotes four to six. The Transcription Certification Institute, via GoTranscript, gives three to five.

These sources agree more than they disagree, and their consensus is worth stating plainly: a skilled transcriptionist on good audio needs three to five hours per audio hour, and everyone else needs more. If you are planning your own time and you have not done this before, budget six to ten hours for your first interview and expect to get faster.

What actually drives the ratio

Audio quality, first and by a distance

Nothing else comes close. A clean recording made with a dedicated microphone in a quiet room transcribes at close to the theoretical minimum. A recording made on a laptop microphone across a table, in a room with air conditioning and corridor noise, can double or triple the time — not because the words are harder to type but because the transcriptionist rewinds constantly. Every unclear passage costs several times its own duration.

Number of speakers

Ditto's published figures show the progression clearly: one speaker in clear dictation takes a professional two to three hours; a two-person interview with clear audio takes three to four; three people about four; four or more, four to five. Poor audio or heavy accents push it to five or six and beyond. Crosstalk is the specific problem — overlapping speech has to be untangled by ear, and speaker attribution becomes guesswork exactly when it matters most.

What the transcript has to capture

A clean-read transcript that removes false starts and filler is much faster to produce than true verbatim. If your analysis depends on hesitations, repetitions, self-corrections or pause lengths — as it does for anything touching disordered speech, therapy process coding, or conversation analysis — you are asking for verbatim, and verbatim is slower. Decide this before you commission the work; converting a clean-read transcript to verbatim afterwards means doing the job twice.

Word count: what an hour of speech becomes

Conversational English runs about 130 to 160 words per minute; the National Center for Voice and Speech puts average US conversational speech near 150. An hour of interview audio therefore produces roughly 8,000 to 9,500 words — about 30 to 40 double-spaced pages. Two podcast hosts at a brisk 160 words per minute generate closer to 9,600 words an hour.

This matters for more than curiosity: word count drives your coding workload downstream. A study with twenty hour-long interviews is roughly 180,000 words to read, code and check — which is the real reason transcription time is worth estimating honestly. The transcript is the cheap part.

AI transcription changes where the time goes, not whether there is any

Machine transcription is effectively instant: one to three minutes of processing per hour of audio. The work does not disappear, it moves to review. Correcting a good automatic transcript of clear audio takes roughly half an hour to an hour and a half per audio hour. Poor recordings, strong accents, overlapping speakers and specialist vocabulary push that up substantially, and clinical vocabulary is exactly the case where automatic systems produce confident, plausible errors.

AI-plus-review is usually the fastest total route, and for most research it is the right default. But an uncorrected machine transcript is not analysis-ready, and treating it as though it were is how errors get laundered into findings. Budget the review honestly. Our guide to audio and speech annotation covers what to check first, and where automatic diarisation typically fails.

Cost, and why quoted rates vary so much

Professional transcription is generally priced per audio minute, not per hour of work, which is how the provider absorbs the variance you are trying to estimate. Published rates commonly sit in the region of $1.50 to $3.00 per audio minute for standard turnaround, with rush delivery and difficult audio charged higher — Ditto, for instance, publishes tiered rates that roughly double for its harder category and rise again for one-to-two-day turnaround.

Treat the figures this calculator produces as indicative market ranges for planning, not as a quote. Rates depend on your language, your turnaround, your verbatim requirements, whether you need timestamps and speaker labels, and whether the provider will sign a data processing agreement — which, for clinical interviews, is usually non-negotiable and narrows the field considerably.

Before you send audio anywhere

Recordings of clinical or research interviews are among the most sensitive data a study holds, and sending them to a transcription service is a disclosure that needs a legal basis. Check what your ethics approval and participant consent actually permit, and what the provider commits to. Our guide to consent and licensing for annotation data covers the HIPAA and GDPR considerations.

If you need to share a transcript more widely than the recording, de-identify it first — the transcript de-identifier runs entirely in your browser, so the text is never uploaded anywhere.

Planning the whole study, not just the typing

Transcription is one line in a budget that also has to cover coding, double-coding for reliability, adjudication of disagreements, and analysis. Our annotation cost and timeline guide works through the rest, and if you intend to report inter-rater reliability, the sample-size planner will tell you how much double-coding you actually need — which is usually less than people fear, and more than they budget.

Privacy

This calculator runs entirely in your browser. No inputs are sent anywhere and there is no account.

Frequently asked questions

How long does it take to transcribe one hour of audio?

The widely cited industry benchmark is about four hours of work per hour of clear audio for a professional transcriptionist — a 4:1 ratio. Experienced specialists working on a single clear speaker can reach 2–3:1. People transcribing for the first time typically need 6–10 hours per audio hour. Multiple speakers, crosstalk, background noise and strong accents each push the ratio up; verbatim conventions that capture every false start push it up further.

How many words are in an hour of recorded speech?

Conversational English runs about 130–160 words per minute, so an hour of interview audio yields roughly 8,000–9,500 words of transcript. Interviews with two speakers and natural pausing sit near the lower end; fast single-speaker narration sits above it. At standard double spacing that is roughly 30–40 pages.

Is AI transcription faster than a human?

Machine transcription itself is effectively instant — typically 1–3 minutes of processing per hour of audio. The real cost moves to review. Correcting a good AI transcript of clear audio takes roughly 0.5–1.5 hours per audio hour, and considerably longer for poor recordings, heavy accents, overlapping speech or specialist vocabulary. AI-plus-review is usually the fastest route overall, but budget the review honestly: an uncorrected transcript is not analysis-ready, and for clinical or verbatim-sensitive work the corrections are the point.

How much does it cost to transcribe an hour of audio?

Professional transcription is normally priced per audio minute rather than per hour of work, commonly in the region of $1.50 to $3.00 per audio minute for standard turnaround — roughly $90 to $180 for an hour of recording. Rush delivery and difficult audio (several speakers, poor recording quality, strong accents) push rates higher, sometimes close to double. Machine transcription costs a fraction of that per minute, but the human review it needs afterwards is the real expense. Treat all of these as indicative market ranges for planning, not a quote: your language, turnaround, verbatim requirements and whether the provider will sign a data processing agreement all move the price.

What affects transcription time the most?

Audio quality first, then the number of speakers. A clean single-speaker recording made with a dedicated microphone transcribes quickly; a multi-party conversation captured on a laptop microphone in a room with background noise can double or triple the time, because the transcriptionist rewinds constantly. After that: verbatim requirements, speaker identification, timestamps, accents, and domain vocabulary the typist has to look up.

Written by Enrique Gutiérrez, PhD (Computer Science) — founder of Tagaroo and Associate Professor of Computer Science, working on inter-rater reliability, measurement and annotation methodology (ORCID).

Last verified: 29 July 2026. Formulas, thresholds and cited figures on this page were checked against their original sources on that date. Every calculation runs in your browser; nothing you enter is transmitted or stored.

Skip the typing, keep the transcript

Tagaroo transcribes your audio and video and hands you a speaker-separated transcript you can code immediately — with every annotation linked to the utterance that justifies it.

Try Tagaroo free