Privacy & data prep
Meeting tools export transcripts for meetings, not for analysis: a timestamp on every line, one turn split across twenty captions, “Speaker 1” instead of a role, and every um intact. Paste or drop the file and get a transcript ready to code.
Free · No sign-up · Runs entirely in your browser
Paste or drop the transcript Zoom, Teams, Otter, Rev or your captioning tool gave you. Timestamps go, caption fragments become turns, speakers get consistent labels, and fillers are removed if you want clean verbatim. It all happens in this tab.
Drop a .vtt, .srt, .txt or .docx, or
Detected: WebVTT captions (Zoom, Teams, Meet) · 10 segments → 6 turns
Transcription style
Speakers
The speaker asking the most questions is taken as the interviewer. Rename labels as you like—names are replaced by these labels in the output.
Dana Whitfield
3 turns · 28% of words
Marcus Lee
3 turns · 72% of words
Turns
6
Words
117
Words removed
15
Duration
00:50
138 words/min
I: Okay, it's recording now. Thanks for joining. So to start, how did you first start caring for your mum?
P: It was about three years ago. She had a fall, and after that she couldn't really manage the stairs on her own. So I kind of just, you know, started going round every evening. [laughs]
I: What did a typical evening look like?
P: Cooking, her meds, making sure she was, settled. Then I'd drive home at like ten. Honestly I was exhausted most of the time, but I didn't want to, ask my brother.
I: Why didn't you want to ask him?
P: He lives in Leeds and he's got the kids. I just felt it was my job, I guess.
Reference · for the curious
Automatic transcription has made the recording-to-text step nearly free, and moved the work downstream. Every meeting tool exports its transcript in a different shape, built for re-watching a meeting rather than for analysing an interview. The cleaner above reshapes those exports into the form qualitative analysis expects; this page explains the choices it makes and the ones it leaves to you.
<v Name>) and gives each cue a long identifier; its Word export puts the name and a time on a line above each paragraph, and adds "started transcription" lines that are not speech.The cleaner detects which of these it has been given and says so, then rebuilds the conversation as one paragraph per turn. If a tool exported captions without speaker names—Zoom does this when speaker identification is off—it keeps the text but cannot split the speakers, and tells you.
True verbatim keeps every filled pause, stutter and false start. Clean verbatim removes those disfluencies while keeping every content word and the speaker's own phrasing—the convention most thematic and content analyses use, and the default here. The readable preset also removes hedges set off by commas (", you know,") and sound tags such as [laughs], which some teams prefer for sharing quotes and others consider a loss of meaning. Each rule is separate, and the "what changed" view colours every removal by the rule that made it, so the decision is visible rather than buried.
Which convention is right depends on the analysis. Conversation analysis, discourse analysis and studies of hesitation or speech in psychosis treat disfluency as data and need true verbatim—often with more detail than an automatic transcript contains. Clinical rating from interviews usually works from clean verbatim. Whichever you use, say so in the methods, along with who checked the automatic transcript against the audio.
The cleaner takes the speaker who asks the most questions per word as the interviewer and labels them I, with participants as P (or P1, P2 in group interviews). Labels are yours to change, and they replace the display names in the output—which matters, because a Zoom display name is often the participant's real name. Consistent labels at the start of each paragraph also let NVivo and MAXQDA auto-code the transcript by speaker, and are what Tagaroo's importer reads.
Interview transcripts are among the most identifying data a study holds. The de-identify option replaces names, contact details and dates with consistent pseudonyms, using the same detector as our transcript de-identifier, and applies it across the whole transcript so that one person keeps one pseudonym. It only applies high- and medium-confidence detections automatically; for a file you are about to share, run it through the de-identifier's review screen, which lists every detection including the uncertain ones.
Cleaning does not fix recognition errors. If you have corrected a sample of a transcript by hand, the word error rate calculator compares the two and shows which kinds of error the tool makes, and the transcription time calculator estimates how long checking the rest will take. Our guide to preparing audio and speech for annotation covers the upstream choices—diarisation, overlapping speech and transcription conventions—in more depth.
It does not transcribe audio; it works on the text your tool already produced. It cannot recover speaker names that were never recorded, and it cannot fix a turn the recogniser attributed to the wrong person—check speaker changes against the recording where they matter. The filler rules are English-language and deliberately conservative: they remove "um", "uh" and "erm", immediate repetitions and cut-off words, and leave "like" and "you know" alone unless they are set off by commas, because in the middle of a sentence they are often meaningful.
Download the transcript as a .vtt file (Zoom: Recordings → the meeting → Audio transcript; Teams: the meeting's Recap or Transcript tab → Download .vtt), then drop it into this cleaner. It strips every timestamp and cue number, joins the caption fragments back into whole turns, and keeps the speaker names so you can relabel them. If you do want a time reference for finding a passage in the recording later, the option to keep one timestamp at the start of each turn does exactly that.
True verbatim keeps everything that was said: filled pauses like um and uh, stutters, false starts and repetitions. Clean verbatim removes those disfluencies but keeps every content word and the speaker's own phrasing, which is the convention most qualitative studies use for thematic analysis. Some analyses need true verbatim—conversation analysis, discourse analysis and many linguistic studies treat hesitations as data—so state which convention you used in your methods.
Yes. Export from Otter as TXT (with or without timestamps) and paste or drop it here. Otter writes each turn under a header line with the speaker name and a time; the cleaner reads those headers, merges consecutive paragraphs by the same speaker, removes the times, and lets you rename “Speaker 1” and “Speaker 2” as the interviewer and participant.
The transcript is processed entirely by JavaScript in your browser; nothing is uploaded and there is no account. You can load the page, disconnect from the internet, and use it offline. For de-identification, the built-in option replaces high- and medium-confidence names, contact details and dates with pseudonyms; for a full review of every detection before sharing a file, use the dedicated transcript de-identifier.
Download the .docx or .txt version. The output puts each turn on its own paragraph starting with the speaker label, which all three packages import as a document, and which NVivo and MAXQDA can auto-code by speaker because the label is the first word of each paragraph. For Tagaroo, use “Annotate this in Tagaroo” to open it straight in the annotation workbench, or download the JSON.
Written by Enrique Gutiérrez, PhD (Computer Science)—founder of Tagaroo and Associate Professor of Computer Science, working on inter-rater reliability, measurement and annotation methodology (ORCID).
Last verified: 25 September 2026. Formulas, thresholds and cited figures on this page were checked against their original sources on that date. Every calculation runs in your browser; nothing you enter is transmitted or stored.
A clean transcript is where analysis starts. Tagaroo opens it turn by turn, lets you tag passages against your codebook or a clinical rating scale, and can code alongside you—with every AI suggestion carrying its rationale, and agreement with human coders measured as you go.
Interview transcript cleaner · tagaroo.ai/materials/transcript-cleaner · figures verified against primary sources 2026-09-25. Educational scoring aid, not a diagnosis, and not reviewed by a licensed clinician.