Privacy & data prep

Interview transcript cleaner

Meeting tools export transcripts for meetings, not for analysis: a timestamp on every line, one turn split across twenty captions, “Speaker 1” instead of a role, and every um intact. Paste or drop the file and get a transcript ready to code.

Free · No sign-up · Runs entirely in your browser

From meeting-tool export to a transcript you can code

Paste or drop the transcript Zoom, Teams, Otter, Rev or your captioning tool gave you. Timestamps go, caption fragments become turns, speakers get consistent labels, and fillers are removed if you want clean verbatim. It all happens in this tab.

Drop a .vtt, .srt, .txt or .docx, or

Detected: WebVTT captions (Zoom, Teams, Meet) · 10 segments → 6 turns

Transcription style

Speakers

The speaker asking the most questions is taken as the interviewer. Rename labels as you like—names are replaced by these labels in the output.

Dana Whitfield

3 turns · 28% of words

Marcus Lee

3 turns · 72% of words

Turns

6

Words

117

Words removed

15

Duration

00:50

138 words/min

I: Okay, it's recording now. Thanks for joining. So to start, how did you first start caring for your mum?

P: It was about three years ago. She had a fall, and after that she couldn't really manage the stairs on her own. So I kind of just, you know, started going round every evening. [laughs]

I: What did a typical evening look like?

P: Cooking, her meds, making sure she was, settled. Then I'd drive home at like ten. Honestly I was exhausted most of the time, but I didn't want to, ask my brother.

I: Why didn't you want to ask him?

P: He lives in Leeds and he's got the kids. I just felt it was my job, I guess.

Reference · for the curious

Preparing interview transcripts for analysis: formats, conventions and what to keep

Automatic transcription has made the recording-to-text step nearly free, and moved the work downstream. Every meeting tool exports its transcript in a different shape, built for re-watching a meeting rather than for analysing an interview. The cleaner above reshapes those exports into the form qualitative analysis expects; this page explains the choices it makes and the ones it leaves to you.

What each tool gives you

  • Zoom cloud recordings produce an audio transcript as a WebVTT (.vtt) file: a caption cue every few seconds, each with a start and end time and the speaker's display name before a colon. One answer can span a dozen cues.
  • Microsoft Teams offers the transcript as .vtt or .docx. Its VTT marks speakers with voice tags (<v Name>) and gives each cue a long identifier; its Word export puts the name and a time on a line above each paragraph, and adds "started transcription" lines that are not speech.
  • Otter exports TXT with a header line per paragraph—"Speaker 1  0:07"—and splits long turns into several paragraphs under repeated headers.
  • Rev and similar services write "Name (00:01):" before each turn.

The cleaner detects which of these it has been given and says so, then rebuilds the conversation as one paragraph per turn. If a tool exported captions without speaker names—Zoom does this when speaker identification is off—it keeps the text but cannot split the speakers, and tells you.

Verbatim, clean verbatim, and what to report

True verbatim keeps every filled pause, stutter and false start. Clean verbatim removes those disfluencies while keeping every content word and the speaker's own phrasing—the convention most thematic and content analyses use, and the default here. The readable preset also removes hedges set off by commas (", you know,") and sound tags such as [laughs], which some teams prefer for sharing quotes and others consider a loss of meaning. Each rule is separate, and the "what changed" view colours every removal by the rule that made it, so the decision is visible rather than buried.

Which convention is right depends on the analysis. Conversation analysis, discourse analysis and studies of hesitation or speech in psychosis treat disfluency as data and need true verbatim—often with more detail than an automatic transcript contains. Clinical rating from interviews usually works from clean verbatim. Whichever you use, say so in the methods, along with who checked the automatic transcript against the audio.

Speaker labels

The cleaner takes the speaker who asks the most questions per word as the interviewer and labels them I, with participants as P (or P1, P2 in group interviews). Labels are yours to change, and they replace the display names in the output—which matters, because a Zoom display name is often the participant's real name. Consistent labels at the start of each paragraph also let NVivo and MAXQDA auto-code the transcript by speaker, and are what Tagaroo's importer reads.

Before the transcript leaves your machine

Interview transcripts are among the most identifying data a study holds. The de-identify option replaces names, contact details and dates with consistent pseudonyms, using the same detector as our transcript de-identifier, and applies it across the whole transcript so that one person keeps one pseudonym. It only applies high- and medium-confidence detections automatically; for a file you are about to share, run it through the de-identifier's review screen, which lists every detection including the uncertain ones.

How accurate was the automatic transcript?

Cleaning does not fix recognition errors. If you have corrected a sample of a transcript by hand, the word error rate calculator compares the two and shows which kinds of error the tool makes, and the transcription time calculator estimates how long checking the rest will take. Our guide to preparing audio and speech for annotation covers the upstream choices—diarisation, overlapping speech and transcription conventions—in more depth.

What this tool does not do

It does not transcribe audio; it works on the text your tool already produced. It cannot recover speaker names that were never recorded, and it cannot fix a turn the recogniser attributed to the wrong person—check speaker changes against the recording where they matter. The filler rules are English-language and deliberately conservative: they remove "um", "uh" and "erm", immediate repetitions and cut-off words, and leave "like" and "you know" alone unless they are set off by commas, because in the middle of a sentence they are often meaningful.

Frequently asked questions

How do I remove timestamps from a Zoom or Teams transcript?

Download the transcript as a .vtt file (Zoom: Recordings → the meeting → Audio transcript; Teams: the meeting's Recap or Transcript tab → Download .vtt), then drop it into this cleaner. It strips every timestamp and cue number, joins the caption fragments back into whole turns, and keeps the speaker names so you can relabel them. If you do want a time reference for finding a passage in the recording later, the option to keep one timestamp at the start of each turn does exactly that.

What is the difference between verbatim and clean verbatim transcription?

True verbatim keeps everything that was said: filled pauses like um and uh, stutters, false starts and repetitions. Clean verbatim removes those disfluencies but keeps every content word and the speaker's own phrasing, which is the convention most qualitative studies use for thematic analysis. Some analyses need true verbatim—conversation analysis, discourse analysis and many linguistic studies treat hesitations as data—so state which convention you used in your methods.

Can I clean an Otter.ai transcript?

Yes. Export from Otter as TXT (with or without timestamps) and paste or drop it here. Otter writes each turn under a header line with the speaker name and a time; the cleaner reads those headers, merges consecutive paragraphs by the same speaker, removes the times, and lets you rename “Speaker 1” and “Speaker 2” as the interviewer and participant.

Is it safe to clean confidential interview transcripts here?

The transcript is processed entirely by JavaScript in your browser; nothing is uploaded and there is no account. You can load the page, disconnect from the internet, and use it offline. For de-identification, the built-in option replaces high- and medium-confidence names, contact details and dates with pseudonyms; for a full review of every detection before sharing a file, use the dedicated transcript de-identifier.

How do I import the cleaned transcript into NVivo, ATLAS.ti or MAXQDA?

Download the .docx or .txt version. The output puts each turn on its own paragraph starting with the speaker label, which all three packages import as a document, and which NVivo and MAXQDA can auto-code by speaker because the label is the first word of each paragraph. For Tagaroo, use “Annotate this in Tagaroo” to open it straight in the annotation workbench, or download the JSON.

Written by Enrique Gutiérrez, PhD (Computer Science)—founder of Tagaroo and Associate Professor of Computer Science, working on inter-rater reliability, measurement and annotation methodology (ORCID).

Last verified: 25 September 2026. Formulas, thresholds and cited figures on this page were checked against their original sources on that date. Every calculation runs in your browser; nothing you enter is transmitted or stored.

The next step is coding it

A clean transcript is where analysis starts. Tagaroo opens it turn by turn, lets you tag passages against your codebook or a clinical rating scale, and can code alongside you—with every AI suggestion carrying its rationale, and agreement with human coders measured as you go.

Try Tagaroo free