Privacy & data prep
Paste an interview transcript and review every candidate identifier the tool finds — names, dates, phone numbers, emails, places, record numbers — then replace them with consistent pseudonyms. The text never leaves your browser.
Free · No sign-up · Runs entirely in your browser
This is an assistive first pass, not a compliance guarantee. Nothing here is legal advice. HIPAA's Safe Harbor method also requires that you have no actual knowledge that the remaining information could identify someone — a judgement about your own study that no pattern matcher can make. Indirect identifiers are invisible to it: a rare diagnosis, an unusual job, a described event, a turn of phrase. Read the whole transcript afterwards, and follow your IRB or ethics committee's approved procedure.
Detection runs in this page. There is no upload step and no request carries your text — open your browser's network tab and watch, or disconnect from the network and keep working.
The names, places and employers specific to your study. Matched literally and case-insensitively — this is how you cover the names the dictionary does not know.
Paste a transcript above, or load the example, and every candidate identifier will be listed here for review.
| §164.514(b)(2)(i) | Category | This tool | What it covers |
|---|---|---|---|
| (A) | Names | Automated pass | Finds common given names, names after a title (Dr Owusu) and names after a relational cue (my brother John). Misses unusual names, most surnames standing alone, and nicknames — add them to the custom list. |
| (B) | Geographic subdivisions smaller than a state | Automated pass | Only postal codes are matched (US ZIP and UK formats). Street addresses, cities, counties, named neighbourhoods and workplaces are not detected; a capitalised word may be flagged as a low-confidence name, nothing more. |
| (C) | All dates except year, and all ages over 89 | Automated pass | Numeric, ISO and written dates are matched, as are ages of 90 and above in the usual phrasings. Ages of 89 and below are deliberately left alone, and a date spelled out in words ("the second Tuesday after Easter") is missed. |
| (D) | Telephone numbers | Automated pass | Matches international, US and UK-style numbers, plus anything after a phone label. A number dictated as words or split across a sentence is missed. |
| (E) | Fax numbers | Automated pass | Found by the same patterns as telephone numbers, and by the word "fax". An unlabelled fax number is indistinguishable from a phone number and is reported as one. |
| (F) | Email addresses | Automated pass | Reliable: the shape of an address is its own evidence. An address dictated as "name at example dot org" is missed. |
| (G) | Social Security numbers | Automated pass | Matches the US 000-00-0000 form. A nine-digit run with no separators, a UK National Insurance number or an NHS number is only found when it follows a label. |
| (H) | Medical record numbers | Automated pass | Matches labelled forms (MRN, chart number, hospital number) and any run of six or more digits, which is where most record numbers land. |
| (I) | Health plan beneficiary numbers | Automated pass | Only when the value follows a label such as "member ID", "policy number" or "beneficiary number". Bare alphanumeric identifiers are missed. |
| (J) | Account numbers | Automated pass | Only when labelled, or when the value is a long run of digits. A short account code embedded in a sentence is missed. |
| (K) | Certificate and licence numbers | Automated pass | Only when the value follows a label such as "licence number" or "certificate number". |
| (L) | Vehicle identifiers, including licence plates | Manual review only | Treat as manual. Plate formats vary by jurisdiction and collide with ordinary words, so nothing recognises an unlabelled plate; only a value written after "VIN" or "licence plate" is caught, incidentally. |
| (M) | Device identifiers and serial numbers | Manual review only | Treat as manual. A serial number has no distinguishing shape; only a value written after an explicit "serial number" or "device ID" label is caught. |
| (N) | Web URLs | Automated pass | Matches http, https and www forms. A bare domain mentioned in passing ("it's on example dot org") is missed. |
| (O) | IP addresses | Automated pass | Matches IPv4 addresses. IPv6 addresses are not matched. |
| (P) | Biometric identifiers, including finger and voice prints | Manual review only | Not detectable from text, by definition. If you are sharing the recording as well as the transcript, the voice itself is an identifier and de-identifying the text does not address it. |
| (Q) | Full-face photographs and comparable images | Manual review only | Not detectable from text. Check any images, screenshots or scanned documents shared alongside the transcript separately. |
| (R) | Any other unique identifying number, characteristic or code | Manual review only | Not detectable, and the category that matters most in qualitative data. A rare diagnosis, an unusual job, a named event, a distinctive turn of phrase — each can re-identify a participant, and only a human who knows the study can see it. |
An automated pass means a detector attempts the category, not that it finds every instance. Category (R) — any other unique identifying characteristic — is the one that matters most in qualitative data and the one only a human can see.
Reference · for the curious
De-identifying a transcript is not the same problem as de-identifying a spreadsheet. In structured data the identifiers live in known columns and can be dropped. In a transcript they are woven through natural speech, they appear in forms nobody anticipated, and the details that actually re-identify a participant are frequently not identifiers at all — they are circumstances. That gap is why this tool is deliberately built as a first pass with a review step, and not as a button that promises a compliant transcript.
It scans the text you paste for patterns that correspond to direct identifiers — personal names, dates, telephone numbers, email addresses, URLs and IP addresses, US social security numbers, medical record and account numbers, postal codes, and ages above 89 — and presents every match for you to accept or reject before anything is replaced. Identifiers without a distinctive shape are only found when a label gives them away: a UK National Insurance number or an NHS number is caught after words like "insurance number", not on its own. Accepted matches are substituted with consistent pseudonyms, so one person becomes the same label everywhere in the transcript.
It runs entirely in your browser. The transcript is never uploaded, there is no account, and the page makes no network request with your text. For clinical material that property is the whole point: pasting a patient interview into a remote service is a disclosure, whatever the service's privacy policy says.
The Safe Harbor method under 45 CFR §164.514(b)(2) treats data as de-identified once 18 categories of identifier have been removed and the covered entity has no actual knowledge that the remaining information could identify the individual. Automated pattern matching can help with most of the first condition. It cannot address the second at all.
Several of the 18 categories are simply not detectable in a text transcript — biometric identifiers, photographs, and the catch-all "any other unique identifying number, characteristic, or code". That last category is the one that matters most in qualitative work, and it is where a human reader is irreplaceable.
A transcript can be entirely free of names and dates and still identify someone unambiguously. A rare diagnosis in a small clinic. An unusual occupation in a named region. A described event that made the local news. The combination of an age, a job and a family structure. None of these trip a regular expression, and all of them have re-identified research participants in practice.
This is why the review step in the tool above is not a formality and why no automated tool should be trusted as a final answer. Read the whole transcript. Ask whether someone who knows the participant's community would recognise them. Where the answer is yes, generalise rather than delete: "a rare autoimmune condition" instead of the diagnosis, "a manufacturing town in the north" instead of the city.
For analysis, replacing each person with a consistent label is almost always better than removing them. Qualitative data is largely about relationships and sequences — who did what to whom, who is mentioned repeatedly, who appears only once. Deletion destroys that structure and often makes passages unreadable. Consistent pseudonyms preserve the account while removing the identity, which is why the tool assigns them by default.
Keep the mapping between real names and pseudonyms in a separate, access-controlled document, and never store it alongside the de-identified transcripts. The pseudonym key is the re-identification key; treat it with the care you would give the original recording.
Safe Harbor requires removing all date elements more specific than the year where they relate to an individual. Applied mechanically, that can destroy a clinical narrative — the sequence and interval between events is frequently the analytically important part. The usual solution is date shifting: move every date for a participant by the same random offset, preserving intervals while breaking the link to real calendar time. Decide your approach before you start and apply it uniformly across the dataset.
Detecting names in text without a language model is genuinely hard, and the tool is honest about it. It reliably finds common given names and names introduced by a title or a relational cue. It misses unusual names, most surnames used alone, and nicknames. It cannot distinguish a surname from an ordinary capitalised word, so bare capitalised tokens are flagged with low confidence rather than replaced silently.
The practical workaround is the custom-terms field: add the names, place names, institutions and study-specific terms you know appear in your data, and they will be matched exactly. Then read the transcript. A tool that flagged every capitalised word would be useless; one that flags a defensible subset and tells you to check is not.
Transcripts carry speaker turns — I: and S:, or "Interviewer:" and "Participant 3:". These are not identifiers and removing them would destroy the transcript's shape, so the tool leaves them alone. If your transcripts label speakers by real name, add those names to the custom-terms list; the replacement will keep the turn structure while anonymising the label.
The terms are often used interchangeably and mean different things. De-identified data has had identifiers removed but may still be re-identifiable, and under GDPR it generally remains personal data — pseudonymised, not anonymous, and still within scope of the regulation. Genuine anonymisation is a higher bar and, for rich qualitative transcripts, frequently unachievable without destroying the data's usefulness. Plan on the basis that your de-identified transcripts are still sensitive, because they are.
Our guides to consent and licensing for annotation data and preparing clinical data for annotation cover the regulatory picture, and clinical transcript annotation tools discusses when to de-identify first versus work under a business associate agreement.
Run the transcript through this tool and review every hit. Add study-specific terms and run it again. Read the output in full, looking for indirect identifiers, and generalise them by hand. Have a second person read it — the same blind spots that let you write an identifying detail let you miss it. Record what you did, because your ethics committee will ask. Then store the pseudonym key separately from everything else.
Nothing on this page is legal advice, and this tool is not a compliance control. Follow the procedure your IRB or ethics committee has approved; where they differ from anything here, they govern.
Once a transcript is clean enough to share, the work is coding it. If you plan to report inter-rater reliability on those codes, the inter-rater reliability calculator computes every coefficient side by side, and the sample-size planner will size the double-coding you need.
The tool is an aid, not a compliance guarantee, and nothing here constitutes legal advice. HIPAA's Safe Harbor method requires removing 18 specified categories of identifier and having no actual knowledge that the remaining information could identify someone. Automated pattern matching finds many identifiers but cannot meet that second condition, and it cannot recognise the indirect details — a rare diagnosis, an unusual job, a described event — that often re-identify a participant in qualitative data. Treat the output as a first pass that a human must read in full, and follow your IRB or ethics committee's approved procedure.
No. Detection and replacement run in JavaScript inside your browser, and the page makes no network requests with your text. That is the reason this tool exists in this form: transcripts of clinical or research interviews are exactly the material you should not paste into a remote service. You can verify it by opening your browser's network tab, or by disconnecting from the network and continuing to use the tool.
Names; geographic subdivisions smaller than a state, including street address, city, county and most ZIP codes; all dates more specific than a year that relate to an individual, plus all ages over 89; telephone numbers; fax numbers; email addresses; Social Security numbers; medical record numbers; health plan beneficiary numbers; account numbers; certificate or licence numbers; vehicle identifiers and licence plates; device identifiers and serial numbers; web URLs; IP addresses; biometric identifiers including fingerprints and voiceprints; full-face photographs and comparable images; and any other unique identifying number, characteristic or code.
For qualitative analysis, consistent pseudonyms are usually better than deletion. Replacing every mention of one person with the same label preserves the structure of the account — who did what to whom, who is referred to repeatedly — which is often central to the analysis, while removing the identity. Deletion flattens that structure and can make a transcript unreadable. Keep the pseudonym key in a separate, secured document, and never store it with the de-identified transcript.
Name detection without a language model is heuristic: it relies on capitalisation, position and a dictionary of common given names, so it reliably finds ordinary first names but misses unusual ones, misreads capitalised sentence openings, and cannot tell a surname from an ordinary capitalised word. Add the names and terms specific to your study to the custom list, which is matched exactly, and read the whole transcript afterwards. The review step is not optional.
Written by Enrique Gutiérrez, PhD (Computer Science) — founder of Tagaroo and Associate Professor of Computer Science, working on inter-rater reliability, measurement and annotation methodology (ORCID).
Last verified: 29 July 2026. Formulas, thresholds and cited figures on this page were checked against their original sources on that date. Every calculation runs in your browser; nothing you enter is transmitted or stored.
Tagaroo annotates clinical and research transcripts against validated scales and coding schemes, with every rating linked to the exact span that justifies it — and inter-rater reliability computed as your team works.