Qualitative rigor

Thematic saturation calculator

“Data were collected until saturation” is the most common unsupported sentence in qualitative methods sections. Paste your coding to show where saturation happened and how strict that claim is—or justify a sample size before you collect anything.

Free · No sign-up · Runs entirely in your browser

Where your analysis stopped producing new themes

Paste the number of new themes each interview added, or a theme × interview matrix exported from NVivo, ATLAS.ti or MAXQDA. You get the saturation point in the notation of Guest, Namey and Chen (2020), how it changes under stricter settings, and a paragraph for your methods section.

Your coding, in analysis order

Base size

The first 4 interviews define the themes new ones are measured against. The paper tested 4, 5 and 6.

4

Run length

How many consecutive further interviews must add little or nothing.

New-information threshold

The share of the base's themes a run may add and still count as saturated.

Unit

Saturation at ≤5% new information · base 4, run 2

6⁺²

Interviews 7–8 added 1 new theme — 2.7% of the 37 in the base — so saturation was reached at interview 6 and confirmed by the next 2.

Base themes

37

Total themes

51

Interviews

20

New themes per interview Cumulative unique themes
base6⁺²1234567891011121314151617181920interview (analysis order)2051

Hover a bar for its counts. The shaded band is the confirming run.

Saturation ratio of each run

New themes in the run ÷ themes in the base. The dashed line is your threshold.

run 5–6run 19–20

How strict is your claim?

The paper's four reporting options at base size 4. Stricter settings mean more confidence that few themes remain undiscovered. Reporting the strictest one you meet is more persuasive than a single "saturation was reached".

≤5% new information0% new informationRun of 2
6⁺²saturated at interview 6
10⁺²saturated at interview 10
Run of 3
9⁺³saturated at interview 9
—not reached

Does the order of analysis matter?

Saturation is measured in the order you analysed the interviews, and a different order can move it. Guest and colleagues validated the method by re-running it on random orderings. That needs to know which themes each interview contained, so paste a theme × interview matrix instead of counts to test it.

Methods paragraph

Thematic saturation was assessed following Guest, Namey and Chen (2020), with a base size of 4 interviews (37 unique themes), a run length of 2 and a new-information threshold of ≤5%. The threshold was first met by interviews 7–8, which contributed 1 new theme (2.7% of the base), indicating saturation at 6⁺² interviews; 20 interviews were conducted in total, yielding 51 unique themes. Under the other reporting options, saturation was 10⁺² (run 2, 0%); 9⁺³ (run 3, ≤5%); not reached (run 3, 0%).

Reference · for the curious

Saturation and sample size in qualitative research: what the evidence supports

Saturation is the most common justification for a qualitative sample size and the least often demonstrated. Morse's complaint from 1995, that saturation is "evident mainly by declaration", still describes most methods sections. The tool above turns the declaration into a reported measurement, and the planning tab covers the earlier question every ethics committee asks: how many participants, and why.

The Guest, Namey and Chen method

Guest, Namey and Chen (2020, PLOS ONE) proposed three parameters that make a saturation claim specific. The base size is how many initial interviews define the themes already known; they tested 4, 5 and 6. The run length is how many consecutive further interviews are examined for new themes, 2 or 3. The new-information threshold is how much novelty a run may add and still count as saturated: ≤5% or 0% of the base's themes.

Their worked example makes the arithmetic concrete. The first four interviews produce 37 themes. Interviews 5 and 6 add seven, a ratio of 19%; interviews 6 and 7 add four, 11%; interviews 7 and 8 add one, 2.7%, which is under the 5% threshold. Saturation is therefore reported at 6⁺²—reached at interview 6 and confirmed by two more. At the stricter 0% threshold the same data saturate at 10⁺². That example is loaded in the tool by default, reconstructed from every figure the paper states.

Two properties make the method more defensible than older approaches. The denominator is the base, not the whole dataset, so saturation is not guaranteed by construction the way it is when new themes are expressed as a share of all themes eventually found. And it works prospectively: you can compute it after each interview and stop recruiting when the threshold is met, rather than only after analysis is finished.

Report the setting, not just the word

Saturation at 6⁺² with a 5% threshold is a weaker claim than 10⁺³ at 0%, and readers can only weigh it if they are told which one it is. The grid in the tool shows all four of the paper's reporting options for your data, and the strictest one you meet is the most persuasive to report. State the base size too: a large base makes the ratio smaller and saturation easier to reach.

Order matters, so test it

The method counts themes in the order interviews were analysed. A dataset where the most talkative participants happened to come first will saturate early; the same interviews in a different order may not. Guest and colleagues addressed this by bootstrapping random orderings of three real datasets. With a theme × interview matrix—the code-by-document matrix NVivo, ATLAS.ti and MAXQDA all export—the tool repeats the assessment on 1,000 seeded random orders and reports how often saturation arrived as early as it did in yours. A figure above 80% is reassuring; a figure near 50% means your saturation point is partly an accident of scheduling.

Before data collection: justifying the number

Three published aids are worth citing in a protocol, and the planning tab combines them.

  • Empirical benchmarks. Hennink and Kaiser's (2022) systematic review of 23 studies found saturation within 9–17 individual interviews or 4–8 focus group discussions—mostly for homogeneous samples with narrow objectives, and mostly for code saturation, which is reached earlier than the fuller "meaning saturation".
  • Information power. Malterud, Siersma and Guassora (2016) argue that the more relevant information a sample holds, the fewer participants it needs, and name five dimensions that decide it: the breadth of the aim, the specificity of the sample, the use of established theory, the quality of dialogue, and case versus cross-case analysis. The model deliberately gives no formula, so the tool does not invent one; it turns your ratings into the argument.
  • A binomial bound. Fugard and Potts (2015) compute how many participants are needed to observe a theme of a given prevalence a given number of times. Their example—29 participants for an 80% chance of seeing a theme held by 10% of the population at least twice—is reproduced exactly. It assumes random sampling, which purposive qualitative samples are not, so it is a sanity check on the order of magnitude rather than a requirement.

When saturation is the wrong criterion

Saturation assumes that themes are discovered in the data and that, after enough interviews, there are no more to find. That fits codebook and coding-reliability approaches, where a stable set of codes is applied across transcripts, and it is the assumption behind every number on this page. It fits reflexive thematic analysis poorly: Braun and Clarke argue that themes are generated by the analyst's engagement with the data, so there is no fixed set to run out of. If that is your approach, use the information power argument and leave the saturation metric out.

Saturation also says nothing about whether themes were coded consistently. If two people applied the codebook, the intercoder reliability calculator pools their agreement from the same NVivo, ATLAS.ti or MAXQDA export, and the codebook generator helps build the definitions that make codes applicable the same way twice. For quantitative reliability studies, where the question is how many subjects are needed to estimate kappa or an ICC precisely, the reliability sample-size planner is the right tool instead.

What this tool does not do

It counts themes, not their depth: two interviews that each mention a theme once contribute the same as one that explores it at length, which is why code saturation arrives earlier than meaning saturation. It trusts the order you give it, so paste interviews in the order they were analysed, not alphabetically. And it cannot tell you whether a theme is important; a single new theme that changes the answer to your research question matters more than any ratio.

Frequently asked questions

How do you calculate thematic saturation?

Guest, Namey and Chen (2020) give the most widely used quantitative method. Count the unique themes in your first few interviews—the base, usually 4 to 6. Then look at the next run of 2 or 3 interviews and count the new themes they add. Divide new themes by base themes: that is the saturation ratio. Slide the run forward one interview at a time until the ratio falls to your threshold, typically ≤5% or 0%. Saturation is reported at the interview before that run, with the run length as a superscript: 6⁺² means saturation at interview 6, confirmed by two more.

How many interviews do I need for saturation?

Hennink and Kaiser's (2022) systematic review of 23 studies found that empirical studies reached saturation within 9–17 individual interviews or 4–8 focus groups—but mostly in homogeneous samples with narrowly defined aims. Heterogeneous samples, several sites, or aiming for the full meaning of each code rather than its mere presence all push the number up. Treat the range as a benchmark for a protocol, and assess saturation in your own data during analysis rather than assuming it.

What is information power in qualitative research?

Information power is Malterud, Siersma and Guassora's (2016) alternative to saturation for planning a sample size. The more information the sample holds that is relevant to the aim, the fewer participants are needed. Five dimensions decide it: a narrow or broad aim, a dense or sparse sample specificity, whether established theory is applied, the quality of dialogue, and whether analysis is case-based or cross-case. The model gives no formula; it structures the argument you make for your sample size.

What does 6⁺² mean in a saturation statement?

It is Guest and colleagues' notation. The number is the interview at which saturation was reached; the superscript is the run length—the number of additional interviews that were conducted to confirm that little or no new information was emerging. 6⁺² therefore means eight interviews in total, with saturation at the sixth. Stating the base size and the threshold alongside it (for example, base size 4, ≤5% new information) makes the claim reproducible.

Does the order of interviews affect saturation?

Yes. The method measures new themes in the order you analyse interviews, and a different order can move the saturation point. Guest and colleagues validated the approach by bootstrapping random orderings of three real datasets. This tool does the same on your data when you paste a theme-by-interview matrix: it re-runs the assessment on 1,000 random orders and reports how often saturation arrives as early as it did in yours.

Should reflexive thematic analysis report saturation?

Usually not. Braun and Clarke argue that saturation assumes themes are waiting to be found in the data, which conflicts with reflexive thematic analysis, where themes are generated by the analyst. If you use reflexive TA, justify your sample with information power or with the depth your question needs. Saturation metrics fit codebook and coding-reliability approaches, where a stable set of codes is applied across interviews.

Written by Enrique Gutiérrez, PhD (Computer Science)—founder of Tagaroo and Associate Professor of Computer Science, working on inter-rater reliability, measurement and annotation methodology (ORCID).

Last verified: 24 September 2026. Formulas, thresholds and cited figures on this page were checked against their original sources on that date. Every calculation runs in your browser; nothing you enter is transmitted or stored.

Let the coding keep the saturation record for you

The theme × interview matrix this tool needs is a by-product of coding in Tagaroo: every transcript you or the AI codes updates it, so the discovery curve and the saturation point are current the moment you finish an interview—not reconstructed from memory at the write-up.

Try Tagaroo free