forensic
Criteria-Based Content Analysis: The 19 CBCA Criteria
Criteria-based content analysis (CBCA) rates 19 content criteria to judge whether a statement reads as experience-based—an aid, not a lie test. See all 19.

Every account of a real event carries a texture that a rehearsed story struggles to fake: a scattered order, a quoted aside, a moment where the speaker stops and corrects a detail. Criteria-based content analysis (CBCA) is the attempt to turn that intuition into a checkable list. It rates a statement against 19 content criteria and asks a narrow question—does this account show the features that research links to experience-based memory?—without ever claiming to catch a liar.
That last point is the one most write-ups get wrong. CBCA is the analytic core of a forensic procedure called Statement Validity Assessment, and it was built to assist a credibility evaluation, not to replace one. Used as a lie detector, it fails; used as a structured lens on what a statement contains, it has held up across decades of study (Vrij, 2005).
What is criteria-based content analysis?
Criteria-based content analysis is a structured technique for judging whether the content of a statement shows features that research associates with truthful, experience-based accounts (Steller & Köhnken, 1989). An analyst reads a transcript of a free-narrative account and rates it against 19 defined criteria; the more criteria present, and the more strongly they appear, the more the statement resembles the profile of a genuine recollection.
The method comes out of German forensic psychology, where it was developed to help courts weigh the testimony of children in suspected sexual-abuse cases—situations where the child’s statement is often the central, and sometimes the only, evidence. Max Steller and Günter Köhnken compiled the modern 19-criterion set in 1989, consolidating earlier German work into a single system that could be taught and applied consistently (Steller & Köhnken, 1989).
One distinction keeps the scope honest. CBCA is a credibility-assessment aid, not a truth machine. A high CBCA score means the account has the hallmarks of experience-based memory; it does not prove the event happened, and a low score does not prove fabrication. That gap between “looks experience-based” and “is true” is where careful practice lives, and where careless use goes wrong.
The Undeutsch hypothesis
CBCA rests on a single idea, the Undeutsch hypothesis: a statement that comes from memory of a real, self-experienced event differs in content and quality from one that is invented or imagined (Undeutsch, 1967). Udo Undeutsch argued that lived experience leaves distinctive traces in how people talk about it—unplanned detail, sensory specifics, unflattering admissions—that are cognitively hard to manufacture on demand.
The 19 criteria are the operational form of that claim. Each one names a feature that, on the hypothesis, should appear more often in truthful than in fabricated accounts. A meta-analytic review of the hypothesis found a large overall effect for CBCA total scores—one its authors report as generalizable to the total score—even as they flagged that several individual criteria discriminate weakly.
Where CBCA fits: the three components of Statement Validity Assessment (SVA)
CBCA is one stage of a larger procedure, and reading it in isolation is the classic error. Statement Validity Assessment (SVA) is the full forensic method: three components wrapped around a preparatory step (Vrij, 2005).
The preparatory step is case-file analysis—the evaluator reviews the file first to understand the allegation and generate specific hypotheses to test, including the hypothesis that the statement is not based on experience. The three components proper then follow:
- A semi-structured interview. A trained interviewer elicits a free-narrative account with open, non-leading questions, so the statement to be analyzed is as complete and uncontaminated as possible. Interviewing quality is not a side issue here; it determines whether there is anything worth analyzing.
- CBCA of the transcript. The interviewer’s transcript is scored against the 19 content criteria—the step this article is about.
- The Validity Checklist. Finally, the evaluator weighs alternative explanations for the CBCA outcome: the witness’s age and verbal skill, possible coaching or motives, the interview’s own suggestiveness, and consistency with other evidence. A rich account from a heavily-led interview is not the same as a rich account from an open one.
The Validity Checklist is what turns a content score into a cautious opinion, and it is also the least-researched part of the method (Vrij, 2005). Skipping it—treating a CBCA tally as a verdict—strips away every safeguard the procedure was designed to carry.
What are the 19 CBCA criteria?
The 19 CBCA criteria are the checkable features an analyst rates in a statement, and they are the reason the method can be applied consistently rather than by gut feel. Steller and Köhnken (1989) organize them into families. The canonical scheme uses five categories; these are often summarized as four broad groups—general characteristics, specific contents (which absorbs the finer “peculiarities of content” set), motivation-related contents, and offense-specific elements. The table below lists all 19 with a one-line definition each, keeping the five-family structure so nothing is lost.
| Category | # | Criterion | What it means (one line) |
|---|---|---|---|
| General characteristics | 1 | Logical structure | The account is coherent and internally consistent; the parts fit together without contradiction. |
| General characteristics | 2 | Unstructured production | Information comes out scattered and non-chronological, as spontaneous recall tends to, not in a rehearsed order. |
| General characteristics | 3 | Quantity of details | The account is rich in specific detail of place, time, people, objects, and actions. |
| Specific contents | 4 | Contextual embedding | Events are anchored in a time and place and tied to everyday routines and circumstances. |
| Specific contents | 5 | Descriptions of interactions | The account contains chains of action and reaction between the people involved. |
| Specific contents | 6 | Reproduction of conversation | Speech from the event is reported, ideally quoting distinct speakers in their own words. |
| Specific contents | 7 | Unexpected complications | The account includes an unplanned interruption or complication during the incident. |
| Peculiarities of content | 8 | Unusual details | Details that are uncommon or idiosyncratic yet plausible. |
| Peculiarities of content | 9 | Superfluous details | Peripheral details not needed for the allegation, reported anyway. |
| Peculiarities of content | 10 | Accurately reported details misunderstood | A detail is reported correctly but its meaning is misinterpreted (typical of a child describing an adult act). |
| Peculiarities of content | 11 | Related external associations | Reference to a related event or conversation outside the incident itself. |
| Peculiarities of content | 12 | Accounts of subjective mental state | The witness describes their own feelings or thoughts during the event. |
| Peculiarities of content | 13 | Attribution of perpetrator's mental state | The witness describes or infers the other person's feelings, thoughts, or motives. |
| Motivation-related contents | 14 | Spontaneous corrections | The witness revises or corrects a detail without being prompted. |
| Motivation-related contents | 15 | Admitting lack of memory | The witness says they do not remember a detail rather than filling the gap. |
| Motivation-related contents | 16 | Raising doubts about one's own testimony | The witness questions whether their own account is convincing or correct. |
| Motivation-related contents | 17 | Self-deprecation | The witness reports self-incriminating or unflattering details about themselves. |
| Motivation-related contents | 18 | Pardoning the perpetrator | The witness excuses the other person or avoids blaming them. |
| Offense-specific elements | 19 | Details characteristic of the offense | Content matching what is professionally known about how this type of offense typically unfolds, even against lay expectations. |
A few notes on reading the list. The criteria are not equally common: general characteristics like quantity of details appear in almost any substantial account, while criteria such as pardoning the perpetrator are rare and, in the samples this meta-analysis pooled, among the weakest discriminators (Amado, Arce & Fariña, 2015). CBCA is scored by presence and strength, not by treating the list as a 19-point quiz where a higher total automatically means “more true.”
How does CBCA scoring work in practice?
To apply CBCA, an analyst reads the transcript and marks where each criterion appears, then rates how strongly it is present—commonly on a 0–2 scale (absent, present, strongly present). The point of tagging each instance is auditability: another analyst can check the call against the exact words rather than the first analyst’s impression.
Consider a short synthetic account (invented for illustration; no real testimony):
Witness: We were in the back office after closing, and he was moving the boxes—no, wait, he’d already stacked them, and then he knocked one over. He kept muttering that the manager would blame him. I remember thinking I should just leave. It was the blue crate, the one with the cracked handle.
Even in three sentences, several criteria surface: unstructured production and a spontaneous correction (“no, wait”), reproduction of the muttered speech, an account of subjective mental state (“I should just leave”), attribution of his mental state (“the manager would blame him”), and an unusual, superfluous detail (“the blue crate… with the cracked handle”). None of that proves the event occurred. It shows the account carries features that experience-based recall tends to produce—which is exactly, and only, what CBCA is designed to surface.
This instance-level tagging is the same discipline used in other discourse-coding schemes. Marking who-said-what and how a story is structured is what the Labov & Waletzky narrative structure model does for oral narratives, and coding utterance by utterance is how the NICHD protocol’s prompt types audit an interviewer’s questions.
Does criteria-based content analysis detect lies?
No. Criteria-based content analysis does not detect lies and does not determine truth—and treating it as if it does is the single most consequential misuse of the method. This is the sharp edge of the scientific and legal debate, and it deserves to be stated plainly rather than buried.
Start with accuracy. In laboratory research, CBCA classifies truthful and fabricated accounts correctly around 70% of the time, which means it misclassifies roughly 30% in both directions (Vrij, 2005). That is well above chance, and genuinely useful as one input among several. It is nowhere near the certainty a court needs from a piece of evidence offered as proof of veracity.
That error rate is why admissibility is contested. In his qualitative review of the first 37 CBCA studies, Aldert Vrij concluded that CBCA’s error rates are too high for it to be used as the sole indicator of veracity, and that Statement Validity Assessment does not meet the Daubert standard for the admissibility of scientific evidence in the United States (Vrij, 2005). In several jurisdictions, CBCA conclusions are not admitted as standalone proof, precisely because of that error rate and the thin research base under the Validity Checklist. The method is most established in the German-speaking forensic tradition where it originated, and even there it functions as expert-assisted evaluation, not a mechanical truth test.
Two more honest caveats round this out. First, individual criteria vary. A meta-analysis found strong, generalizable overall support for the Undeutsch hypothesis (δ = 0.79 for the total score), yet several motivation-related criteria—such as self-deprecation and pardoning the perpetrator—were weak discriminators even in the child samples it pooled (Amado, Arce & Fariña, 2015).
Second, CBCA is sensitive to who is speaking. Verbally skilled adults and coached witnesses can produce high-scoring accounts, and young or less verbal children can produce genuine accounts that score low. The Validity Checklist exists to catch these confounds, which is why CBCA without it is not really CBCA.
Tagging CBCA criteria for research and training
For research and training—not legal determinations—CBCA is a natural fit for span-level annotation: mark each place in a transcript where a criterion appears, attach the criterion label and a strength rating, and you get a reviewable, countable record instead of a one-line impression. Coders in training can compare their tags against a reference, and researchers can compute how reliably two analysts agree on each criterion.
That reliability question matters, because criteria differ in how consistently people can spot them. Concrete, surface features like reproduction of conversation are easier to agree on than interpretive ones like details characteristic of the offense. Computing inter-rater reliability with Cohen’s kappa per criterion is the honest way to show which parts of a coding scheme are solid and which are shaky—and CBCA has both.
One point of reconciliation, so the numbers here line up. Tagaroo’s Scale Library currently curates 8 of the 19 CBCA criteria—Logical Structure, Unstructured Production, Quantity of Details, Contextual Embedding, Reproduction of Conversation, Accounts of Subjective Mental State, Spontaneous Corrections, and Admitting Lack of Memory—so the card above shows “8 items” while this article documents the full 19; the remaining criteria are being added as the library expands. The eight shipped first are the concrete, high-frequency criteria that annotate most reliably, which is deliberate. Whatever the tool, the same guardrail holds: this is for research and training annotation, not for reaching legal conclusions about a real person.
One caution on data handling belongs here. Real statements in credibility cases—especially those involving children—are among the most sensitive records that exist, governed by strict legal and safeguarding controls. No annotation workflow replaces those controls.
The sane default is de-identified, synthetic, or already-lawfully-held text. Tagaroo supports a browser-side anonymous mode so content can stay local, the same privacy discipline that applies to any clinical transcript annotation work. And where CBCA sits next to the interview itself, the NICHD investigative-interview prompt types post covers the questioning side that produces a codable statement in the first place.
Common mistakes when using CBCA
Most errors trace back to forgetting what the method is for:
- Treating a criteria count as a verdict. CBCA describes content; it does not decide truth. The total is an input to a cautious opinion, not the opinion (Vrij, 2005).
- Dropping the Validity Checklist. Without it, age, verbal skill, coaching, and a leading interview all masquerade as signal. CBCA scores mean little detached from the checklist (Vrij, 2005).
- Ignoring the interview quality. A suggestive interview can inflate detail and conversation criteria. Garbage in, high score out.
- Over-summing rare criteria. Offense-specific and motivation-related criteria are rare and, in this meta-analysis’s child samples, weak discriminators—counting them like the common ones distorts the picture (Amado, Arce & Fariña, 2015).
- Quoting a single accuracy number as if it settled admissibility. The defensible framing is a moderate discriminator with a ~30% lab error rate that does not meet the Daubert standard—not “X% accurate” (Vrij, 2005).
The practical upshot: criteria-based content analysis earns its place as a disciplined way to describe what a statement contains, and loses it the moment anyone treats the 19 criteria as a lie test. Code the content carefully, keep the Validity Checklist and the interview in the frame, and use the criteria for what they are—an investigative aid, never the last word.
References
- Steller, M., & Köhnken, G. (1989). Criteria-Based Content Analysis. In D. C. Raskin (Ed.), Psychological Methods in Criminal Investigation and Evidence (pp. 217–245). Springer.
- Undeutsch, U. (1967). Beurteilung der Glaubhaftigkeit von Zeugenaussagen. In Handbuch der Psychologie, Vol. 11: Forensische Psychologie (pp. 26–181). Hogrefe.
- Undeutsch, U. (1982). Statement reality analysis. In A. Trankell (Ed.), Reconstructing the Past: The Role of Psychologists in Criminal Trials (pp. 27–56). Kluwer.
- Vrij, A. (2005). Criteria-Based Content Analysis: A qualitative review of the first 37 studies. Psychology, Public Policy, and Law, 11(1), 3–41. doi:10.1037/1076-8971.11.1.3
- Amado, B. G., Arce, R., & Fariña, F. (2015). Undeutsch hypothesis and Criteria Based Content Analysis: A meta-analytic review. European Journal of Psychology Applied to Legal Context, 7(1), 3–12. doi:10.1016/j.ejpal.2014.11.002
- Lamb, M. E., Orbach, Y., Hershkowitz, I., Esplin, P. W., & Horowitz, D. (2007). A structured forensic interview protocol improves the quality and informativeness of investigative interviews with children. Child Abuse & Neglect, 31(11–12), 1201–1231. doi:10.1016/j.chiabu.2007.03.021
If you code statements for research or interviewer training, Tagaroo turns the CBCA criteria in its Scale Library into a guided, source-anchored annotation workflow—with inter-rater reliability computed as your coders work, and criteria for research and training use, never legal determinations.
Frequently asked questions
- What is criteria-based content analysis (CBCA)?
- Criteria-based content analysis is a structured method for judging whether the content of a statement shows features that research associates with truthful, experience-based accounts. An analyst rates a transcript against 19 content criteria—such as quantity of detail, reproduction of conversation, and admitting lack of memory (Steller & Köhnken, 1989). CBCA is the content-analysis stage of a wider procedure called Statement Validity Assessment, and it is an investigative aid, not a standalone test of truth or deception (Vrij, 2005).
- What are the 19 CBCA criteria?
- The 19 criteria fall into five families: general characteristics (logical structure, unstructured production, quantity of details); specific contents (contextual embedding, descriptions of interactions, reproduction of conversation, unexpected complications); peculiarities of content (unusual details, superfluous details, accurately reported details misunderstood, related external associations, accounts of subjective mental state, attribution of the perpetrator's mental state); motivation-related contents (spontaneous corrections, admitting lack of memory, raising doubts about one's own testimony, self-deprecation, pardoning the perpetrator); and one offense-specific element (details characteristic of the offense) (Steller & Köhnken, 1989).
- Does CBCA detect lies?
- No. CBCA does not detect lies and does not determine truth. It is a qualitative aid that flags whether an account contains features more typical of experience-based statements (Vrij, 2005). In laboratory research CBCA misclassifies roughly 30% of both truthful and fabricated accounts, an error rate high enough that Statement Validity Assessment does not meet the Daubert standard for scientific evidence in the United States, and its conclusions are not admitted as standalone proof in several jurisdictions (Vrij, 2005).
- What is the Undeutsch hypothesis?
- The Undeutsch hypothesis holds that a statement based on memory of a real, self-experienced event differs in content and quality from a statement that is fabricated or imagined (Undeutsch, 1967). CBCA operationalizes that idea as 19 checkable content features. A meta-analysis of the hypothesis found a large overall effect (δ = 0.79) for CBCA total scores discriminating true from fabricated accounts—an effect its authors report as generalizable to the total score—while several individual motivation-related criteria discriminated weakly (Amado, Arce & Fariña, 2015).
- How is CBCA different from the NICHD protocol?
- They code different things. CBCA scores the content of a witness's statement against 19 criteria to assess whether it reads as experience-based (Steller & Köhnken, 1989). The NICHD investigative interview protocol codes the interviewer's prompts—invitation, directive, option-posing, suggestive—to keep questioning open and non-leading (Lamb et al., 2007). In practice they are complementary: a clean, open interview is what produces a statement worth analyzing in the first place.
Put this into practice
Tagaroo turns any rating scale or coding scheme into a guided annotation workflow — with inter-rater reliability computed as you go.