tagaroo

forensic

CBCA vs Reality Monitoring: Two Credibility Methods

CBCA vs reality monitoring: how two statement-credibility methods compare on criteria, accuracy, and court limits—and why neither one detects lies.

Enrique Gutiérrez15 min readUpdated July 2026
A single block of abstract statement-marks viewed through two side-by-side analytic lenses, one casting a long checklist of ticks and the other a short checklist whose single coral mark points the opposite way.

Two people describe an event they say happened to them. One is telling the truth; one made it up. Read only the words—no polygraph, no demeanor, no case file—and the question both criteria-based content analysis and reality monitoring try to answer is the same: does this account carry the fingerprints of a real, remembered experience, or of an invented one? They just look for different fingerprints.

That shared goal is why CBCA vs reality monitoring is a real comparison and not a false one. Both are structured, criterion-by-criterion ways to read a statement’s content. And both share the single most important caveat in this field: neither is a lie detector. Neither outputs “true” or “false.” Used as if they did, both fail (Masip et al., 2005; Vrij, 2005).

What is the difference between CBCA and reality monitoring?

The difference between CBCA and reality monitoring is where each method’s criteria come from. CBCA is built on a forensic hypothesis about lived experience; reality monitoring is built on an experimental model of how memory encodes its own source. Same question, two different theories of what a genuine account looks like.

Criteria-based content analysis (CBCA) is the content-analysis core of Statement Validity Assessment, a procedure developed in German forensic psychology to help courts weigh children’s testimony in suspected abuse cases. Max Steller and Günter Köhnken consolidated the modern 19-criterion set in 1989, operationalizing Udo Undeutsch’s hypothesis that a statement drawn from memory of a real event differs in content and quality from one that is fabricated (Undeutsch, 1967; Steller & Köhnken, 1989). If you want the full 19, we cover them in the CBCA criteria deep dive.

Reality monitoring (RM) comes from a different building. Marcia Johnson and Carol Raye’s 1981 model in Psychological Review described how people decide whether a memory came from an external source (something perceived) or an internal one (something imagined or reasoned). Their claim: externally derived memories tend to carry more perceptual, spatial, temporal, and affective detail, while internally generated ones carry more traces of the cognitive operations that produced them (Johnson & Raye, 1981). Researchers—chiefly María Luisa Alonso-Quecuty in Spain and Siegfried Sporer in Germany—later extrapolated that memory model into a tool for separating truthful from invented statements (Sporer, 2004).

CBCA vs reality monitoring at a glance

At a glance, the two methods differ in origin, criterion count, and, most consequentially, in whether they can flag a lie or only affirm the truth. The table below lines them up on the dimensions that actually change how you’d use one.

DimensionCBCAReality Monitoring (RM)
RootsGerman forensic psychology; the Undeutsch hypothesis (Undeutsch, 1967)Experimental cognitive psychology; source-monitoring memory model (Johnson & Raye, 1981)
Consolidated bySteller & Köhnken (1989)Alonso-Quecuty; Sporer (JMCQ operationalization, 2004)
Core premiseReal experiences produce richer, more detailed content than fabricationsPerceived memories carry different detail than imagined ones
Number of criteria19 content criteria~8 memory-characteristic criteria
Direction of criteriaAll point toward truthfulness (presence = more experience-based)7 point toward truth; cognitive operations points toward deception
Sits insideStatement Validity Assessment (with a Validity Checklist)No formal wrapper; typically a standalone content score
Ease of useMore criteria, more training, more timeFewer criteria; faster and easier to teach (Vrij et al., 2004)
Lab accuracy~70% correct, ~30% error (Vrij, 2005)Above chance, similar to CBCA, ~30% error (Masip et al., 2005)
Court statusSVA fails the Daubert standard (Vrij, 2005)Cautioned against for applied forensic use (Masip et al., 2005)
CBCA vs reality monitoring on the dimensions that change practice. Both are investigative aids; neither is a test of truth or deception.

The theoretical split: richness versus memory source

Here’s where it gets subtle. CBCA and reality monitoring can flag many of the same words, but they justify the flag with opposite logics, and that shapes what each method can and can’t claim.

CBCA reads for richness. On the Undeutsch hypothesis, a truthful account should be denser in the kinds of detail that are cognitively hard to invent on demand: quoted conversation, unstructured production, spontaneous corrections, subjective mental states, unexpected complications. A meta-analytic review found a large overall effect for CBCA total scores discriminating true from fabricated accounts (δ = 0.79), even as several individual criteria discriminated weakly (Amado, Arce & Fariña, 2015). Every criterion is a marker of truth; a high count nudges toward “experience-based.”

Reality monitoring reads for memory source. Its criteria ask whether the statement looks like a perception or a construction. Seven criteria—clarity, perceptual detail, spatial and temporal information, affect, reconstructability, and realism—are expected to run higher in genuine, perceived memories. The eighth, cognitive operations (explicit references to inferences, reasoning, or thoughts about the event), is expected to run higher in fabricated accounts, because invented stories lean on internal construction (Sporer, 2004; Vrij et al., 2004).

That single structural difference is the most cited practical advantage of reality monitoring. As Vrij and colleagues put it, “unlike CBCA which consists of criteria solely related to truth telling, Reality Monitoring contains both truth telling criteria and a criterion indicative of deception” (Vrij et al., 2004). A method that can point both ways has, in principle, a more balanced decision rule than one that can only accumulate evidence of truth.

What are the reality monitoring criteria?

The reality monitoring approach to deception is most commonly operationalized as about eight criteria, drawn from Sporer’s Judgment of Memory Characteristics Questionnaire (Sporer, 2004):

  1. Clarity/vividness: how clear and sharp the account is.
  2. Perceptual information: sensory detail (sights, sounds, smells, textures, tastes).
  3. Spatial information: where things were, and how they were arranged.
  4. Temporal information: when things happened, and in what order.
  5. Affective information: the feelings experienced during the event.
  6. Reconstructability of the story: whether the event can be mentally pieced back together.
  7. Realism: whether the account is plausible and coherent.
  8. Cognitive operations: references to inferences, reasoning, and thoughts about the event.

Criteria 1 through 7 are hypothesized to be more present in truthful, perceived accounts; cognitive operations is hypothesized to be more present in fabricated ones. But the evidence for individual criteria is uneven. A 2021 meta-analysis of 40 studies found the total RM score separated perceived from imagined memories with a small-to-medium effect (d ≈ 0.54); temporal information was the single most reliable cue, while affective information did not discriminate at all and the cognitive-operations effect was very small (Gancedo et al., 2021).

The lesson mirrors CBCA: the total is more trustworthy than any one criterion, and no single cue is a tell.

How accurate are CBCA and reality monitoring?

Both methods classify truthful and fabricated statements correctly well above chance and well below certainty—and at roughly the same rate. This is the number that should anchor any honest comparison, and it is the one most write-ups either inflate or omit.

For CBCA, Aldert Vrij’s qualitative review of the first 37 studies put correct classification around 70%, meaning it misclassifies roughly 30% of accounts in both directions (Vrij, 2005). For reality monitoring, Masip and colleagues reviewed the empirical evidence and reached a strikingly parallel conclusion.

The approach as a whole appears to discriminate above chance level, reaching accuracy rates that are similar to those of criteria-based content analysis (CBCA). — Masip, Sporer, Garrido & Herrero, 2005

Single studies do sometimes separate the two. In one experiment, 60% of statements were correctly classified on CBCA scores alone versus 74% on reality monitoring scores alone (Vrij et al., 2004). But that gap reflects one sample and one paradigm; across the literature the two methods track each other, and the discipline’s own reviewers treat their accuracy as effectively comparable (Masip et al., 2005). The data points somewhere less comfortable than a winner: two moderate tools, each wrong about a third of the time.

Which is easier to apply, CBCA or reality monitoring?

Reality monitoring is the lighter instrument. With about eight criteria against CBCA’s 19, it takes less training and less time, and researchers consistently report that RM scoring is easier to teach and learn than CBCA scoring (Vrij et al., 2004, citing Sporer, 1997). For a lab building a coding team, that lower training cost is a genuine consideration.

But “easier” is not “reliable,” and reliability is where both methods demand care. Not every criterion is equally codable. In reality monitoring, coders tend to agree well on sensory, spatial, temporal, and affective information, but agreement on clarity, reconstructability, realism, and cognitive operations often needs improvement and benefits from training (Masip et al., 2005).

CBCA shows the same pattern: concrete criteria are easier to agree on than interpretive ones. Whichever method you use, reporting inter-rater reliability with Cohen’s kappa per criterion is the honest way to show which parts of the scheme are solid and which are shaky.

Do CBCA or reality monitoring hold up in court?

Neither method holds up as standalone proof of truthfulness, and the reason is the shared ~30% error rate. This is the legal fulcrum, and it deserves to be stated plainly rather than softened.

For CBCA, Vrij concluded that its error rates are too high for it to serve as the sole indicator of veracity, and that Statement Validity Assessment does not meet the Daubert standard for the admissibility of scientific evidence in the United States (Vrij, 2005). For reality monitoring, Masip and colleagues were more pointed still: despite above-chance discrimination, “the relatively high risk of mis-classifying witnesses’ accounts advises against its use in applied forensic settings” (Masip et al., 2005).

Both are also sensitive to who is speaking and how they were questioned. Verbally skilled or coached witnesses can produce high-scoring accounts; young or less verbal witnesses can produce genuine accounts that score low. And a leading interview can inflate the very detail these criteria reward—which is why the questioning that produces a statement matters as much as the scoring.

The NICHD protocol’s prompt types exist precisely to keep that interview open and non-leading. In CBCA’s home procedure, the Validity Checklist is meant to weigh exactly these confounds before any content score becomes an opinion (Vrij, 2005).

Coding CBCA and RM criteria for research

For research and training—not legal determinations—both methods are a natural fit for span-level annotation: mark where each criterion appears, attach the label and a strength rating, and you get a reviewable, countable record instead of a one-line impression. It also lets you do the honest thing and see the two lenses disagree.

Consider a short synthetic account (invented for illustration; no real testimony):

Witness: We were in the stockroom after closing—it must’ve been near nine, the delivery bay lights were already off. He picked up the blue crate with the cracked handle, and I remember thinking he’d drop it. He kept saying the manager would blame him. I figured he was just nervous. No, wait, he set it down first, then knocked it.

Read through CBCA, this account is rich: unstructured production and a spontaneous correction (“no, wait”), reproduction of speech, a subjective mental state, an unusual superfluous detail (the cracked handle). Read through reality monitoring, the same lines split apart: strong spatial and temporal information (“stockroom,” “near nine”), perceptual detail (the lights, the blue crate)—but “I figured he was just nervous” is a cognitive operation, an inference, which RM treats as a cue that leans the other way. One statement, two scoring logics, and a criterion the two methods read in opposite directions.

One point of reconciliation, so the numbers line up. Tagaroo’s Scale Library currently curates 8 of the 19 CBCA criteria—Logical Structure, Unstructured Production, Quantity of Details, Contextual Embedding, Reproduction of Conversation, Accounts of Subjective Mental State, Spontaneous Corrections, and Admitting Lack of Memory—so the card above shows “8 items” while this article discusses the full 19; the rest are being added as the library expands. Reality monitoring is not a curated scale in the library, because it is a memory-source framework rather than a fixed published instrument. The NICHD interview scale and the CBCA criteria set are both available as guided coding schemes for research and training use, never for legal conclusions.

A caution on data handling belongs here. Real statements in credibility cases—especially those involving children—are among the most sensitive records that exist, governed by strict legal and safeguarding controls. No annotation workflow replaces those controls. The sane default is de-identified, synthetic, or already-lawfully-held text; Tagaroo supports a browser-side anonymous mode so content can stay local—see our privacy policy for specifics.

Common mistakes comparing the two methods

Most errors come from treating either method as more than it is:

  • Calling either one a lie detector. Both describe content; neither decides truth. The score is an input to a cautious opinion, not the opinion (Vrij, 2005; Masip et al., 2005).
  • Declaring a winner on accuracy. Across the literature the two are comparable; a single study favoring one does not generalize (Masip et al., 2005).
  • Summing criteria mechanically. For both methods, total scores are more trustworthy than individual criteria, and rare or interpretive criteria distort a naive sum (Gancedo et al., 2021; Amado, Arce & Fariña, 2015).
  • Ignoring cognitive operations’ direction. In reality monitoring it points toward deception, not truth—reverse it and you invert your own result (Vrij et al., 2004).
  • Forgetting the interview. A suggestive interview inflates the detail both methods reward. Garbage in, high score out.

The practical upshot of CBCA vs reality monitoring: pick the instrument for the question, not the other way round. Reach for CBCA when you want the richer, forensically established 19-criterion read inside Statement Validity Assessment; reach for reality monitoring when you want a lighter, memory-grounded scheme with a built-in deception cue. Just don’t ask either one to do the thing neither can—decide, on its own, whether a person is telling the truth.

References

  • Johnson, M. K., & Raye, C. L. (1981). Reality monitoring. Psychological Review, 88(1), 67–85. doi:10.1037/0033-295X.88.1.67
  • Steller, M., & Köhnken, G. (1989). Criteria-Based Content Analysis. In D. C. Raskin (Ed.), Psychological Methods in Criminal Investigation and Evidence (pp. 217–245). Springer.
  • Undeutsch, U. (1967). Beurteilung der Glaubhaftigkeit von Zeugenaussagen. In Handbuch der Psychologie, Vol. 11: Forensische Psychologie (pp. 26–181). Hogrefe.
  • Vrij, A. (2005). Criteria-Based Content Analysis: A qualitative review of the first 37 studies. Psychology, Public Policy, and Law, 11(1), 3–41. doi:10.1037/1076-8971.11.1.3
  • Masip, J., Sporer, S. L., Garrido, E., & Herrero, C. (2005). The detection of deception with the reality monitoring approach: A review of the empirical evidence. Psychology, Crime & Law, 11(1), 99–122. doi:10.1080/10683160410001726356
  • Sporer, S. L. (2004). Reality monitoring and detection of deception. In P. A. Granhag & L. A. Strömwall (Eds.), The Detection of Deception in Forensic Contexts (pp. 64–102). Cambridge University Press. doi:10.1017/cbo9780511490071.004
  • Vrij, A., Akehurst, L., Soukara, S., & Bull, R. (2004). Let me inform you how to tell a convincing story: CBCA and reality monitoring scores as a function of age, coaching, and deception. Canadian Journal of Behavioural Science, 36(2), 113–126. doi:10.1037/h0087222
  • Gancedo, Y., Fariña, F., Seijo, D., Vilariño, M., & Arce, R. (2021). Reality Monitoring: A Meta-analytical Review for Forensic Practice. European Journal of Psychology Applied to Legal Context, 13(2), 99–110. doi:10.5093/ejpalc2021a10
  • Amado, B. G., Arce, R., & Fariña, F. (2015). Undeutsch hypothesis and Criteria Based Content Analysis: A meta-analytic review. European Journal of Psychology Applied to Legal Context, 7(1), 3–12. doi:10.1016/j.ejpal.2014.11.002

If you code statements for research or interviewer training, Tagaroo turns the CBCA criteria in its Scale Library into a guided, source-anchored annotation workflow—with inter-rater reliability computed as your coders work, for research and training use, never legal determinations.

Frequently asked questions

What is the difference between CBCA and reality monitoring?
Criteria-based content analysis (CBCA) and reality monitoring (RM) are two verbal methods for judging whether a statement reads as experience-based, and they come from different traditions. CBCA grew out of German forensic psychology and rates a statement against 19 content criteria on the Undeutsch hypothesis that memories of real events are richer than fabricated ones (Steller & Köhnken, 1989). RM grew out of cognitive memory science and rates roughly eight memory-characteristic criteria on Johnson and Raye's model that perceived memories differ from imagined ones in sensory, contextual, and cognitive detail (Johnson & Raye, 1981; Sporer, 2004). Both are investigative aids, not lie detectors.
Is reality monitoring more accurate than CBCA?
No method is reliably more accurate than the other. In their review of the empirical evidence, Masip and colleagues concluded that the reality monitoring approach as a whole discriminates above chance, reaching accuracy rates similar to those of CBCA, with error rates around 30% (Masip et al., 2005). Individual studies sometimes favor one method—one experiment classified 60% of statements correctly on CBCA alone versus 74% on reality monitoring alone (Vrij, Akehurst, Soukara & Bull, 2004)—but such single-sample differences do not generalize, and both methods sit well short of the certainty a legal decision requires.
What are the reality monitoring criteria?
The reality monitoring approach to deception is usually operationalized with about eight criteria: clarity/vividness, perceptual (sensory) information, spatial information, temporal information, affective information, reconstructability of the story, realism, and cognitive operations (Sporer, 2004). Seven of these are expected to appear more in truthful, experience-based accounts; cognitive operations—references to inferences and reasoning—is expected to appear more in fabricated accounts, which makes it the one criterion pointing toward deception rather than truth (Vrij et al., 2004).
Can CBCA or reality monitoring be used to prove someone is lying in court?
No. Neither CBCA nor reality monitoring determines truth or detects lies, and neither should be used as standalone proof of veracity. Vrij concluded that CBCA's error rates are too high for it to be the sole indicator of truthfulness and that Statement Validity Assessment does not meet the Daubert standard for scientific evidence in the United States (Vrij, 2005). Masip and colleagues explicitly advised against using reality monitoring in applied forensic settings because of its roughly 30% misclassification rate (Masip et al., 2005).
Which method is easier to apply, CBCA or reality monitoring?
Reality monitoring is generally the simpler of the two to learn and apply. It uses about eight criteria against CBCA's 19, and researchers report that reality monitoring scoring is faster and easier to teach than CBCA scoring (Vrij, Akehurst, Soukara & Bull, 2004, citing Sporer, 1997). Reality monitoring also includes a criterion indicative of deception (cognitive operations), whereas every CBCA criterion points only toward truthfulness—a structural difference that some researchers argue makes reality monitoring's decision rule more balanced.

Put this into practice

Tagaroo turns any rating scale or coding scheme into a guided annotation workflow — with inter-rater reliability computed as you go.