Consensus Auditory-Perceptual Evaluation of Voice, Revised (CAPE-Vr)

Gail B. Kempster, Kathleen F. Nagle, Nancy Pearl Solomon · 2025

Six perceptual attributes of voice quality — overall severity, roughness, breathiness, strain, pitch and loudness — rated by ear from a recorded voice sample.

The CAPE-V is the standard auditory-perceptual voice evaluation in English-speaking speech-language pathology: a clinician listens to a voice sample and rates six attributes, giving a shared vocabulary for something that had been described in incompatible in-house terms for decades. This entry follows the 2025 revision (CAPE-Vr) by the original authors, which was published open access — the reason we can host it at all.

This is an ear scale, not a transcript scale. Nothing here can be read off the words. Roughness, breathiness and strain are qualities of the sound; a transcript of a severely dysphonic speaker and a healthy one are identical. Rate from the audio in the player, over a stretch you have actually listened to.

How severity works here. The paper form uses a 100 mm visual-analogue line with printed mildly / moderately / severely deviant markers. Tagaroo stores an ordinal rating, so the line is banded into the four published landmarks — normal, mildly, moderately, severely deviant. If you need the raw millimetre value, record it in the annotation rationale; the band is what supports agreement between raters.

Consistency matters as much as degree. CAPE-V asks whether a deviation is consistent or intermittent. In a timeline tool you express that by where you put the annotation: tag the stretches where the attribute is present rather than the whole file, and intermittency becomes visible in the coverage instead of being flattened into one global number.

Sample the same tasks the protocol specifies — sustained vowels, the six standard sentences, and running speech — and say in the rationale which task the rating came from. Ratings taken from different tasks are not interchangeable.

Domains (6)

Overall SeverityOSV

Global impression of how deviant the voice is, taking every attribute together — not the sum of the others but a whole-voice judgement.

Curated skill
RoughnessROU

Perceived irregularity in the voicing source — a rasping, uneven, gravelly quality.

Curated skill
BreathinessBRE

Audible air escaping through the glottis during voicing — a whispery or leaky quality.

Curated skill
StrainSTR

Perceived effort in phonating — the voice sounds squeezed, pressed, or hard-won.

Curated skill
PitchPIT

Whether the perceived pitch deviates from what is expected for this speaker's age and sex; note the direction in the rationale.

Curated skill
LoudnessLOU

Whether the perceived loudness deviates from what the speaking situation calls for; note the direction in the rationale.

Curated skill
Kempster GB, Nagle KF, Solomon NP. The Consensus Auditory-Perceptual Evaluation of Voice, Revised (CAPE-Vr). J Voice. 2025. doi:10.1016/j.jvoice.2024.11.017

Reproduced from the 2025 CAPE-Vr revision, published open access under CC BY 4.0 with the complete form and instructions as appendices — attribute Kempster, Nagle & Solomon 2025. The original 2009 CAPE-V is an ASHA copyright with photocopy-for-clinical-use permission only and is NOT reproduced here; the ordinal bands below are Tagaroo's discretization of the published visual-analogue markers.