The CAPE-V is the standard auditory-perceptual voice evaluation in English-speaking speech-language pathology: a clinician listens to a voice sample and rates six attributes, giving a shared vocabulary for something that had been described in incompatible in-house terms for decades. This entry follows the 2025 revision (CAPE-Vr) by the original authors, which was published open access — the reason we can host it at all.
This is an ear scale, not a transcript scale. Nothing here can be read off the words. Roughness, breathiness and strain are qualities of the sound; a transcript of a severely dysphonic speaker and a healthy one are identical. Rate from the audio in the player, over a stretch you have actually listened to.
How severity works here. The paper form uses a 100 mm visual-analogue line with printed mildly / moderately / severely deviant markers. Tagaroo stores an ordinal rating, so the line is banded into the four published landmarks — normal, mildly, moderately, severely deviant. If you need the raw millimetre value, record it in the annotation rationale; the band is what supports agreement between raters.
Consistency matters as much as degree. CAPE-V asks whether a deviation is consistent or intermittent. In a timeline tool you express that by where you put the annotation: tag the stretches where the attribute is present rather than the whole file, and intermittency becomes visible in the coverage instead of being flattened into one global number.
Sample the same tasks the protocol specifies — sustained vowels, the six standard sentences, and running speech — and say in the rationale which task the rating came from. Ratings taken from different tasks are not interchangeable.