Vocal Affect & Prosody (Audio Coding)

James A. Russell, Klaus R. Scherer, Roddy Cowie · 1980

Dimensional coding of the emotion carried by the voice — arousal, valence, emphatic stress, affect bursts, and vocal-verbal incongruence.

The voice carries emotion the words do not. "I'm fine" is a complete sentence in a transcript and an unresolved question in a recording. This is the audio sibling of the schemes Tagaroo already hosts for the other modalities — Ekman and Plutchik for emotion words in text, Facial Affect for still frames — and it deliberately uses the dimensional tradition rather than a basic-emotion category list, because that is what vocal-affect research overwhelmingly uses. Listeners agree well on how activated and how positive a voice sounds; they agree far less on whether it is contempt or disgust.

Arousal and valence are rated as dimensions, not severities. The slider runs from very low to very high in both cases, and the midpoint is neutral, not "mild". A flat, exhausted voice is a 1 on arousal — that is a strong, meaningful code, not an absence of one. This differs from every clinical severity scale in the library, so say so in your codebook and make sure your raters have internalized it before you compute agreement.

Prosodic incongruence is the clinically interesting code. Vocal affect contradicting verbal content — bad news delivered brightly, distress reported in a level tone, agreement voiced with audible reluctance — is a staple of psychotherapy-process research and is invisible to any transcript-based method. It is also the code where inter-rater agreement is hardest to establish, so treat a high-agreement incongruence dataset as a real achievement and pilot it before scaling.

Rate short stretches, not whole recordings. Continuous-trace methods (Cowie's FEELTRACE and its descendants) exist because affect moves on a scale of seconds. A single arousal rating for a fifty-minute session is close to meaningless. Tag the segment you actually listened to and rate that.

Ear only, by construction. These codes are about how something sounded. The curated skills instruct the AI agent to abstain rather than infer emotion from word choice, because inferring vocal affect from a transcript is exactly the error the scheme exists to prevent — and doing it silently would produce confident, unfalsifiable, wrong annotations.

Domains (5)

Vocal ArousalARO

How activated the voice sounds, from flat and de-energized to highly aroused. A dimension: low is a code, not an absence.

Curated skill
Vocal ValenceVAL

How positive or negative the voice sounds, independent of what the words say. A dimension, midpoint neutral.

Curated skill
Emphatic StressEMP

Prosodic marking of focus or importance — pitch, loudness or duration used to make one element stand out.

Curated skill
Affect BurstAFB

A nonverbal vocal emotion event: laughter, sigh, sob, gasp, groan, tut. Coded as an event, not a rating.

Curated skill
Prosodic IncongruencePRI

Vocal affect that contradicts the verbal content — the mismatch itself is the code.

Curated skill
Russell JA. A circumplex model of affect. J Pers Soc Psychol. 1980;39(6):1161-1178. Banse R, Scherer KR. Acoustic profiles in vocal emotion expression. J Pers Soc Psychol. 1996;70(3):614-636. Cowie R, Douglas-Cowie E, Savvidou S, McMahon E, Sawey M, Schröder M. FEELTRACE: an instrument for recording perceived emotion in real time. Proc ISCA Workshop on Speech and Emotion. 2000:19-24. doi:10.1037/h0077714

Conceptual framework from the published literature — free to operationalize with citation. Dimension definitions and anchor wording are Tagaroo's own operational phrasing for time-range audio coding. No licensed instrument is reproduced: this scheme does not include the Geneva Emotion Wheel (which requires a licence for commercial use) or the Self-Assessment Manikin artwork (permission-per-use).