The voice carries emotion the words do not. "I'm fine" is a complete sentence in a transcript and an unresolved question in a recording. This is the audio sibling of the schemes Tagaroo already hosts for the other modalities — Ekman and Plutchik for emotion words in text, Facial Affect for still frames — and it deliberately uses the dimensional tradition rather than a basic-emotion category list, because that is what vocal-affect research overwhelmingly uses. Listeners agree well on how activated and how positive a voice sounds; they agree far less on whether it is contempt or disgust.
Arousal and valence are rated as dimensions, not severities. The slider runs from very low to very high in both cases, and the midpoint is neutral, not "mild". A flat, exhausted voice is a 1 on arousal — that is a strong, meaningful code, not an absence of one. This differs from every clinical severity scale in the library, so say so in your codebook and make sure your raters have internalized it before you compute agreement.
Prosodic incongruence is the clinically interesting code. Vocal affect contradicting verbal content — bad news delivered brightly, distress reported in a level tone, agreement voiced with audible reluctance — is a staple of psychotherapy-process research and is invisible to any transcript-based method. It is also the code where inter-rater agreement is hardest to establish, so treat a high-agreement incongruence dataset as a real achievement and pilot it before scaling.
Rate short stretches, not whole recordings. Continuous-trace methods (Cowie's FEELTRACE and its descendants) exist because affect moves on a scale of seconds. A single arousal rating for a fifty-minute session is close to meaningless. Tag the segment you actually listened to and rate that.
Ear only, by construction. These codes are about how something sounded. The curated skills instruct the AI agent to abstain rather than infer emotion from word choice, because inferring vocal affect from a transcript is exactly the error the scheme exists to prevent — and doing it silently would produce confident, unfalsifiable, wrong annotations.