Fluency research runs on two kinds of number: a count of disfluent events in a speech sample, and a global scale rating by a listener. This scheme gives you both, built out of the openly published measurement literature rather than a licensed protocol — the widely used SSI-4 is a proprietary Pro-Ed instrument that requires a licensing agreement, so it is not reproduced here.
Count the events by tagging them. Each stuttering-like disfluency gets
its own annotation on the timeline, typed as part-word repetition,
single-syllable word repetition, or dysrhythmic phonation. Percent
stuttered syllables (%SS) and weighted SLD then fall out of the tags rather
than being estimated: weighted SLD is
[(part-word repetitions + single-syllable word repetitions) × mean repetition units] + (2 × dysrhythmic phonations), which is why the severity
rating on repetition codes is the number of repetition units, not an
impression of badness.
Type the typical disfluencies too. Interjections, phrase repetitions and revisions are coded as a separate, presence-only class. This is not busywork: the single most common measurement error in fluency work is sweeping normal disfluency into the stuttered count, and the fix is to give raters somewhere else to put it.
The two global ratings are deliberately kept on their published scales. Stuttering severity is the nine-point scale of O'Brian et al. (2004), and speech naturalness is the nine-point scale of Martin et al. (1984), where 1 is highly natural and 9 highly unnatural. Naturalness exists because fluency-inducing treatments can produce speech that is technically fluent and obviously odd; reporting severity without it hides the trade-off. Both are conventionally averaged over several raters — three for severity, five for naturalness in the classic protocols — so run them as inter-rater-reliability tasks, not single-coder tasks.
A caution about agent assistance. Repetitions and blocks are visible in a time-stamped transcript, but only if transcription preserved them. Deepgram's default output cleans disfluencies away; unless the recording was transcribed with verbatim/filler-word options enabled, the transcript Joey sees has already deleted the phenomenon. Check that before treating agent counts as a starting point, and keep the audio as the authority.