The BFCRS is the instrument catatonia research runs on, and almost every sign in it is something you have to see: posturing, staring, grimacing, stereotypy, mannerisms, waxy flexibility. Its authors' own training resource pairs every item with an example video, so a video annotation workflow is not an adaptation of this scale — it is how the field already teaches it.
Twenty-three items, two rating patterns. Seventeen are graded 0 to 3 with item-specific anchors. Six are binary, scored 0 or 3 only: waxy flexibility, mitgehen, gegenhalten, ambitendency, grasp reflex and perseveration. Their presence is what matters; there is no meaningful "mild gegenhalten". Note that combativeness and autonomic abnormality are graded 0 to 3 despite often being listed among the binary items in secondary sources.
The screening instrument is the first fourteen items (BFCSI): excitement, immobility/stupor, mutism, staring, posturing/catalepsy, grimacing, echopraxia/echolalia, stereotypy, mannerisms, verbigeration, rigidity, negativism, waxy flexibility and withdrawal. Two or more present is the commonly used screening threshold — cite that to the reliability literature, not to the manual, and note it is a screen rather than a diagnosis.
This scale requires the examination, not just an interview. Mitgehen, gegenhalten, grasp reflex, rigidity and waxy flexibility are elicited by the examiner's hands; autonomic abnormality needs vital signs. They are rateable from a recording of the standardized examination and not from a recording of a conversation. If your footage does not contain the manoeuvre, do not infer the sign — the manual's own rule is that an item you are unsure of is rated 0.
Declare which window you rated. The published form distinguishes a state examination from an interval examination over a stated number of hours, with a floor of about five minutes or the duration of the assessment. On a timeline this becomes the span you annotate, so tag the assessment window and say which kind it is; a state and an interval rating of the same patient are not the same measurement.
Some items are audible, not visual. Mutism, verbigeration, echolalia and stereotyped speech show up in a transcript with timings, which is what makes them the few items an AI agent can genuinely help with here. The motor signs are for your eyes only, and the curated skills say so.