tagaroo

methods

The OPTION Scale: Measuring Shared Decision-Making

The OPTION scale rates how far a clinician involves the patient in a decision. See the behaviors it scores, why real scores are low, and how to code it.

Enrique Gutiérrez10 min readUpdated July 2026
A patient-involvement gauge reading low, illustrating that observed shared decision-making is consistently low.

Everyone in medicine agrees patients should share in decisions about their care. Almost no one measures whether it actually happens—and when someone does, the numbers are humbling. The OPTION scale is the instrument that turns “shared decision-making” from a slogan into an observable, codeable behavior, and its central finding, repeated across dozens of studies, is that clinicians involve patients far less than anyone would like (Elwyn et al., 2003; Couët et al., 2015).

What is the OPTION scale?

The OPTION scale, short for “Observing Patient Involvement,” is a third-party observer measure of how far a clinician involves the patient in shared decision-making during a consultation (Elwyn et al., 2003). A trained rater watches or reads a real encounter and scores the clinician’s behavior against a fixed set of items. It is not a questionnaire the patient or doctor fills in afterward.

That observer basis is the whole point. Self-reports of “we decided together” are unreliable, so Glyn Elwyn and colleagues built an instrument that scores what a clinician actually did in the room—whether they named the decision, laid out the options, and checked what mattered to the patient. Because it codes conduct rather than perception, OPTION measures offered involvement, which is a different quantity from the perceived involvement that patient-reported tools capture.

What does the OPTION scale measure?

The OPTION scale measures a set of specific clinician behaviors that, together, involve a patient in a decision (Elwyn et al., 2005). The canonical revised instrument, OPTION-12, rates twelve such behaviors; Tagaroo’s library groups them into eight annotatable categories for transcript coding, but the underlying scale is the twelve-item one. The table lists the behaviors and what each looks like in a consultation.

BehaviorWhat the clinician does
Problem definitionNames and defines the problem that needs a decision
Options existConveys that more than one reasonable option exists (including doing nothing)
Pros and consExplains the benefits and harms of the options
Patient expectationsExplores the patient's ideas, concerns, and expectations
Preferred involvementAsks how involved the patient wants to be in deciding
Check understandingConfirms the patient has understood the information
Opportunity to askInvites the patient's questions
Defer / reviewOffers to defer the decision and signals it can be reviewed
The clinician behaviors the OPTION scale rates, as grouped in Tagaroo's library (eight categories from the twelve-item OPTION-12; Elwyn et al., 2005). Every item scores the clinician's conduct, not the patient's.

Two features define how it is scored. Every item is rated 0–4 for the extent to which the behavior was performed, where 0 means it was not observed and higher values mean it was carried out more completely (Elwyn et al., 2005). Item scores are then usually summed and rescaled to a 0–100 total, which is the form most studies report.

Why are OPTION scores usually so low?

Observed shared decision-making is consistently, strikingly low—this is the best-established empirical finding about the instrument. A systematic review of 33 studies using the OPTION scale found a mean score of just 23 out of 100 in consultations without any decision-support intervention (Couët et al., 2015).

That number is the reason the scale matters. Even with active interventions to promote involvement, average scores climbed only to around 34 out of 100 (Couët et al., 2015). Shared decision-making, in other words, is widely endorsed and rarely performed—and the gap is invisible until someone codes the behavior.

This is exactly where a measure earns its keep: it converts a comfortable assumption (“of course we involve patients”) into a checkable claim. The same logic runs through the empathic communication coding system, which finds that clinicians often bypass the emotional openings patients offer.

OPTION-12 vs OPTION-5: which version?

There are two current observer versions, and they differ in length and design rather than in what they rate. The revised OPTION-12 (twelve items, each 0–4) is reliable but effortful to apply, and its items performed unevenly, which motivated a shorter tool (Elwyn et al., 2005; Couët et al., 2015).

In 2013, Elwyn and colleagues used a “talk” model of shared decision-making—choice talk, option talk, decision talk—to derive Observer OPTION-5, a five-item version that keeps the 0–4 rating and aligns more tightly with the stages of the decision conversation (Elwyn et al., 2013a; the model is set out in Elwyn et al., 2012). For a large coding project, OPTION-5 is faster; for fine-grained behavioral detail, OPTION-12 captures more. Pick the version to match the question, not the other way round.

How reliable is the OPTION scale?

OPTION reaches acceptable inter-rater reliability when coders are trained, and the revised version improved on the original. The revised OPTION-12 reported an intraclass correlation of 0.77 for the total score (Elwyn et al., 2005), up from 0.62 in the original scale, which also showed internal consistency of Cronbach’s alpha 0.79 (Elwyn et al., 2003).

Reliability is not automatic, though. It depends on trained raters and a clear coding manual, and inter-rater agreement is exactly the kind of thing a coding project should compute rather than assume. If you are calculating it, the Cohen’s kappa and inter-rater reliability guide covers which coefficient fits an ordinal 0–4 scale (a weighted kappa or an intraclass correlation, not plain percent agreement).

Observer-rated vs patient-reported: which shared decision-making measure?

OPTION is observer-rated; the other common instruments are patient-reported, and the distinction changes what you can conclude. A trained coder scoring OPTION captures what the clinician offered; a patient completing CollaboRATE or the SDM-Q-9 reports how involved they felt (Elwyn et al., 2013b; Barr et al., 2014).

MeasureWho rates itWhat it captures
OPTION-12 / OPTION-5Trained observer (from recording/transcript)Clinician behaviors that offer involvement
CollaboRATEPatient (after the visit)Patient's perception of being involved
SDM-Q-9Patient (after the visit)Patient-perceived shared decision-making process
Observer-rated versus patient-reported shared decision-making measures. OPTION indexes offered involvement; CollaboRATE and SDM-Q-9 index perceived involvement (Elwyn et al., 2013b; Barr et al., 2014).

The two families often disagree, and neither is “wrong.” A clinician can perform every OPTION behavior while a patient still feels unheard, or vice versa. Blending an observer score with a patient-report score hides that tension, so keep them separate and report which one you used.

How do you code shared decision-making from a transcript?

To code the OPTION scale, mark the passage where the clinician performs each behavior and record how fully it was done—this is evidence-mode annotation, where you find the behavior anywhere in the encounter rather than tagging every turn. A behavior is scored once, on the strongest evidence for it in the consultation.

Consider a short synthetic consultation:

Clinician: Your knee arthritis has reached the point where we should decide on next steps. There are really three reasonable paths here—keep managing it with physiotherapy and painkillers, try a steroid injection, or look at a knee replacement. Each has real trade-offs. Before I go through those, what matters most to you about how we handle this?

That single turn shows three OPTION behaviors: problem definition (“reached the point where we should decide”), options exist (“three reasonable paths”), and the start of exploring patient expectations (“what matters most to you”). Marking the passage that evidences each behavior keeps the score auditable, so a second coder can check the call against the clinician’s actual words. That is what makes the OPTION coding scheme reproducible enough to compute agreement across raters, and it pairs naturally with coding empathy through the ECCS in the same consultation.

On data handling: consultation transcripts are sensitive clinical records, so the sane default is de-identified text and a privacy-first setup. Tagaroo supports a browser-side anonymous mode, so transcript content can stay local rather than being uploaded—worth checking against your ethics approval and information-governance rules before any real consultation data touches a tool. Coding the interviewer’s conduct rather than the patient’s is a pattern OPTION shares with the NICHD forensic interview protocol.

Common mistakes when coding OPTION

The recurring errors come from misreading what the scale is for:

  • Scoring intent instead of behavior. OPTION rates observable conduct, not whether the clinician “seemed collaborative.” If the behavior isn’t in the transcript, it scores 0.
  • Treating a high score as good care. OPTION measures involvement in the decision process, not whether the decision was right or the patient satisfied. It is one dimension, not a verdict.
  • Blending observer and patient-report scores. OPTION and CollaboRATE measure different things; reporting a single “SDM score” from both is a category error (Barr et al., 2014).
  • Forgetting the instrument is 12 items. Groupings and short forms exist, but cite the canonical OPTION-12 (or OPTION-5) so your method is reproducible.

The practical upshot: the OPTION scale matters because it makes a universally endorsed ideal—involving patients in their own care—into something you can actually measure, and the measurement keeps coming back low. Code the clinician’s behaviors from the transcript, keep observer and patient-report scores apart, and let the 0–4 evidence, not an impression, carry the score.

References

  • Elwyn, G., Edwards, A., Wensing, M., Hood, K., Atwell, C., & Grol, R. (2003). Shared decision making: developing the OPTION scale for measuring patient involvement. Quality and Safety in Health Care, 12(2), 93–99. doi:10.1136/qhc.12.2.93
  • Elwyn, G., Hutchings, H., Edwards, A., Rapport, F., Wensing, M., Cheung, W.-Y., & Grol, R. (2005). The OPTION scale: measuring the extent that clinicians involve patients in decision-making tasks. Health Expectations, 8(1), 34–42. doi:10.1111/j.1369-7625.2004.00311.x
  • Elwyn, G., Tsulukidze, M., Edwards, A., Légaré, F., & Newcombe, R. (2013a). Using a ‘talk’ model of shared decision making to propose an observation-based measure: Observer OPTION-5. Patient Education and Counseling, 93(2), 265–271. doi:10.1016/j.pec.2013.08.005
  • Elwyn, G., Barr, P. J., Grande, S. W., Thompson, R., Walsh, T., & Ozanne, E. M. (2013b). Developing CollaboRATE: a fast and frugal patient-reported measure of shared decision making in clinical encounters. Patient Education and Counseling, 93(1), 102–107. doi:10.1016/j.pec.2013.05.009
  • Elwyn, G., Frosch, D., Thomson, R., Joseph-Williams, N., et al. (2012). Shared decision making: a model for clinical practice. Journal of General Internal Medicine, 27(10), 1361–1367. doi:10.1007/s11606-012-2077-6
  • Couët, N., Desroches, S., Robitaille, H., Vaillancourt, H., Leblanc, A., Turcotte, S., Elwyn, G., & Légaré, F. (2015). Assessments of the extent to which health-care providers involve patients in decision making: a systematic review of studies using the OPTION instrument. Health Expectations, 18(4), 542–561. doi:10.1111/hex.12054
  • Barr, P. J., Thompson, R., Walsh, T., Grande, S. W., Ozanne, E. M., & Elwyn, G. (2014). The psychometric properties of CollaboRATE: a fast and frugal patient-reported measure of the shared decision-making process. Journal of Medical Internet Research, 16(1), e2. doi:10.2196/jmir.3085

If you code shared decision-making from consultation transcripts, Tagaroo turns the OPTION scale into a guided, evidence-anchored annotation workflow—with inter-rater reliability computed as your coders work.

Frequently asked questions

What is the OPTION scale?
The OPTION scale ('Observing Patient Involvement') is a third-party observer measure of how far a clinician involves the patient in shared decision-making during a consultation (Elwyn et al., 2003). A trained rater scores a recording or transcript on a set of clinician behaviors—such as defining the problem, saying that options exist, and describing their pros and cons—rather than relying on patient or clinician self-report.
How is the OPTION scale scored?
In the revised OPTION-12, each clinician behavior is rated 0–4 for the extent to which it was performed, where 0 means the behavior was not observed and higher values mean it was carried out more completely (Elwyn et al., 2005). Item scores are usually summed and rescaled to a 0–100 total. A shorter five-item version, Observer OPTION-5, uses the same 0–4 rating (Elwyn et al., 2013a).
Why are OPTION scores usually low?
Observed shared decision-making is consistently low in routine care. A systematic review of studies using the OPTION instrument found a mean score of 23 out of 100 in consultations without a decision-support intervention, rising only to about 34 with one (Couët et al., 2015). The robust, repeated finding is that clinicians involve patients far less than the ideal, even when trained.
What is the difference between OPTION and CollaboRATE?
OPTION is observer-rated: a trained coder scores the clinician's behavior from a recording or transcript. CollaboRATE and SDM-Q-9 are patient-reported: the patient rates how involved they felt (Elwyn et al., 2013b; Barr et al., 2014). They measure related but distinct things—offered involvement versus perceived involvement—and their scores should not be blended.
How many items does the OPTION scale have?
The canonical revised instrument, OPTION-12, has twelve items, each rated 0–4 (Elwyn et al., 2005). A later five-item observer measure, OPTION-5, was derived from a 'talk' model of shared decision-making (Elwyn et al., 2013a). Tagaroo's library groups the twelve behaviors into eight annotatable categories for transcript coding; the underlying instrument remains the 12-item scale.

Put this into practice

Tagaroo turns any rating scale or coding scheme into a guided annotation workflow — with inter-rater reliability computed as you go.