therapy process
MITI Motivational Interviewing: A Fidelity Coding Guide
How MITI motivational interviewing coding works: behavior counts, four global scores, and the R:Q ratio and %complex reflections that grade MI fidelity.

A clinician can be certain they “did motivational interviewing” and be wrong. What actually happened in the room is a specific pattern of moves: how often they reflected versus asked, whether they argued for change or evoked it, whether they honored the client’s autonomy or steered. That gap between intention and behavior is exactly what the MITI motivational interviewing fidelity code was built to measure, by turning a session into countable, checkable data (Moyers, Manuel & Ernst, 2014).
What does the MITI motivational interviewing code measure?
The MITI 4.2.1 is a one-pass behavioral coding system that measures how faithfully a clinician delivers motivational interviewing, scored from a single random 20-minute segment of a session (Moyers, Manuel & Ernst, 2014). It was derived from the longer Motivational Interviewing Skill Code and is now the most widely used MI fidelity measure (Moyers et al., 2016). MI itself produces a small but reliable benefit—about g = 0.28 against weak or no-treatment comparisons in a meta-analysis of 25 years of studies (Lundahl et al., 2010)—which is exactly why delivering it faithfully, and being able to verify that you did, is worth the coding effort.
Two things make the MITI distinctive, and both matter for how you read a score. First, it has two components: holistic global scores and discrete behavior counts. Second, it codes the clinician only.
The MITI deliberately leaves out client change talk and sustain talk, which are captured by the fuller MISC change-talk codes and the SCOPE (Moyers et al., 2016). If your question is about what the client said, the MITI is the wrong instrument.
A designated change goal is set before coding begins. The MITI assumes the conversation has a specific target behavior (say, reducing drinking), and coders are told what it is so they can judge whether the clinician is evoking and responding to change about that goal (Moyers, Manuel & Ernst, 2014). Sessions with no change target are a poor fit for the tool.
What are the four MITI global scores?
The four global scores are each a single 1–5 judgment of the whole segment, capturing the coder’s overall impression rather than a tally (Moyers, Manuel & Ernst, 2014). The coder starts at a default of 3 and moves up or down; a 5 is withheld when there are prominent examples of poor practice. The four globals fall into two families.
| Global score | Component | What the 1–5 rating captures |
|---|---|---|
| Cultivating Change Talk | Technical | How actively the clinician evokes and responds to the client's own arguments for change |
| Softening Sustain Talk | Technical | How well the clinician avoids reinforcing the client's arguments for the status quo |
| Partnership | Relational | How much the clinician works with the client as an equal, rather than as the expert-in-charge |
| Empathy | Relational | How much the clinician understands and conveys the client's perspective |
The two families roll up into two summary numbers: the Technical Global is the average of Cultivating Change Talk and Softening Sustain Talk, and the Relational Global is the average of Partnership and Empathy (Moyers, Manuel & Ernst, 2014). This split is not cosmetic. It maps onto the two competing accounts of why MI works—a technical hypothesis (skillful evoking of change talk drives outcomes) and a relational hypothesis (the empathic, collaborative stance does). A meta-analysis of MI’s causal chain found that clinician MI-consistent skills correlate with more client change talk (r = .26), while client sustain talk predicts worse outcomes (r = −.24)—so how the clinician behaves measurably shapes what the client says (Magill et al., 2014).
The MITI behavior counts
The behavior counts are running tallies of specific clinician utterances, assigned by decision rules rather than overall impression (Moyers, Manuel & Ernst, 2014). MITI 4.2.1 defines 10 behavior counts. They divide into codes that feed the technical ratios (questions and reflections), the MI-adherent group, the MI non-adherent group, and two neutral codes.
| Behavior count | Group | What the clinician does |
|---|---|---|
| Giving Information | Neutral | Educates, gives feedback, or offers an opinion without persuading |
| Persuade with Permission | Neutral | Advises or argues for change after asking permission or emphasizing autonomy |
| Question | Technical (feeds R:Q) | Asks the client a question, open or closed |
| Simple Reflection | Technical (feeds R:Q, %CR) | Reflects back the client's meaning, adding little or nothing |
| Complex Reflection | Technical (feeds R:Q, %CR) | Reflects with added meaning, emphasis, or direction |
| Affirm | MI-adherent | Accentuates the client's strengths, efforts, or worth |
| Seeking Collaboration | MI-adherent | Shares power; asks permission; negotiates the agenda |
| Emphasizing Autonomy | MI-adherent | Highlights the client's control and freedom of choice |
| Persuade | MI non-adherent | Argues for change with logic or facts, without emphasizing autonomy |
| Confront | MI non-adherent | Disagrees, warns, moralizes, shames, or corrects |
A parsing rule keeps the tallies honest. Reflections are coded once per volley—if any reflection in a clinician’s turn is complex, the whole volley scores a single Complex Reflection; otherwise a single Simple Reflection—and only one Question is coded per volley (Moyers, Manuel & Ernst, 2014). This stops a talkative clinician from inflating their own counts, and it is one reason MITI reflection and question counts code reliably.
Reflection-to-question ratio, %complex reflections, and MI-adherence
The counts become useful once they are turned into summary scores, because raw frequencies depend on how much the clinician talked (Moyers, Manuel & Ernst, 2014). Four summary scores carry most of the interpretive weight, and the manual attaches suggested thresholds to some of them.
| Summary score | Formula | Fair | Good |
|---|---|---|---|
| Relational Global | (Partnership + Empathy) ÷ 2 | 3.5 | 4 |
| Technical Global | (Cultivating CT + Softening ST) ÷ 2 | 3 | 4 |
| % Complex Reflections | CR ÷ (SR + CR) | 40% | 50% |
| Reflection-to-Question (R:Q) | Total reflections ÷ total questions | 1:1 | 2:1 |
| Total MI-Adherent | Seeking Collaboration + Affirm + Emphasizing Autonomy | — | — |
| Total MI Non-Adherent | Confront + Persuade | — | — |
The reflection-to-question ratio is the headline number: reflections divided by questions, with 1:1 marking fair practice and 2:1 marking good practice (Moyers, Manuel & Ernst, 2014). Percent complex reflections asks how many of those reflections did real work—complex reflections over all reflections—with 40% fair and 50% good.
Two cautions before anyone treats these as targets. First, the manual is explicit that the thresholds rest on expert opinion and lack normative or validity data, so they belong alongside other evidence, not as a pass/fail gate (Moyers, Manuel & Ernst, 2014). Second, gaming a ratio is easy and meaningless—a clinician can hit 2:1 with shallow reflections while the globals stay low. The numbers are a lens, not a verdict.
MI-consistent vs MI-inconsistent: the behavior split
The clearest signal in the behavior counts is the split between behaviors that support the MI spirit and behaviors that work against it (Moyers, Manuel & Ernst, 2014). MITI 4.2.1 sums Total MI-Adherent from Seeking Collaboration, Affirm, and Emphasizing Autonomy, and Total MI Non-Adherent from Confront and Persuade. Giving Information and Persuade with Permission sit outside both as neutral.
One thing MITI 4 deliberately does not give you is a percentage. MITI 3.1.1 reported a percent MI-consistent score; the MITI 4 authors judged it uninformative and dropped it, so 4.2.1 reports the raw Total MI-Adherent and Total MI Non-Adherent counts and intentionally leaves their thresholds unset for lack of data (Moyers et al., 2016).
The terminology trips people up, so it is worth pinning down. Earlier versions (MITI 3.1.1) called these MI-consistent and MI-inconsistent behaviors; MITI 4 renamed them MI-adherent and MI non-adherent, and made Persuade with Permission a separate, non-penalized code (Moyers et al., 2016). If you read “MI-consistent behaviors” in an older paper and “MI-adherent” in a newer one, they are pointing at the same idea from different manual versions—so cite the version you actually used.
Why the Persuade-with-Permission carve-out matters: giving advice is not inherently against MI. Arguing for change without permission (Persuade) counts against fidelity; asking first, or framing the advice around the client’s own choice (Persuade with Permission), does not. That single distinction is where a lot of real clinical practice lives, and collapsing it—as any 9-category simplification does—loses a genuinely important line.
How reliable is MITI coding?
MITI motivational interviewing coding is reliable enough for research when raters are trained and the sample is adequate, but reliability is uneven across codes (Moyers et al., 2016). In the MITI 4 validation, four undergraduate, non-professional raters coded 50 audiotaped MI sessions, and inter-rater reliability was estimated with intraclass correlations.
Using all four raters, items landed in the good-to-excellent range—Questions at ICC = .97, Simple Reflection .93, Complex Reflection .91, and the R:Q ratio .93 (Moyers et al., 2016). Two codes were the weak spots: Emphasizing Autonomy and percent complex reflections. Those are the ones to watch, and the reason is instructive.
Low-frequency behaviors are fragile. When Moyers and colleagues re-estimated reliability from just two coders on a 20% subsample, Emphasizing Autonomy collapsed to ICC = .06 and Persuade with Permission to .22—not because the code is bad, but because rare events give agreement statistics almost nothing to work with (Moyers et al., 2016). The practical lesson for any coding project: budget enough sessions and enough double-coding, especially for the sparse codes, and establish reliability rather than assuming it. The same “code the exact utterance so a second rater can check it” discipline underlies related consultation-coding schemes such as the Empathic Communication Coding System, the OPTION shared decision-making scale, and the NICHD interview-prompt coding, where question type is central just as it is to the MITI’s R:Q ratio.
How do you code the MITI from a transcript?
To code the MITI, you tag each clinician utterance with its behavior count and rate the four globals across the segment—instance-level annotation for the counts, a holistic judgment for the globals (Moyers, Manuel & Ernst, 2014). The counts live in the exact words; the globals live in the whole.
Consider a short synthetic exchange, with a change goal of reducing drinking:
Client: I know the drinking is a problem, but honestly a couple of beers is the only way I switch off after work.
Clinician: So it feels like the one reliable way to unwind right now—and part of you already sees the cost. [Complex Reflection]
Clinician: Would it be okay if I shared what tends to help with the winding-down part? [Seeking Collaboration]
The first clinician turn reflects both sides of the client’s ambivalence and adds meaning, so it scores a Complex Reflection; the second asks permission before offering information, an MI-adherent Seeking Collaboration rather than an unbidden Persuade. Tag the utterance, and the count is anchored to text a second coder can check—which is what makes the MITI behavior counts reproducible enough to compute agreement across raters.
This is where hand-coding hurts. A single 20-minute segment can hold dozens of volleys, each needing a parsing decision and a code, and a study needs many segments double-coded.
Tagaroo’s value here is a first pass: the agent tags candidate reflections, questions, affirmations, and MI non-adherent moves across the transcript, so your coders start from a draft and spend their time on the hard judgment calls—simple versus complex reflection, Persuade versus Persuade with Permission—instead of the mechanical tally. The human still decides; the machine removes the drudgery. For the wider set of options here, see our survey of clinical transcript annotation tools.
On data handling: MI session transcripts are sensitive clinical records, so the sane default is de-identified text and a privacy-first setup. Tagaroo supports a browser-side anonymous mode, so transcript content can stay local rather than being uploaded—worth checking against your ethics approval and information-governance rules before any real session data touches a tool.
Common mistakes when coding the MITI
The recurring errors come from treating the MITI as more than it claims to be, or from ignoring its parsing rules:
- Reading the MITI as a measure of the client. It codes the clinician; client change talk belongs to the MISC change-talk codes, not the MITI (Moyers et al., 2016).
- Treating the thresholds as validated cutoffs. R:Q 2:1 and %CR 50% are expert opinion, not empirically set pass marks (Moyers, Manuel & Ernst, 2014).
- Over-parsing reflections. Only one reflection code per volley, and any complex reflection makes the volley Complex—counting each phrase separately breaks reliability (Moyers, Manuel & Ernst, 2014).
- Confusing Persuade with Persuade with Permission. Advising with permission is neutral; arguing without it is MI non-adherent. Collapsing the two mislabels ordinary good practice as a fidelity failure.
- Trusting sparse codes on small samples. Emphasizing Autonomy and other rare behaviors need enough sessions and double-coding to reach dependable agreement (Moyers et al., 2016).
The practical upshot: MITI motivational interviewing coding works because it converts “was that MI?” into countable moves and holistic ratings that two trained people can agree on—reflections over questions, complex over simple, adherent over non-adherent. Code the utterances, respect the parsing rules, and read the ratios as a lens rather than a scoreboard.
References
- Moyers, T. B., Manuel, J. K., & Ernst, D. (2014). Motivational Interviewing Treatment Integrity Coding Manual 4.2.1. University of New Mexico, Center on Alcoholism, Substance Abuse, and Addictions (CASAA).
- Moyers, T. B., Rowell, L. N., Manuel, J. K., Ernst, D., & Houck, J. M. (2016). The Motivational Interviewing Treatment Integrity Code (MITI 4): Rationale, preliminary reliability and validity. Journal of Substance Abuse Treatment, 65, 36–42. doi:10.1016/j.jsat.2016.01.001
- Miller, W. R., & Rollnick, S. (2013). Motivational Interviewing: Helping People Change (3rd ed.). New York: Guilford Press.
- Magill, M., Gaume, J., Apodaca, T. R., Walthers, J., Mastroleo, N. R., Borsari, B., & Longabaugh, R. (2014). The technical hypothesis of motivational interviewing: A meta-analysis of MI’s key causal model. Journal of Consulting and Clinical Psychology, 82(6), 973–983. doi:10.1037/a0036833
- Lundahl, B. W., Kunz, C., Brownell, C., Tollefson, D., & Burke, B. L. (2010). A meta-analysis of motivational interviewing: Twenty-five years of empirical studies. Research on Social Work Practice, 20(2), 137–160. doi:10.1177/1049731509347850
- Miller, W. R., Moyers, T. B., Ernst, D., & Amrhein, P. (2008). Manual for the Motivational Interviewing Skill Code (MISC), Version 2.x. University of New Mexico, CASAA.
If you code MI sessions for fidelity, Tagaroo turns the MITI behavior counts into a guided, evidence-anchored annotation workflow—with a first-pass draft to review and inter-rater reliability computed as your coders work.
Frequently asked questions
- What is the MITI in motivational interviewing?
- The MITI (Motivational Interviewing Treatment Integrity code, version 4.2.1) is a behavioral coding system for rating how faithfully a clinician delivers motivational interviewing (Moyers, Manuel & Ernst, 2014). A coder listens to a random 20-minute segment of a session and produces two kinds of data: four global scores (Cultivating Change Talk, Softening Sustain Talk, Partnership, Empathy), each rated 1–5, and behavior counts that tally specific clinician utterances such as reflections, questions, and affirmations. It codes the clinician only, not the client.
- What is the difference between MITI global scores and behavior counts?
- Global scores are holistic 1–5 judgments of an entire segment—the coder's overall impression of Cultivating Change Talk, Softening Sustain Talk, Partnership, and Empathy—while behavior counts are running tallies of discrete utterances (questions, simple and complex reflections, affirmations, and so on) made without judging overall quality (Moyers, Manuel & Ernst, 2014). Globals capture the gestalt; behavior counts capture the countable mechanics, and the two are combined into summary scores like the reflection-to-question ratio.
- What is a good reflection-to-question ratio in MI?
- The MITI 4.2.1 manual suggests a reflection-to-question (R:Q) ratio of 1:1 for 'fair' practice and 2:1 for 'good' practice, and a percentage of complex reflections of 40% (fair) and 50% (good) (Moyers, Manuel & Ernst, 2014). These thresholds are explicitly based on expert opinion and lack normative validity data, so the manual advises using them alongside other evidence rather than as pass/fail cutoffs.
- What are MI-adherent and MI-inconsistent behaviors?
- In MITI 4.2.1, MI-adherent behaviors are Seeking Collaboration, Affirm, and Emphasizing Autonomy (summed as Total MI-Adherent), and MI non-adherent behaviors are Confront and Persuade (summed as Total MI Non-Adherent) (Moyers, Manuel & Ernst, 2014). Earlier versions (MITI 3.1.1) called these MI-consistent and MI-inconsistent; MITI 4 renamed them MI-adherent and MI non-adherent and split off Persuade with Permission and Giving Information as neutral.
- Does the MITI measure the client or the clinician?
- The MITI codes the clinician's behavior only. It deliberately excludes client change talk and sustain talk, which are process variables captured by the more labor-intensive Motivational Interviewing Skill Code (MISC) and the SCOPE (Moyers et al., 2016). If you need the client side of the conversation—the balance of change talk versus sustain talk—you code with the MISC, not the MITI.
Put this into practice
Tagaroo turns any rating scale or coding scheme into a guided annotation workflow — with inter-rater reliability computed as you go.