methods
Rating Scale to Codebook: Operationalize Any Instrument
A step-by-step method to go from rating scale to codebook: turn each item into a code, anchors, examples, and decision rules coders apply consistently.

A published rating scale and an annotation codebook look alike—both are lists of named things with definitions—but they do different jobs, and the gap between them is exactly where inter-coder disagreement lives. A rating scale tells a trained clinician how severe a symptom is. A codebook tells any coder, human or agent, which span of text counts as that symptom in the first place. The rating-scale-to-codebook conversion is the work of making that second decision explicit: operationalizing each item into a code, an operational definition, positive and negative examples, and a decision rule a coder can apply the same way twice.
What a codebook adds that a rating scale does not
A codebook adds the operational layer a rating scale leaves implicit: it names each construct as a short code and then spells out, in terms of observable text, when that code applies and when it does not. A rating scale item like Reported Sadness comes with a severity anchor and a clinician’s trained intuition about what counts. A codebook cannot rely on that intuition, because its whole purpose is to make two different coders converge—so it has to write the intuition down.
That distinction is not academic. In team-based coding, MacQueen and colleagues at the CDC found that a structured codebook—each code carrying a mnemonic, a short and full definition, guidance on when to use it, guidance on when not to, and examples—was the practical instrument for improving agreement among coders working at dispersed sites (MacQueen et al., 1998). The scale gives you the first two fields for free. The codebook is the other four.
The annotation-for-machine-learning tradition frames the same move differently but arrives at the same place. Pustejovsky and Stubbs describe building an annotation task as turning your abstract model of a phenomenon into a concrete specification, and they treat “how do you take someone else’s standard and use it to create a specification” as a first-class question (Pustejovsky & Stubbs, 2012). A rating scale is exactly that someone-else’s-standard. Operationalizing it is writing the spec.
| Field | The rating-scale item gives you | The codebook entry must add |
|---|---|---|
| Name / label | An item name (e.g. "Reported Sadness") | A short code + mnemonic a coder types (RSD) |
| Definition | A construct description written for a clinician | An operational definition tied to observable text |
| When it applies | Implicit in rater training | An explicit inclusion rule: what the span must show |
| When it does not | Usually unstated | An explicit exclusion rule + the neighbor code that fits instead |
| Examples | Rarely provided | Positive and near-miss examples (synthetic, for boundaries) |
| Severity | Anchored 0–6 by the manual | A mapping from span features to each anchor band |
| Unit of analysis | The whole interview → one score | The span, utterance, or turn a code attaches to |
The rating-scale-to-codebook procedure in seven steps
The rating-scale-to-codebook procedure is seven steps, and each one has a home in the methods literature. Follow them in order; the output is a coding manual from the rating scale you started with, plus the examples and boundary rules the scale never wrote down.
- Fix the scope and read the manual. Decide which construct and which items you are operationalizing, then read the instrument’s own glossary first—the item definitions and severity anchors are your raw material (Montgomery & Åsberg, 1979; Andreasen, 1986). Note the rater type and the unit the scale rates: most clinical scales rate a whole interview, and you are about to move that decision down to the span.
- Extract each item as a candidate code. Every scale item becomes one theory-driven code—a code taken deductively from an existing framework rather than induced from the data (DeCuir-Gunby et al., 2011). Give each a short label and a mnemonic you can type while coding.
- Write an operational definition per code. Convert the clinician-facing description into a rule about text: what a span must express to earn the code. This is the step where the model becomes a specification (Pustejovsky & Stubbs, 2012). Keep it to one sentence a new coder can hold in their head.
- Add inclusion, exclusion, and examples. For each code, state when to apply it, when not to, and give at least one positive example and one near-miss. This four-part structure—definition, inclusion, exclusion, example—is the part MacQueen and colleagues tie directly to higher intercoder agreement (MacQueen et al., 1998).
- Map severity anchors to observable features. A scale’s 0–6 anchors become the codebook’s severity guide: for each band, write what a span has to show to land there. Where two codes are neighbors, add a tie-break rule—the TLC manual does this explicitly for pairs like circumstantiality versus loss of goal (Andreasen, 1986).
- Pilot on a sample and measure agreement. Have two coders independently apply the draft codebook to the same synthetic transcripts, then compute an agreement coefficient: Cohen’s kappa for two coders on nominal codes, Krippendorff’s alpha for more coders, ordinal severity, or missing data (Landis & Koch, 1977; Hayes & Krippendorff, 2007).
- Revise and version. Treat the first codebook as a draft. The annotation literature builds this in as the MAMA loop—model, annotate, model, annotate—iterating the guidelines until agreement stabilizes (Pustejovsky & Stubbs, 2012). Bump a version number on every pass so a coded dataset can name the codebook that produced it.
Steps 1 through 5 are authoring; steps 6 and 7 are where the codebook earns trust. A codebook that has never been piloted is a hypothesis, not a coding manual.
A worked rating-scale-to-codebook transformation for the MADRS
Take a concrete instrument. The Montgomery-Åsberg Depression Rating Scale is a 10-item, clinician-rated measure of depression severity, each item scored 0 to 6 for a total of 0 to 60, designed to be sensitive to change (Montgomery & Åsberg, 1979). Its items are already close to observable speech, which makes it a clean rating-scale-to-codebook example. Below, five of the ten items are carried through the transformation: each item becomes a code with an anchor rule (what the span must show) and a decision rule (the boundary that keeps it distinct from its neighbors).
| MADRS item (0–6) | Code | Anchor: what the span must show | Decision rule |
|---|---|---|---|
| Reported Sadness | RSD | The subject's own words describing low or flat mood | Applies to self-reported mood; if sadness is only inferred from described affect, code Apparent Sadness (ASD) instead |
| Inability to Feel | INF | Loss of interest or of emotional response to things once felt | Applies when pleasure or emotion is absent ("felt nothing"); do not apply to low mood alone—that is RSD |
| Concentration Difficulties | CNC | Reported trouble collecting or holding a train of thought | Applies to subject-reported focus problems; do not code the coder's own impression that the subject seemed distracted |
| Pessimistic Thoughts | PES | Guilt, self-blame, worthlessness, or hopelessness about the future | Applies to negative self- or future-appraisal; escalate to Suicidal Thoughts (SUI) only if death or self-harm is named |
| Suicidal Thoughts | SUI | Any statement that life is not worth living, death wishes, or self-harm | Applies and is flagged for clinician review regardless of severity band, even at anchor 0 |
Read the decision-rule column carefully, because that is where a rating scale is silent and a codebook has to speak. Reported Sadness and Apparent Sadness both describe low mood, but one lives in the subject’s words and the other in observed affect; a text-only codebook has to send the coder to the right one. Pessimistic Thoughts and Suicidal Thoughts sit on a continuum, so the boundary rule tells the coder precisely where one ends and the other begins. None of that is in the scale manual’s severity anchors—it is the operational scaffolding you add.
A copy-paste codebook entry skeleton
Every code in the table expands into the same fixed shape. Fill one of these per item and you have a codebook; the skeleton also maps one-to-one onto a Tagaroo phenomenon skill file (name, mnemonic, definition, examples), which is why a library scale can ship as a ready codebook.
CODE: <SHORT_LABEL> (<MNEMONIC>)
Source item: <instrument, item name/number, rating range>
Definition: <one sentence — what a span must express to earn this code>
Apply when: <the observable cue in the transcript>
Do NOT apply when: <the boundary; name the neighbor code that fits instead>
Positive examples (synthetic):
- "<invented quote that clearly fits>"
Near-miss examples (synthetic):
- "<invented quote that looks close>" -> code as <OTHER_CODE>
Severity anchors -> span features:
0: <what a 0 span shows>
2: <what a 2 span shows>
4: <what a 4 span shows>
6: <what a 6 span shows>
Unit of analysis: <utterance | span | speaker turn>
Source: <author, year, DOI>
Version: <n> Last revised: <YYYY-MM-DD>
The Do NOT apply when and Near-miss fields do the heaviest lifting. DeVellis and Thorpe make the general measurement point that a construct is defined as much by what it excludes as by what it includes (DeVellis & Thorpe, 2021); in a codebook, the exclusion rule is where that exclusion becomes something a coder can act on. The full item walk-throughs, anchor by anchor, are in the MADRS scoring guide, and the same evidence-first habit—find the span, then rate—drives scoring a scale from the transcript rather than a global impression.
Pilot the codebook and measure agreement
A codebook is validated by piloting, not by inspection. Have two coders independently apply the draft to the same sample of synthetic transcripts, compare their labels span by span, and compute an agreement coefficient before you trust a single downstream number. DeCuir-Gunby and colleagues fold this into codebook development itself: you draft, you train other coders on the codebook, and you establish reliability as part of building it, not as an afterthought (DeCuir-Gunby et al., 2011).
Which coefficient depends on your design. Cohen’s kappa fits two coders on nominal codes; Krippendorff’s alpha is the more general choice because it holds regardless of the number of coders, the level of measurement, sample size, and the presence of missing data (Hayes & Krippendorff, 2007). A widely cited convention reads kappa of 0.61 to 0.80 as substantial and 0.81 to 1.00 as almost perfect agreement, but Landis and Koch offered those labels as arbitrary benchmarks, not statistical law (Landis & Koch, 1977). The mechanics of both are worked through in the guides to Cohen’s kappa and to computing Krippendorff’s alpha.
Low agreement is a signal, not a failure. It usually points at a specific code whose boundary is underspecified—two coders reading the same span into different categories—which is precisely the disagreement that annotation guidelines that reduce disagreement are written to fix. That is the seventh step in action: revise the ambiguous entry, re-pilot, and version.
When the library already ships the codebook
For standard instruments, the operationalized codebook may already exist, and rebuilding it from scratch is wasted effort. Tagaroo’s Scale Library ships curated, span-level codebooks—each scale item pre-written as a code with a definition, ready to re-anchor with your own examples. Two facts a careful user should reconcile before relying on a card: how many items the library exposes, and how many the published instrument defines.
The MADRS is the clean case. Tagaroo’s curated MADRS exposes all ten published items (Montgomery & Åsberg, 1979), so the card and the instrument agree one-to-one—the codebook is the whole scale.
The TLC is the case that needs a note. Andreasen’s published Scale for the Assessment of Thought, Language, and Communication defines 18 subtypes of thought-language disorder plus a global rating (Andreasen, 1986); Tagaroo’s curated card exposes 12 of them—the core, frequently occurring, span-identifiable disorders (derailment, tangentiality, circumstantiality, poverty of speech, poverty of content, pressure of speech, distractible speech, illogicality, clanging, neologisms, perseveration, loss of goal).
It omits six that are either rare or hard to fix to a single span: incoherence, word approximations, echolalia, blocking, stilted speech, and self-reference. If you report a TLC assessment, use the published 18-item scale as your reference; the curated 12 are a span-annotation starting point, not a re-labeling of the full instrument. The full item glossary is in the TLC thought-language guide.
Where the rating-scale-to-codebook mapping breaks down
The mapping is not always clean, and pretending it is would set up a coder to fail. Three failure modes recur.
Observational items resist text. Some scale items live in what a clinician sees, not in what a subject says. MADRS Apparent Sadness is read from affect; several Hamilton items are objective or observational. A text-only span cannot fully justify those, which is why the curated HAM-D codebook is smaller than the published 17-item scale—the purely observational and measured items do not map to a transcript span.
Count and threshold logic does not live in a single utterance. Criteria-style instruments often require a symptom to persist for a duration or a count of items to cross a threshold. A codebook labels spans; the counting and thresholding is a separate scoring rule layered on top, and conflating the two produces codes that quietly encode a diagnosis they should not. Keep the span label and the scoring logic in different places.
Provenance is not validity. A beautifully operationalized codebook built on a poorly chosen scale still measures the wrong thing precisely. Operationalization makes coding consistent; it does not make the underlying instrument valid for your question (DeVellis & Thorpe, 2021). Choose the instrument first, then operationalize it. The same discipline—exhaustive, mutually exclusive, boundary-ruled categories—governs designing an image-labeling taxonomy, which is the rating-scale-to-codebook problem in a different modality.
From an instrument to a working coding manual
The practical upshot: a rating scale hands you names and severity anchors, and a codebook is what you build to make those names mean the same thing to two different coders. The rating-scale-to-codebook conversion is mostly the four fields the scale leaves out—operational definition, inclusion, exclusion, and examples—plus a pilot that tells you whether the boundaries actually hold. Do that once per item, measure agreement, version the result, and you have turned a published instrument into a coding manual a team, or an assisted agent, can apply consistently.
References
- MacQueen, K. M., McLellan, E., Kay, K., & Milstein, B. (1998). Codebook development for team-based qualitative analysis. Cultural Anthropology Methods, 10(2), 31–36. doi:10.1177/1525822X980100020301
- DeCuir-Gunby, J. T., Marshall, P. L., & McCulloch, A. W. (2011). Developing and using a codebook for the analysis of interview data. Field Methods, 23(2), 136–155. doi:10.1177/1525822X10388468
- Pustejovsky, J., & Stubbs, A. (2012). Natural Language Annotation for Machine Learning. O’Reilly Media. ISBN 9781449332693
- Saldaña, J. (2025). The Coding Manual for Qualitative Researchers (5th ed.). SAGE Publications.
- DeVellis, R. F., & Thorpe, C. T. (2021). Scale Development: Theory and Applications (5th ed.). SAGE Publications.
- Landis, J. R., & Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1), 159–174. doi:10.2307/2529310
- Hayes, A. F., & Krippendorff, K. (2007). Answering the call for a standard reliability measure for coding data. Communication Methods and Measures, 1(1), 77–89. doi:10.1080/19312450709336664
- Krippendorff, K. (2004). Reliability in content analysis: Some common misconceptions and recommendations. Human Communication Research, 30(3), 411–433. doi:10.1111/j.1468-2958.2004.tb00738.x
- Montgomery, S. A., & Åsberg, M. (1979). A new depression scale designed to be sensitive to change. British Journal of Psychiatry, 134(4), 382–389. doi:10.1192/bjp.134.4.382
- Andreasen, N. C. (1979). Thought, language, and communication disorders. I. Clinical assessment, definition of terms, and evaluation of their reliability. Archives of General Psychiatry, 36(12), 1315–1321. doi:10.1001/archpsyc.1979.01780120045006
- Andreasen, N. C. (1986). The Scale for the Assessment of Thought, Language, and Communication (TLC). Schizophrenia Bulletin, 12(3), 473–482. doi:10.1093/schbul/12.3.473
If you code symptoms from interviews, Tagaroo turns instruments like the MADRS and the TLC into guided, span-level codebooks where each item ships as a code with its definition, examples stay editable, and inter-rater agreement is computed as your coders work. Start from the operationalized scale, pilot it on your own material, and version the codebook that comes out.
Frequently asked questions
- What is the difference between a rating scale and an annotation codebook?
- A rating scale tells a trained rater how severe a symptom is; an annotation codebook tells any coder which span of text counts as that symptom in the first place. The scale assumes rater training and rates a whole interview to a number, while a codebook makes the decision explicit at the level of a span or utterance, with an operational definition, inclusion and exclusion rules, and examples (MacQueen et al., 1998). The codebook is what lets two coders, or a coder and an assisted agent, apply the same label the same way twice.
- How do you operationalize a rating scale item into a code?
- You convert the item's clinician-facing description into a rule about text: what a span must express to earn the code, when not to apply it, and which neighboring code fits instead. In annotation terms this is turning an existing standard into a specification, the concrete form of your model (Pustejovsky & Stubbs, 2012). Each scale item becomes one theory-driven code, taken deductively from the instrument rather than discovered in the data (DeCuir-Gunby et al., 2011).
- How many examples should each codebook entry include?
- At least one positive example and one near-miss for every code. MacQueen and colleagues found that pairing each code's definition with explicit inclusion criteria, exclusion criteria, and examples is what raises intercoder agreement, because most disagreement comes from ambiguous boundaries rather than obvious cases (MacQueen et al., 1998). Near-miss examples that name the correct alternative code are more useful than extra positive examples.
- How do you know a codebook is reliable?
- Pilot it: have two coders independently apply the draft to the same sample, then compute an agreement coefficient. Use Cohen's kappa for two coders on nominal codes and Krippendorff's alpha when you have more than two coders, ordinal severity, or missing data (Hayes & Krippendorff, 2007). A common benchmark treats kappa of 0.61 to 0.80 as substantial and 0.81 to 1.00 as almost perfect, though those labels are conventions, not hard thresholds (Landis & Koch, 1977).
- Do I have to build a codebook from scratch for a standard scale?
- Not always. For widely used instruments, the operationalized codebook may already exist. Tagaroo's Scale Library ships curated, span-level codebooks for scales like the MADRS and the TLC, with each item pre-written as a code and definition; you can start from those and re-anchor the examples to your own material rather than writing the operational definitions from a blank page.
Put this into practice
Tagaroo turns any rating scale or coding scheme into a guided annotation workflow — with inter-rater reliability computed as you go.