tagaroo

methods

Localize a Rating Scale: Annotating in a Second Language

Localize a rating scale for coding in a second language: what back-translation validates, what it misses, and how to re-anchor the coding examples.

Enrique Gutiérrez14 min readUpdated July 2026
Two parallel segmented speech-waveform ribbons, one per language, aligned span for span except a single coral bracket sitting at a different height, illustrating a severity anchor whose meaning shifts across languages.

A team hands you a Spanish-language interview and an English rating scale and asks for severity codes. The obvious first move—translate the instrument—is necessary and nowhere near sufficient. To localize a rating scale for annotation, you have to do something a translation never does on its own: re-anchor the examples a coder matches against, and re-earn the reliability numbers in the target language.

This is where scale translation validation and transcript coding part ways. The published methods for adapting an instrument were built for a patient filling in a form. Coding is a different act—an annotator reading what the subject actually said and binding a span of it to a severity anchor. The same translated wording can pass every back-translation check and still push a coder toward the wrong number.

What does it mean to localize a rating scale for annotation?

To localize a rating scale is to adapt the whole coding system—items, response anchors, examples, and the reliability evidence—so it measures the same construct in a new language and culture, not merely to translate its words. Cross-cultural adaptation researchers name four targets: semantic, idiomatic, experiential, and conceptual equivalence (Guillemin et al., 1993; Beaton et al., 2000). Translation reaches the first two on a good day. Coding lives or dies on the last two.

There are two layers to any coded scale, and they localize differently. The first is the instrument layer: the item definitions and response options a patient or rater reads. The second is the coding layer: the aliases, cue phrases, and worked examples an annotator matches against a subject’s actual speech.

Tagaroo’s library ships several scales bilingually—the GAD-7, the PHQ-9, and the Empathic Communication Coding System all carry English and Spanish—so the instrument layer is already adapted. The coding layer is the work that remains, and standard translation guidance was never designed to cover it. For how that layer is built from a scale in the first place, see turning a scale into a codebook.

Forward translation, back-translation, or harmonization: what does each step validate?

Each step guards a different failure, and none of them guards the one that breaks coding. Forward translation produces fluent target-language wording; back-translation checks that wording didn’t distort the source; harmonization keeps multiple language versions consistent with each other. Read across the table and the gap is visible: every column’s “what it can’t catch” is some version of meaning.

StepDirectionWhat it validatesWhat it can't catch
Forward translationSource → targetFluent, natural target wording by ≥2 independent translators, then reconciled (semantic + idiomatic equivalence)Whether the meaning drifted from the source
Back-translationTarget → source (blind)That the translation did not distort the source words, via a translator who never saw the original (Beaton et al., 2000; Guillemin et al., 1993)Whether the item means the same thing to a patient, or codes the same way against real speech
HarmonizationAcross all language versionsConsistency across a multilingual or multi-site study, versions checked against the source and each other (Wild et al., 2005)A culture-specific meaning shift that every version happens to share
The three moves at the heart of scale translation validation. Forward and back-translation are the classic pair (Beaton et al., 2000); harmonization is the step most guidelines under-specify (Wild et al., 2005). Survey methodology reaches the same place through a team model—TRAPD: Translation, Review, Adjudication, Pretesting, Documentation (Harkness, 2003).

The fuller pipelines make the point by omission. The ISPOR linguistic-validation process runs ten steps—preparation, forward translation, reconciliation, back translation, back-translation review, harmonization, cognitive debriefing, review and finalization, proofreading, and a final report (Wild et al., 2005). Nine of the ten govern the instrument. Only cognitive debriefing—sitting real respondents in front of the draft—tests whether the words land as intended, and even that tests comprehension of a form, not the coding of a conversation. The International Test Commission’s guidelines span 18 recommendations across six categories, yet a systematic review found most published adaptations did not follow them (International Test Commission, 2018).

When does a translated severity anchor mean something different?

A translated anchor means something different whenever the target word carries a construct the source word doesn’t—which is exactly what back-translation is blind to. The four equivalences pull apart here. Semantic and idiomatic equivalence are about surface language; experiential and conceptual equivalence are about whether the thing being described exists and means the same in the target culture (Beaton et al., 2000).

EquivalenceQuestion it asksWhere coding trips
SemanticDo the words carry the same meaning?The Spanish 'sueño' means both sleep and dream—a PHQ-9 sleep item must disambiguate
IdiomaticDo idioms survive translation?'Feeling blue' has no literal Spanish equivalent; it needs a natural idiom, not a word-for-word render
ExperientialDoes the described situation exist in the target culture?An item keyed to one health system's paperwork may have no counterpart in another
ConceptualDoes the underlying construct match?'Nervios' indexes a broader distress idiom than the trait worry a scale item intends (Lewis-Fernández et al., 2010)
The four equivalences a cross-cultural scale adaptation targets (Guillemin et al., 1993; Beaton et al., 2000). Back-translation mainly protects the first two; the last two decide whether coding transfers.

The clearest worked example is the GAD-7’s first item. In English it reads “feeling nervous, anxious, or on edge.” Now take a synthetic Spanish reply a coder might meet:

Entrevistador: ¿Cómo se ha sentido estas dos semanas?

Sujeto: Ay, doctor, ando de los nervios—cualquier cosa me altera y siento que no me puedo controlar.

A coder localizing the GAD-7 sees “de los nervios” and reaches for that first item. But “andar de los nervios,” and the related “ataque de nervios,” is a cultural idiom of distress among many Latin American communities that can index anger, loss of control, or a grief reaction rather than the future-oriented, uncontrollable worry the GAD-7 is built to measure (Lewis-Fernández et al., 2010). Back-translate the phrase and you get “I’m on my nerves / feeling nervous”—which maps back to item 1 cleanly. The words round-trip. The coding decision does not.

This is not an argument that adaptation is hopeless—it is an argument for doing it, and for reusing the versions that already have. When the GAD-7 was culturally adapted into Spanish, pre-testing forced the team to modify not just wording but the header and the response-category anchors, and the resulting version reached an optimal screening cut-off of 10 (sensitivity 86.8%, specificity 93.4%; García-Campayo et al., 2010). The anchors had to be re-anchored, not merely translated—and that step, not the raw translation, is what made the threshold hold.

Why translating the items isn’t enough to localize a rating scale for coding

Because a coder is not filling in the form—they are matching a subject’s spontaneous words to an anchor, and translation validates the form, not the match. This is the difference between a patient answering “nearly every day” on a printed scale and an annotator deciding whether “casi no duermo, doctor” belongs at anchor 2 or anchor 3. The instrument can be perfectly localized and the second decision still drift.

Practically, that means the aliases and examples in the codebook need their own adaptation pass. An English PHQ-9 coding note might list “can’t get going,” “everything is an effort,” “no energy” as cues for the fatigue item. The Spanish coding layer needs the phrases Spanish speakers actually use—“no tengo ganas,” “me cuesta todo,” “ando sin fuerzas”—chosen by bilingual clinicians, not translated from the English cue list. Miss this and your coders are pattern-matching against phrases no subject will ever say. This is the same discipline as writing annotation examples that hold up, done twice, once per language, with the second pass grounded in the target culture rather than the first.

How do you re-establish inter-rater reliability in the target language?

You re-run the reliability study from scratch on target-language transcripts—an English kappa tells you nothing about a Spanish codebook. Reliability is a property of the raters, the materials, and the population together, not of the construct in the abstract, so it does not travel across a translation (International Test Commission, 2018). A codebook that scored a Cohen’s kappa of 0.82 in English can land far lower in its localized form if the re-anchored examples are thin or the idioms are ambiguous.

The fix is a target-language pilot annotation round: have at least two bilingual coders independently code a sample of real target-language transcripts, compute agreement, and revise the coding layer where they disagree—before any production coding starts. Then report the language of the reliability sample alongside the coefficient. “κ = 0.79” means little if a reader can’t tell whether it was earned on English or Spanish material. Making that provenance visible is part of keeping the score auditable: every rating tied to its span, in the language it was spoken.

How do you annotate code-switched or bilingual transcripts?

Code-switched transcripts—where a subject moves between languages, sometimes mid-sentence—are common and fully codeable, with three rules. First, document the working languages and use bilingual coders who rate in the language actually spoken. Second, anchor every rating to the original-language span, never to a coder’s on-the-fly translation, so the evidence stays intact. Third, measure cross-linguistic annotation agreement across the bilingual team, because two coders can split on a switched phrase for reasons that have nothing to do with severity.

The trap is “translate then code”: rendering a Spanish-English span into one language first and coding the translation. That buries the exact judgment you are trying to make reliable, and it launders a cultural idiom into whatever the translator’s word choice implied. Keep the code close to the words. The point of coding from transcripts, as opposed to a live clinician rating, is that the evidence is still there to check—so don’t overwrite it with a translation.

A checklist to localize a rating scale you’ll code in a second language

The steps below assemble the published adaptation methods and the coding-specific layer into one order of operations:

  1. Start from a validated version if one exists—the Spanish PHQ (Díez-Quevedo et al., 2001) or Spanish GAD-7 (García-Campayo et al., 2010)—rather than translating from scratch.
  2. Run the full adaptation, not a translation: forward translation by two or more translators, reconciliation, back-translation, expert-committee review or harmonization, and cognitive debriefing (Beaton et al., 2000; Wild et al., 2005).
  3. Re-anchor the coding examples and aliases, not only the item text—the cue phrases coders match spans against, chosen by bilingual clinicians in the target language.
  4. Localize and pilot the response anchors too; the Spanish GAD-7 had to rework its response categories after pre-testing (García-Campayo et al., 2010).
  5. Flag culturally bound idioms of distress (such as “nervios”) with explicit coding rules so annotators don’t map them at face value (Lewis-Fernández et al., 2010).
  6. Re-run a pilot and recompute inter-rater reliability on target-language transcripts; never inherit the source-language coefficient.
  7. Document everything—translators, decisions, and the language of the reliability sample (International Test Commission, 2018; the “D” for Documentation in Harkness’s TRAPD, 2003).

Limitations: when to use an existing validated version instead

The honest limit of this piece is that most teams should not be adapting a scale at all. If a validated target-language version exists, use it—developing and validating a new adaptation is a multi-year psychometric project, and a home-brewed translation with a re-anchored codebook is not a substitute for one. Adaptation also cannot rescue a construct that doesn’t hold cross-culturally: if the underlying phenomenon is defined differently in the target culture, no amount of re-anchoring makes the scores comparable, and forcing them is worse than reporting them separately (Lewis-Fernández et al., 2010). And a localized codebook is a claim you have to defend with data, not a translation you can quietly ship.

Within those limits, the discipline is straightforward and it pays off. To localize a rating scale you’ll code from interviews, adapt the instrument by the book, then do the part the book skips—rebuild the coding examples in the target language and re-earn your reliability there. Tagaroo is built for that second half: bilingual scales like the GAD-7 and PHQ-9 with evidence-anchored coding, so every rating is pinned to the span that justifies it, in the language it was spoken, with inter-rater reliability computed as your coders work. For the broader question of when a rating scale should become a coding scheme at all, see rating scales versus coding schemes and mapping diagnostic criteria to annotations.

References

  • Beaton, D. E., Bombardier, C., Guillemin, F., & Ferraz, M. B. (2000). Guidelines for the process of cross-cultural adaptation of self-report measures. Spine, 25(24), 3186–3191. doi:10.1097/00007632-200012150-00014
  • Guillemin, F., Bombardier, C., & Beaton, D. (1993). Cross-cultural adaptation of health-related quality of life measures: literature review and proposed guidelines. Journal of Clinical Epidemiology, 46(12), 1417–1432. doi:10.1016/0895-4356(93)90142-N
  • Wild, D., Grove, A., Martin, M., Eremenco, S., McElroy, S., Verjee-Lorenz, A., & Erikson, P. (2005). Principles of good practice for the translation and cultural adaptation process for patient-reported outcomes (PRO) measures: report of the ISPOR Task Force for Translation and Cultural Adaptation. Value in Health, 8(2), 94–104. doi:10.1111/j.1524-4733.2005.04054.x
  • International Test Commission. (2018). ITC guidelines for translating and adapting tests (second edition). International Journal of Testing, 18(2), 101–134. doi:10.1080/15305058.2017.1398166
  • García-Campayo, J., Zamorano, E., Ruiz, M. A., Pardo, A., Pérez-Páramo, M., López-Gómez, V., Freire, O., & Rejas, J. (2010). Cultural adaptation into Spanish of the generalized anxiety disorder-7 (GAD-7) scale as a screening tool. Health and Quality of Life Outcomes, 8, 8. doi:10.1186/1477-7525-8-8
  • Díez-Quevedo, C., Rangil, T., Sánchez-Planell, L., Kroenke, K., & Spitzer, R. L. (2001). Validation and utility of the patient health questionnaire in diagnosing mental disorders in 1003 general hospital Spanish inpatients. Psychosomatic Medicine, 63(4), 679–686. doi:10.1097/00006842-200107000-00021
  • Lewis-Fernández, R., Hinton, D. E., Laria, A. J., Patterson, E. H., Hofmann, S. G., Craske, M. G., et al. (2010). Culture and the anxiety disorders: recommendations for DSM-V. Depression and Anxiety, 27(2), 212–229. doi:10.1002/da.20647
  • Harkness, J. A. (2003). Questionnaire translation. In J. A. Harkness, F. J. R. van de Vijver, & P. Ph. Mohler (Eds.), Cross-Cultural Survey Methods (pp. 35–56). Hoboken, NJ: Wiley.
  • World Health Organization. (n.d.). Process of translation and adaptation of instruments. WHO management of substance abuse

Frequently asked questions

Is translating a rating scale enough to use it for annotation in another language?
No. A careful translation plus cross-cultural adaptation validates the questionnaire a patient reads (Beaton et al., 2000; Wild et al., 2005), but coding transcripts adds two requirements translation never covers: you must re-anchor the examples and aliases a coder matches spans against, and re-establish inter-rater reliability on target-language transcripts. Back-translation checks that the words line up; it cannot confirm that the coding decisions line up.
What is the difference between forward translation, back-translation, and harmonization?
Forward translation renders the source into the target language, typically by two or more independent translators whose drafts are reconciled (Wild et al., 2005). Back-translation sends the target version back into the source language via a translator blind to the original, to surface drift (Beaton et al., 2000; Guillemin et al., 1993). Harmonization compares all language versions against the source and each other so a multilingual or multi-site study stays consistent (Wild et al., 2005). The first two protect wording; harmonization protects consistency; none of the three settles culture-specific meaning.
Can I reuse my English codebook's kappa in another language?
No. Reliability is a property of the raters, the materials, and the population, not of the construct in the abstract, so an English inter-rater agreement figure does not carry over. Re-run a pilot round and recompute agreement on target-language transcripts before you trust the localized codebook (International Test Commission, 2018).
How do I code a transcript where the subject switches between languages?
Code-switching is common and codeable. Document the working languages, use bilingual coders who rate in the language actually spoken, anchor every rating to the original-language span rather than to a coder's on-the-fly translation, and measure cross-linguistic annotation agreement across the bilingual team. Translating the span first and coding the translation buries the exact decision you are trying to make reliable.
Is there a validated Spanish version of the PHQ-9 and GAD-7?
Yes. The Patient Health Questionnaire was validated in Spanish against clinician diagnosis (Cohen's κ = 0.74; Díez-Quevedo et al., 2001), and the GAD-7 was culturally adapted into Spanish with an optimal screening cut-off of 10 (García-Campayo et al., 2010). Start from a published, validated version rather than translating one yourself.

Put this into practice

Tagaroo turns any rating scale or coding scheme into a guided annotation workflow — with inter-rater reliability computed as you go.