Study workflow

Annotation codebook generator

The hard part of a codebook is not the format — it is knowing which categories to use and where their boundaries sit. Twenty-five published instruments already answer that for a lot of constructs. Describe what you want to code and take the result away as a file.

Free · No sign-up · Runs entirely in your browser

Describe what you want to code

Searches 187 categories from 25 published instruments, bundled into this page — no account, no upload, and no network request.

  • Conceptual DisorganizationCDS

    Brief Psychiatric Rating Scale (BPRS) · 1962

    Disorganized thought shown by disconnected, tangential, or incoherent speech.

    Matched on code name: “Conceptual Disorganization”

  • Distractible SpeechDST

    SAPS: Positive Formal Thought Disorder Subscale · 1984

    Stopping mid-sentence and changing the subject in response to a nearby stimulus.

    Matched on code name: “Distractible Speech”

  • Distractible SpeechDST

    Scale for the Assessment of Thought, Language and Communication (TLC) · 1986

    Stopping mid-sentence and changing topic in response to a nearby stimulus.

    Matched on code name: “Distractible Speech”

  • Poverty of Content of SpeechPOC

    SANS: Alogia Subscale · 1984

    Adequate amount of speech that conveys little information — vague, over-abstract, or repetitive.

    Matched on code name: “Poverty of Content of Speech”

  • Speech (Rate and Amount)SPE

    Young Mania Rating Scale (YMRS) · 1978

    Increased rate and volume of speech that is pressured and hard to interrupt.

    Matched on code name: “Speech (Rate and Amount)”

  • Reproduction of ConversationCNV

    Criteria-Based Content Analysis (CBCA) · 1989

    Speech during the event is reproduced, ideally quoting distinct speakers verbatim.

    Matched on alias: “quoted speech”

  • RetardationRET

    Hamilton Depression Rating Scale (HDRS-17) · 1960

    Reported or observed slowness of thought, speech, and movement.

    Matched on alias: “slowed speech”

  • Unusual Thought ContentUTC

    Brief Psychiatric Rating Scale (BPRS) · 1962

    Unusual, odd, or bizarre beliefs and delusional ideas voiced in speech.

    Matched on definition: “Unusual, odd, or bizarre beliefs and delusional ideas voiced in speech.”

  • Pressure of SpeechPRS

    SAPS: Positive Formal Thought Disorder Subscale · 1984

    Increased, rapid, hard-to-interrupt speech, often with an increased amount of talk.

    Matched on code name: “Pressure of Speech”

  • Poverty of Content of SpeechPOC

    Scale for the Assessment of Thought, Language and Communication (TLC) · 1986

    Adequate amount of speech that conveys little information; vague, repetitive, empty.

    Matched on code name: “Poverty of Content of Speech”

  • Poverty of SpeechPOS

    SANS: Alogia Subscale · 1984

    Restriction in the amount of spontaneous speech; replies are brief, terse, and unelaborated.

    Matched on code name: “Poverty of Speech”

  • Language-Thought DisorderLTD

    Young Mania Rating Scale (YMRS) · 1978

    Disorganized, tangential, circumstantial, or flight-of-ideas language; distractible, hard to follow.

    Matched on definition: “Disorganized, tangential, circumstantial, or flight-of-ideas language; distractible, hard to follow.”

Reference · for the curious

What a codebook has to contain before two coders can agree

In our experience most annotation projects end up writing their codebook twice: once before coding, quickly and optimistically, and once after the first pilot round has shown which entries two people read differently. The second version is the real one. What makes the difference is not the template but the specificity — and specifically, whether each entry says where its category stops.

The seven fields, and which two matter most

A usable entry names the code, defines in one sentence what a span must express to earn it, states the observable cue that triggers it, states the boundary where a neighbouring code applies instead, gives a synthetic example that clearly fits and one that looks close but does not, and records the unit of analysis, the source and the version.

The two exclusion fields — do not apply when and the near-miss example — carry most of the weight. DeVellis and Thorpe make the general measurement point that a construct is defined as much by what it excludes as by what it includes; in a codebook, the exclusion rule is where that becomes something a coder can act on. Every disagreement you will have to adjudicate is a case that sat near a boundary somebody left implicit. Our full method for turning a rating scale into a codebook walks the same skeleton item by item.

Why start from a published instrument

Because the categories and their boundaries are the expensive part, and for many constructs somebody has already done that work and had it reviewed. Starting from the TLC's twelve signs of disordered speech, or the MADRS's ten items, gives you definitions that survived peer review and a citation that tells a reader what you operationalised. It also gives you published reliability figures, which are a useful warning: in Andreasen's original reliability study four of the eighteen thought-and-language items fell below a kappa of 0.6, so if your own coders struggle with tangentiality you are in known territory rather than doing something wrong.

The library here covers twenty-five instruments across thought and language disorder, mood and anxiety, emotion taxonomies, therapy and consultation process, forensic content analysis, interview protocols and discourse structure. Each is browsable in full in our scale library, with the item-level detail this generator condenses.

Codebook vs annotation guidelines

These are related documents with different jobs, and conflating them is why some projects end up with neither. The codebook is the per-code reference: one entry per category, definition, boundary, examples. The guidelines are the surrounding protocol: what unit to annotate, what to do when nothing applies, how to handle overlapping spans, when to flag rather than guess, and the running log of edge-case rulings the team has made.

A codebook without guidelines leaves coders inventing process; guidelines without a codebook leave them inventing categories. Our companion piece on writing annotation guidelines that reduce disagreement covers the second half, including the dated edge-case decision log that is the single cheapest quality intervention available to a coding team.

Piloting is not optional, and the number tells you which entry to fix

A codebook is validated by piloting rather than by inspection. Two coders apply the draft independently to the same sample, you compare span by span, and you compute an agreement coefficient before trusting anything downstream. DeCuir-Gunby and colleagues fold this into codebook development itself — you draft, you train coders on the draft, and you establish reliability as part of building it rather than as an afterthought.

Which coefficient depends on the design: Cohen's kappa for two coders on nominal codes, Krippendorff's alpha when you have more coders, ordered categories or missing ratings, an ICC for continuous ratings. Our inter-rater reliability calculator computes those side by side, the ICC calculator handles the continuous case, and if your codes mark stretches of text rather than whole items you want the span agreement calculator instead, because chance-corrected agreement does not apply when coders choose the units themselves.

The useful habit is to read low agreement as a pointer rather than a verdict. It almost always localises to one code whose boundary is underspecified — two coders reading the same span into different categories — which tells you precisely which entry to revise before re-piloting. That is what the version and revision-date fields are for.

What this tool does not do

It does not write your exclusion rules or your example quotes. That is a deliberate refusal rather than a missing feature: an invented boundary rule would read as authoritative while being a guess, and example quotes have to be synthetic — never lifted from a real transcript — which is a judgement only you can make for your data. The audit names each gap instead.

It also covers only the twenty-five instruments in the bundled library, all of which are freely reproducible with citation. Many rating scales are licensed, and reproducing a copyrighted scoring manual is a different matter from citing an item definition — if your construct needs an instrument that is not here, the entry shape still applies, but the wording has to come from your own licensed copy.

Frequently asked questions

What is an annotation codebook?

A codebook is the document that turns a construct into something two people can code the same way. Each entry names a code, defines it in one sentence, states the observable cue that triggers it, states the boundary where a neighbouring code applies instead, gives synthetic examples of both a clear fit and a near miss, and records the unit of analysis, the source and the version. The exclusion rule and the near-miss examples do the heaviest lifting: a construct is defined as much by what it rules out as by what it includes, and in a codebook that exclusion is where the definition becomes something a coder can act on.

How do I turn a rating scale into a codebook?

Take each item of the instrument as a candidate code, keep the published definition as the entry's definition, then add the three things the instrument does not give you: the observable cue in a transcript, the boundary against the neighbouring item, and example spans. If the item has severity anchors, map each anchor onto what a span at that rating actually shows. That last step is where most of the work sits, because a published anchor like "moderate" describes a clinical impression rather than a stretch of text. This generator seeds the first part from the library and leaves the rest to you, because inventing an exclusion rule would be guessing.

Which instruments does the library cover?

Twenty-five, spanning thought and language disorder (TLC, SAPS-FTD, SANS-alogia), mood and anxiety (PHQ-9, GAD-7, HAM-D, HAM-A, MADRS, YMRS, BPRS), emotion taxonomies (Ekman, Plutchik), therapy and consultation process (MITI, MISC, ECCS, OPTION, VR-CoDES), forensic and content analysis (CBCA, Gottschalk-Gleser), interview protocols (NICHD), CBT cognitive distortions, and discourse and argument structure (Searle, Labov-Waletzky, Toulmin, IRF). Each carries its own citation, DOI where one exists, and an explicit statement of the terms under which its definitions may be reproduced.

Can I use these definitions in my own research?

For every instrument in this library, yes — that is why these twenty-five were curated. Each one is either public domain, published open access, distributed free for research and education by its authors, or a conceptual framework whose operational wording here was written in-house rather than copied from a manual. Each export carries the relevant reproduction terms into the file so the provenance travels with the codebook, and every entry keeps its citation. What you should not do is assume the same applies to instruments outside this library: many rating scales are licensed, and reproducing a copyrighted scoring manual is a different matter from citing an item definition.

Do I still need to pilot the codebook?

Yes, and no amount of care in drafting replaces it. A codebook is validated by piloting rather than by inspection: two coders apply the draft independently to the same sample, you compare their labels span by span, and you compute an agreement coefficient before trusting a single downstream number. Low agreement is a signal rather than a failure — it usually points at one specific code whose boundary is underspecified, which tells you exactly which entry to revise. Then re-pilot and increment the version, which is why the entry shape carries a version and a revision date.

Written by Enrique Gutiérrez, PhD (Computer Science) — founder of Tagaroo and Associate Professor of Computer Science, working on inter-rater reliability, measurement and annotation methodology (ORCID).

Last verified: 30 July 2026. Formulas, thresholds and cited figures on this page were checked against their original sources on that date. Every calculation runs in your browser; nothing you enter is transmitted or stored.

Skip the export and code against it directly

The file this generator produces is the thing Tagaroo loads natively. Curated scales arrive as pre-written anchored codebooks, coders work against them span by span, and inter-rater reliability is computed as they go — so the pilot round this page tells you to run is a normal week of work rather than a separate project.

Try Tagaroo free