Rating ScalesAuditable Clinical Ratings: Score From the Transcript
Build an auditable clinical rating by tying every scale item to a quoted line, not a global impression. See how span-grounded scoring works.
The Tagaroo Blog
Field notes on annotation: inter-rater reliability, clinical rating scales, qualitative coding, image labeling, and running annotation campaigns that hold up.
0 matching articles
Validated clinician- and self-report instruments — items, anchors, cutoffs, and how to code them from a transcript.
Rating ScalesBuild an auditable clinical rating by tying every scale item to a quoted line, not a global impression. See how span-grounded scoring works.
Rating ScalesBPRS vs PANSS compared: item overlap, the positive-negative-general structure, trial usage, and score linking. See which psychosis scale to use.
Rating ScalesDepression rating scales compared: HAM-D, MADRS, PHQ-9, and content analysis by rater, item count, and change-sensitivity. See which fits your study.
Rating ScalesGAD-7 vs HAM-A compared: the 7-item self-report anxiety screen against the 14-item clinician-rated severity endpoint. See which anxiety scale to use.
Rating ScalesTLC vs SAPS thought disorder, compared: the 18-item TLC glossary vs the 8-item SAPS positive-FTD subscale, with an overlap grid and when to use each.
Rating ScalesMADRS vs HAM-D compared: items, somatic weighting, sensitivity to change, score conversion, and remission cutoffs. See which depression scale to use.
Rating ScalesThe Brief Psychiatric Rating Scale (BPRS) rates psychopathology on 7-point items. See the scoring, the factor structure, BPRS vs PANSS, and how to code it.
Rating ScalesThe Hamilton Anxiety Rating Scale (HAM-A) is a 14-item clinician measure. See the psychic and somatic items, the scoring, the cutoffs, and how to code it.
Rating ScalesThe GAD-7 is a seven-item anxiety screen. See the 5/10/15 severity bands, the ≥10 cutoff, and how to code it from an interview transcript.
Rating ScalesHow the Hamilton Depression Rating Scale (HAM-D) is scored and coded: the 17 items, the 0–52 range, remission and response cutoffs, and its known flaws.
Rating ScalesMADRS scoring explained: the 10 clinician-rated items, the 0-6 anchors, severity bands, and the response and remission cutoffs. See how it's coded.
Rating ScalesPHQ-9 scoring made clear: what the nine items measure, the 0-3 anchors, the severity bands, the ≥10 cutoff, and how to code a PHQ-9 from an interview.
Rating ScalesThe Young Mania Rating Scale (YMRS) explained: its 11 items, the four double-weighted items, severity bands, and how to rate mania from interview speech.
What the constructs actually look like in speech — and how to annotate them consistently.
Clinical PhenomenaEkman vs Plutchik for annotation: 6 discrete emotions or 8 with intensities and dyads, and why more categories lower agreement. See which to pick.
Clinical PhenomenaThe cognitive distortions list, explained: all 10 CBT thinking errors with definitions and examples, where the list came from, and why they resist coding.
Clinical PhenomenaEkman's basic emotions are six universal states: anger, disgust, fear, happiness, sadness, surprise. See why they anchor most emotion datasets.
Clinical PhenomenaPlutchik's wheel of emotions maps eight primaries into four opposing pairs, with intensity rings and dyads. See how to label emotion in text.
Clinical PhenomenaPoverty of speech is the core of alogia. See the four SANS alogia items, how they differ from poverty of content, and how to code each from a transcript.
Clinical PhenomenaPositive formal thought disorder is disorganized speech you can see. Meet the SAPS FTD items, how they differ from the TLC and SANS, and how to code them.
Clinical PhenomenaThe Thought Language and Communication scale (TLC) is Andreasen's glossary of disordered speech. See the items, the reliability data, and how to code them.
Coding schemes, taxonomies, and protocols for turning language into structured data.
Methods & FrameworksHow the VR-CoDES coding system codes patient emotional cues versus concerns and the provider response, using the providing-space vs reducing-space axis.
Methods & FrameworksAnnotation fatigue erodes label quality within a session. See the vigilance-decrement evidence and the breaks, caps, and catch-trials that hold it steady.
Methods & FrameworksAnnotator burnout has clear causes and warning signs. Learn the drivers, symptoms to watch for, and the workload and rotation practices that prevent it.
Methods & FrameworksEvidence mode rates the whole transcript once and cites its supporting spans; instance mode tags each occurrence. How the choice reshapes reliability math.
Methods & FrameworksEthical data annotation means fair pay, conditions, contracts, management, and representation. Audit any vendor against the Fairwork standard.
Methods & FrameworksConcrete protections for sensitive content annotation: informed consent, exposure limits, grayscale and blur defaults, and debriefing. Get the protocol.
Methods & FrameworksLocalize a rating scale for coding in a second language: what back-translation validates, what it misses, and how to re-anchor the coding examples.
Methods & FrameworksMulti-layer annotation stacks several coding schemes and rating scales on one transcript. Keep layers independent, handle overlaps, score each separately.
Methods & FrameworksRating scale vs coding scheme: how dimensional severity and categorical coding change your unit of analysis, reliability statistic, and annotation UI.
Methods & FrameworksHow to turn many annotator labels into one gold label: majority vote, Dawid-Skene, MACE, STAPLE, and expert adjudication, with a table for when each fits.
Methods & FrameworksHow teams apply text annotation use cases beyond the clinic: support-ticket intent, legal clause tagging, and survey coding, with span evidence and IRR.
Methods & FrameworksAn annotation cost calculator: turn items, labels per item, time, wage, redundancy, and QA overhead into a labeling budget and timeline. See the math.
Methods & FrameworksData annotation consent explained: HIPAA de-identification vs a BAA, GDPR special-category data, who owns the labels, and how to license a dataset.
Methods & FrameworksHow ground truth for medical AI is built for FDA review: reference standards, expert panels, MRMC reader studies, and the SegAgree agreement idea.
Methods & FrameworksWrite annotation guidelines that reduce annotator disagreement: operational definitions, positive and negative examples, and an edge-case decision log.
Methods & FrameworksAudio annotation turns raw speech into labeled data: transcription, speaker diarization, timestamped events, and prosody tags. Compare the types.
Methods & FrameworksCBCA vs reality monitoring: how two statement-credibility methods compare on criteria, accuracy, and court limits—and why neither one detects lies.
Methods & FrameworksA finding-by-finding chest X-ray annotation example: build a CheXpert-style schema, handle uncertainty labels, and pick image-level vs region labels.
Methods & FrameworksClinician-rated vs self-report scales compared: who rates, the bias each carries, reliability and cost trade-offs, and transcript coding as a third path.
Methods & FrameworksCriteria-based content analysis (CBCA) rates 19 content criteria to judge whether a statement reads as experience-based—an aid, not a lie test. See all 19.
Methods & FrameworksData cascades are compounding downstream failures from upstream data problems. 92% of AI teams hit them—here's why fixing labels beats tuning the model.
Methods & FrameworksDice vs IoU, the Jaccard index, and Hausdorff distance compared: formulas, the Dice-IoU identity, and when overlap metrics mislead. See which to report.
Methods & FrameworksA DICOM annotation and data-prep guide: strip header and burned-in pixel PHI, window images for display, and store labels as DICOM-SEG, SR, or JSON.
Methods & FrameworksForcing one gold label can throw away real information. See why annotator disagreement is signal, not noise, and when to reconcile versus preserve it.
Methods & FrameworksHow document annotation works for ML: layout regions, document entities, key-value pairs, and table structure, plus the datasets. See the tasks.
Methods & FrameworksDoes pay improve annotation quality? Higher pay lifts speed, participation, and fairness—but rarely accuracy. See what really drives label quality.
Methods & FrameworksTurn DSM-5 and ICD-11 criteria into observable, taggable phenomena: a worked MDD-to-PHQ-9 mapping and why you annotate symptoms, not a diagnosis.
Methods & FrameworksExpert vs crowd annotation: aggregated non-experts match experts on many tasks, but clinical judgment breaks the ratio. Get the decision guide.
Methods & FrameworksLabel errors hide in almost every dataset. Learn how to find label errors in dataset audits with disagreement, model confidence, and confident learning.
Methods & FrameworksGeospatial image annotation covers satellite, aerial, and microscopy imagery—where huge rasters, rare classes, and expert-only labels change the rules.
Methods & FrameworksGold questions, honeypots, and attention checks catch bad annotations—if thresholds spare legitimate disagreement. See how each QA method works.
Methods & FrameworksHow many annotators medical images need depends on clinical risk, from spot-checking low-risk labels to blind multi-reader reads with adjudication.
Methods & FrameworksHow many annotators per item do you need? One label is risky. Compare majority vote, Dawid-Skene, and MACE, and see where extra labels stop paying off.
Methods & FrameworksThe human cost of data labeling: the underpaid, invisible ghost workers behind AI, the toll of the work, and what responsible sourcing looks like.
Methods & FrameworksWho should annotate your data? Match annotator background to the task—expert, trained non-expert, or crowd—using a clear selection framework and matrix.
Methods & FrameworksBounding box, polygon, semantic vs instance vs panoptic masks, or keypoints? Compare image annotation types by cost, precision, and downstream task.
Methods & FrameworksWrite image labeling guidelines annotators actually apply: mutually exclusive vs overlapping classes, visual examples, and clear when-in-doubt rules.
Methods & FrameworksWhy Cohen's kappa fails on pixels, and how Dice, IoU, and the STAPLE algorithm measure image segmentation agreement across annotators. See the math.
Methods & FrameworksIRF classroom discourse is the three-move Initiation-Response-Feedback exchange behind most teacher talk. Learn to code it and read the F move.
Methods & FrameworksMedical image annotation, end to end: annotation types, who should label, consensus and adjudication, DICOM de-identification, and QA tied to risk.
Methods & FrameworksChange talk vs sustain talk, coded: how the MISC classifies client language, why the balance predicts outcomes, and how to segment it reliably.
Methods & FrameworksMISC vs MITI, disambiguated: the MITI codes the clinician's MI fidelity; the MISC codes the client's change and sustain talk. Compare purpose and use.
Methods & FrameworksHow MITI motivational interviewing coding works: behavior counts, four global scores, and the R:Q ratio and %complex reflections that grade MI fidelity.
Methods & FrameworksModel-assisted labeling lets a model draft image labels so annotators correct, not create. How pre-labeling and active learning cut work honestly.
Methods & FrameworksAnnotate aligned text, image, and audio in one schema for foundation models: cross-modal alignment, unified taxonomies, and QA that keeps labels in sync.
Methods & FrameworksHow to build multimodal image-text annotation datasets: pair reports to images, ground report phrases to regions, and link findings across modalities.
Methods & FrameworksWhole-slide image annotation for computational pathology: the gigapixel challenge, the slide-to-cell label hierarchy, and pathologist-in-the-loop QA.
Methods & FrameworksRun a pilot annotation round before you scale: scope, sample size, double-coding, measuring agreement, and revising the codebook. See the checklist.
Methods & FrameworksA step-by-step method to go from rating scale to codebook: turn each item into a code, anchors, examples, and decision rules coders apply consistently.
Methods & FrameworksTo improve annotator quality, screen candidates for aptitude and train them before coding—filtering bad labels later loses. See the evidence.
Methods & FrameworksSynthetic data vs human annotation: when LLM-generated labels are cheap and good, where they miss the tail, and how to build a hybrid you can trust.
Methods & FrameworksThe Empathic Communication Coding System codes empathy in clinical talk as a two-step sequence: patient opportunities, then a 0–6 provider response ladder.
Methods & FrameworksLabov and Waletzky narrative analysis breaks an oral story into six parts. See each part, why evaluation is the heart, and how reliably it can be coded.
Methods & FrameworksThe NICHD protocol classifies a forensic interviewer's prompts from open to suggestive. See the five types, why open prompts win, and how to code them.
Methods & FrameworksThe OPTION scale rates how far a clinician involves the patient in a decision. See the behaviors it scores, why real scores are low, and how to code it.
Methods & FrameworksSearle's speech acts sort every utterance into five illocutionary types. See the classes, direction of fit, examples, and the tie to dialogue-act tagging.
Methods & FrameworksThe Toulmin model of argumentation breaks an argument into six functional parts. See the elements, a worked example, and why warrants are hardest to code.
Methods & FrameworksHow Gottschalk-Gleser content analysis scores affect in speech, clause by clause—the seven depression subscales, the weighted coding, and its NLP legacy.
Inter-rater reliability done honestly — kappa, alpha, ICC, and the study design behind them.
Reliability & AgreementWhy a bare agreement number misleads, and how to build a confidence interval for kappa, alpha, and the ICC. See analytic vs bootstrap methods.
Reliability & AgreementRaw percent agreement isn't useless. When to chance-correct, when kappa misleads, and when agreement plus a confidence interval is the honest report.
Reliability & AgreementCohen's kappa needs a fixed item set that span and NER annotation never has. Measure span-level agreement with pairwise F1 and IoU instead—see how.
Reliability & AgreementWhich ICC to use for continuous ratings: choose the model (1, 2, or 3), single vs average measures, and absolute agreement vs consistency, from one table.
Reliability & AgreementWhy Cohen's kappa collapses under skewed prevalence, and how Gwet's AC1 and AC2 chance-correction fixes it—with the formula and a worked example.
Reliability & AgreementHow to compute Krippendorff's alpha: the observed-over-expected disagreement formula for nominal, ordinal, and missing data, with a worked example.
Reliability & AgreementSet the sample size for inter-rater reliability: how many subjects and raters pin kappa or an ICC to a target CI width. See the planning tables.
Reliability & AgreementA fill-in-the-blanks methods template for how to report inter-rater reliability: coefficient, raters, unit of analysis, CI, and software, per GRRAS.
Reliability & AgreementWhich inter-rater reliability coefficient to use, by data type, rater count, and missing data—a decision guide to kappa, alpha, AC1, and ICC.
Reliability & AgreementHow to compute Cohen's kappa and weighted kappa by hand, read the result honestly, and know when to use Krippendorff's alpha, ICC, or Gwet's AC1 instead.
Reliability & AgreementField notes on annotation—inter-rater reliability, clinical rating scales, qualitative coding, and running labeling projects that actually hold up.
Annotation and qualitative-analysis platforms compared, without the marketing gloss.
Tools & ComparisonsAn RLHF data annotation playbook: pairwise preference vs Likert rubrics, reward-model data, eval sets, and rater agreement. See the decision guide.
Tools & ComparisonsCompare RLHF annotation tools by task—SFT demos, preference ranking, red-teaming, eval rubrics—and find which tool fits each. See the matrix.
Tools & ComparisonsCompare the best data annotation tools of 2026 by open-source status, pricing, modalities, and workforce—one master table, honest picks by use case.
Tools & ComparisonsWhich clinical transcript annotation tool fits psychiatric, therapy, or SLP research? Compare options on HIPAA, speaker roles, severity scales, and IRR.
Tools & ComparisonsCVAT vs Labelbox vs V7 for computer-vision annotation: open-source vs enterprise, auto-labeling, DICOM, and pricing models, compared. See which fits.
Tools & ComparisonsLabel Studio vs Prodigy vs doccano for text annotation: open-source status, active learning, team QA, and pricing model, compared. See which one fits.
Tools & ComparisonsLooking for an NVivo alternative? Compare the best free, modern, and AI-guided qualitative coding tools by price, platform, and clinical fit.
Tools & ComparisonsOpen-source vs commercial annotation tools: the real trade-offs in total cost, support, compliance, and lock-in, plus a decision framework for your team.
Tools & ComparisonsA fair, four-way qualitative data analysis software comparison of NVivo, ATLAS.ti, MAXQDA and Dedoose: pricing, AI coding and built-in IRR, side by side.
Try a different term — a scale, a method, or a tool name.
Turn these field notes into practice. Annotate transcripts against validated scales, with an agent that learns your codebook and keeps raters in agreement.
Try Tagaroo free