phenomena
Ekman vs Plutchik: Which Emotion Model to Label With?
Ekman vs Plutchik for annotation: 6 discrete emotions or 8 with intensities and dyads, and why more categories lower agreement. See which to pick.

If you are choosing an emotion taxonomy for a labeling scheme, the Ekman vs Plutchik question usually gets framed as “six emotions or eight?” That framing misses the decision that actually matters. Ekman gives you six discrete categories that annotators agree on and that map cleanly to sentiment. Plutchik gives you eight primaries with intensities and blends—far more nuance, but measurably harder to apply consistently. The real trade-off is granularity against inter-annotator agreement, and it decides the quality of every label you collect after it.
Ekman vs Plutchik at a glance
Ekman and Plutchik answer different questions. Ekman asks “which of a few universal states is this?”; Plutchik asks “where does this sit in a structured space of primaries, intensities, and blends?” The table below lines them up on the axes that decide a labeling scheme.
| Axis | Ekman (basic emotions) | Plutchik (wheel) |
|---|---|---|
| Model type | Discrete / categorical | Structured: categories + geometry (opposites, intensity) |
| Categories | 6 (anger, disgust, fear, happiness, sadness, surprise); contempt sometimes 7th | 8 primaries in 4 opposing pairs |
| Intensity | Not built in | 3 tiers per primary (e.g. annoyance → anger → rage) |
| Mixed emotions | Not modeled | Dyads (e.g. joy + trust = love) |
| Native modality | Faces / vision (FACS lineage) | General; widely adapted for text |
| Agreement tendency | Higher (fewer categories) | Lower (more categories + intensities) |
| Best for | Coarse, high-agreement labels; cross-dataset compatibility | Fine-grained nuance with trained annotators + adjudication |
Read down the “agreement tendency” row before anything else. It is the axis most teams underweight, and it is the one that quietly determines whether your labels are usable.
What is Ekman’s model?
Ekman’s basic emotions are six states he argued are recognized across cultures from distinct facial signals: anger, disgust, fear, happiness, sadness, and surprise (Ekman, 1992). The empirical root is forced-choice facial-recognition work in literate and preliterate cultures, including the Fore of Papua New Guinea, where observers picked the predicted emotion above chance (Ekman, Sorenson & Friesen, 1969). Ekman later described contempt as a candidate seventh universal expression (Ekman & Friesen, 1986). It is a flat, discrete taxonomy: a small set of qualitatively distinct kinds, with no built-in notion of intensity or blending.
The full label set and coded examples live on the Ekman’s six basic emotions explainer and the Ekman scale page.
What is Plutchik’s wheel?
Plutchik’s wheel organizes affect into eight primary emotions arranged as four opposing pairs: joy↔sadness, trust↔disgust, fear↔anger, and surprise↔anticipation (Plutchik, 2001). Each primary comes in three intensities (serenity → joy → ecstasy), and adjacent primaries combine into named dyads such as joy plus trust equals love. The model is psychoevolutionary: emotions are adaptive reactions conserved across species, and the wheel encodes their similarity (adjacency), polarity (opposites), and intensity (radius). Unlike Ekman’s flat list, Plutchik supplies an explicit geometry, which is why it is partly dimensional.
The eight primaries, intensity rings, and dyads are worked through on the Plutchik’s wheel of emotions explainer and the Plutchik scale page.
The trade-off that decides it: granularity vs agreement
The single most important fact for choosing between these models is that inter-annotator agreement falls as you add categories, and it falls unevenly. More labels mean more boundary decisions, and subtle or social emotions are genuinely ambiguous in short text. Plutchik’s richer scheme buys nuance at a direct cost in reliability that you pay on every item.
The clearest evidence comes from GoEmotions, a 58,000-comment dataset labeled with 27 emotions plus neutral. Even with three to five raters per example, only 31% of examples had three or more raters agree on at least one label (Demszky et al., 2020). Agreement did not degrade uniformly; it varied enormously by emotion. Annotators lined up well on lexically cued states like gratitude and amusement, and barely at all on subtle ones like grief and nervousness.
The chart shows per-emotion interrater correlation in GoEmotions: raters agree strongly on emotions with explicit lexical cues and weakly on subtle ones (Demszky et al., 2020). The lesson is not “avoid fine-grained schemes.” It is that the hard categories are hard for everyone, so a scheme with more of them concentrates your disagreement in exactly the labels you probably care about.
How far can this go? One fine-grained Arabic emotion dataset asked annotators to choose among 37 emotions and saw Cohen’s kappa collapse to roughly 0.008; the team abandoned the 37-way scheme and reverted to seven emotions (ArPanEmo, 2023). That is the granularity-agreement curve at its extreme.
To interpret any agreement number, anchor it to the Landis and Koch (1977) bands: kappa of 0.00–0.20 is slight, 0.21–0.40 fair, 0.41–0.60 moderate, 0.61–0.80 substantial, and 0.81–1.00 almost perfect. Most fine-grained emotion corpora sit in the slight-to-fair range; coarse, Ekman-style sets can reach moderate to substantial. Those bands are widely used but arbitrary, so report them as a guide, not a verdict. If reliability is the whole game for your project, our guide to Cohen’s kappa and inter-rater reliability covers why a bare coefficient without base rates and a confidence interval is close to meaningless.
Which model fits which modality?
Match the emotion model to the signal you are labeling. Ekman’s six are rooted in the face: the Facial Action Coding System and the posed-photo recognition studies are vision work, and the six-way scheme dominates facial-expression datasets. If you are labeling faces or building a multimodal dataset that must interoperate with existing ones, Ekman-6 is the compatible default.
Text is a different signal, and pure Ekman-6 is rarely what text teams actually use. They tend either to coarsen to sentiment or to expand toward Plutchik, because written emotion is frequently mixed and valenced in ways six discrete kinds cannot hold. SemEval-2018’s affect-in-tweets task is a good tell: its multi-label emotion subtask used eleven labels—Plutchik’s eight primaries plus love, optimism, and pessimism (Mohammad et al., 2018). The dyad idea, in other words, earns its keep in text even when the full wheel does not.
Is basic-emotion theory even settled?
No, and a labeling team should know it. The premise that six emotions are universal natural kinds you can read off a face is seriously contested. The most systematic review to date spanned cultures, newborns, and congenitally blind individuals (Barrett et al., 2019). It concluded that people do smile when happy and scowl when angry more often than chance, but with “substantial variation” across cultures and situations, and that a given facial configuration often communicates something other than an emotional state.
The constructionist account goes further: emotions are not fixed biological types triggered by dedicated circuits but are constructed from bodily sensation plus learned concepts (Barrett, 2017). For annotation, this reframes low agreement as expected rather than as annotator failure. If category boundaries are partly cultural and linguistic conventions, then two careful raters can legitimately disagree, and your codebook is defining a convention rather than discovering a fact. That is a reason to document your taxonomy precisely and, sometimes, to abandon categories altogether.
The third option most teams miss
Often the right move in the Ekman vs Plutchik debate is to pick neither and use a dimensional or collapsible scheme instead. A dimensional model rates each item on continuous axes rather than sorting it into a bin. Russell’s circumplex places emotion on two axes, valence (pleasant-unpleasant) and arousal (activation), which sidesteps category boundaries and yields regression targets (Russell, 1980). The valence-arousal-dominance extension adds a third axis (Mehrabian, 1996).
The middle path is a collapsible fine-grained set. Cowen and Keltner (2017) found self-reported emotion is best captured by 27 categories bridged by continuous gradients—richer than valence-arousal, yet not strictly discrete. That taxonomy seeded GoEmotions, which ships built-in hierarchical groupings so you annotate once and report at whichever level clears your agreement bar (Demszky et al., 2020). EmoBank shows the two worlds can coexist, pairing valence-arousal-dominance ratings with a categorical subset for interoperability (Buechel & Hahn, 2017).
| Level | Labels | Agreement |
|---|---|---|
| Fine-grained (GoEmotions) | 27 emotions + neutral | Lowest |
| Ekman grouping | anger, disgust, fear, joy, sadness, surprise (+ neutral) | Higher |
| Sentiment | positive, negative, ambiguous, neutral | Highest |
One 2026 wrinkle changes this calculus. LLM-assisted pre-labeling applies a fine-grained scheme consistently, if not always correctly, so the practical bottleneck shifts from annotator disagreement toward auditing the model’s systematic errors. That makes a collapsible taxonomy even more useful: you can let a model draft at 27 categories, verify the ones it is unsure of, and still report at the Ekman or sentiment level where humans and the model actually agree.
Ekman vs Plutchik: so which should you use?
Choose the model that matches your modality, your annotators, and the agreement you can afford. The evidence points to three clean rules.
- Choose Ekman’s six (or seven with contempt) when you are labeling faces or vision, you need coarse and high-agreement labels that map to sentiment, your annotators are non-expert or crowdsourced, or cross-dataset compatibility matters. The cost is that it misses blends, mixed valence, and social emotions.
- Choose Plutchik’s eight (with intensities and dyads) when you need structured nuance in text, you have trained annotators and an adjudication step, and you value the built-in geometry for downstream modeling. The cost is measurably lower inter-annotator agreement, and intensity tiers that annotators rarely apply reliably.
- Choose a dimensional or collapsible scheme when category agreement is stuck in the slight-to-fair range, or when your emotions are subtle, mixed, or continuous. Russell’s valence-arousal or a 27-to-Ekman-to-sentiment hierarchy will usually serve you better than forcing a choice between six and eight.
The one thing not to do is pick a category count by taste and discover your agreement problem after 10,000 labels are in.
How do you label emotion from a transcript?
To label emotion from an interview, tag the span where the subject expresses the emotion in their own words, assign the category, and record the intensity if your scheme has one. Both Ekman and Plutchik are instance-mode schemes in this setting: you mark the utterance that is the emotion, not a global impression of the speaker. Emotion words that appear only in the interviewer’s question are not coded, and neutral or purely informational speech is left untagged.
Consider a short synthetic exchange:
Interviewer: How did you feel when you heard the news?
Subject: At first I just froze, and then honestly I was furious that nobody had warned me.
Under Ekman, “I just froze” supports fear and “furious” supports anger, giving two discrete tags on one turn. Under Plutchik, the same span invites fear and anger primaries and, if you code intensity, “furious” reads as rage rather than annoyance. Tagging the exact evidence span, rather than rating the turn from memory, is what lets a reviewer check each label and lets you compute inter-annotator agreement per category as coders work. That surfaces the low-agreement emotions early, while your codebook can still be fixed.
If you are building a labeled emotion dataset for a model rather than coding clinical interviews, the same span-and-evidence discipline applies, and the tooling question is separate from the taxonomy question—see our comparison of Label Studio, Prodigy, and doccano for the text-annotation side. For the historical roots of reading affect out of language, the Gottschalk-Gleser content analysis method is the clause-level ancestor of today’s emotion labeling.
The practical upshot of the Ekman vs Plutchik choice: it is not really six versus eight. It is how much agreement you are willing to trade for nuance, whether your signal is faces or text, and whether a dimensional or collapsible scheme would serve better than either flat taxonomy. Pick the representation your data and your annotators can actually support, and design it so you can always collapse down rather than wishing you had labeled up.
References
- Ekman, P. (1992). An argument for basic emotions. Cognition & Emotion, 6(3–4), 169–200. doi:10.1080/02699939208411068
- Ekman, P., Sorenson, E. R., & Friesen, W. V. (1969). Pan-cultural elements in facial displays of emotion. Science, 164(3875), 86–88. doi:10.1126/science.164.3875.86
- Ekman, P., & Friesen, W. V. (1986). A new pan-cultural facial expression of emotion. Motivation and Emotion, 10(2), 159–168. doi:10.1007/BF00992253
- Plutchik, R. (2001). The nature of emotions. American Scientist, 89(4), 344–350. Publisher page
- Russell, J. A. (1980). A circumplex model of affect. Journal of Personality and Social Psychology, 39(6), 1161–1178. doi:10.1037/h0077714
- Mehrabian, A. (1996). Pleasure-arousal-dominance: A general framework for describing and measuring individual differences in temperament. Current Psychology, 14(4), 261–292. doi:10.1007/BF02686918
- Cowen, A. S., & Keltner, D. (2017). Self-report captures 27 distinct categories of emotion bridged by continuous gradients. PNAS, 114(38), E7900–E7909. doi:10.1073/pnas.1702247114
- Demszky, D., Movshovitz-Attias, D., Ko, J., Cowen, A., Nemade, G., & Ravi, S. (2020). GoEmotions: A dataset of fine-grained emotions. ACL 2020, 4040–4054. doi:10.18653/v1/2020.acl-main.372
- Mohammad, S. M., Bravo-Marquez, F., Salameh, M., & Kiritchenko, S. (2018). SemEval-2018 Task 1: Affect in Tweets. Proceedings of SemEval-2018, 1–17. ACL Anthology
- Buechel, S., & Hahn, U. (2017). EmoBank: Studying the impact of annotation perspective and representation format on dimensional emotion analysis. EACL 2017, 578–585. doi:10.18653/v1/E17-2092
- Barrett, L. F. (2017). The theory of constructed emotion. Social Cognitive and Affective Neuroscience, 12(1), 1–23. doi:10.1093/scan/nsw154
- Barrett, L. F., Adolphs, R., Marsella, S., Martinez, A. M., & Pollak, S. D. (2019). Emotional expressions reconsidered. Psychological Science in the Public Interest, 20(1), 1–68. doi:10.1177/1529100619832930
- Landis, J. R., & Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1), 159–174. doi:10.2307/2529310
If you are choosing a taxonomy for an emotion dataset, Tagaroo lets you annotate against Ekman’s six or Plutchik’s eight with span-level evidence and inter-rater reliability computed as your coders work—so you catch the low-agreement categories before they cost you a relabel.
Frequently asked questions
- What is the difference between Ekman and Plutchik?
- Ekman's model is a flat list of six discrete basic emotions—anger, disgust, fear, happiness, sadness, and surprise (Ekman, 1992). Plutchik's wheel is a structured model of eight primary emotions in four opposing pairs, each with three intensity levels, that blend into named dyads such as joy plus trust equals love (Plutchik, 2001). Ekman gives you fewer, cleaner categories; Plutchik gives you more nuance and an explicit geometry of opposites and intensities.
- How many emotions should you label?
- Label as few categories as your task can tolerate, because inter-annotator agreement falls as you add categories. Ekman's six reach higher agreement and map cleanly to sentiment; Plutchik's eight-plus-intensities capture nuance but are harder to apply consistently. A practical pattern is to annotate a fine-grained set but keep an explicit path to collapse it to a coarser one, as GoEmotions does with its 27-to-Ekman-to-sentiment hierarchy (Demszky et al., 2020).
- Which emotion model is best for NLP and text?
- For text there is no single winner: Ekman's six are the most compatible across face and multimodal datasets, but text teams often adopt a Plutchik-derived set to capture mixed and valenced states. SemEval-2018's emotion task, for example, used eleven labels—Plutchik's eight primaries plus love, optimism, and pessimism (Mohammad et al., 2018). Many teams now annotate a larger set with a documented mapping down to Ekman or to valence-arousal.
- Why does inter-annotator agreement drop as you add more emotions?
- More categories mean more boundary decisions, and subtle or social emotions are genuinely ambiguous in text. In GoEmotions, only 31% of examples had three or more raters agree on at least one label, and per-emotion agreement ranged from strong on lexically cued emotions like gratitude down to weak on subtle ones like grief (Demszky et al., 2020). One 37-emotion Arabic dataset saw agreement collapse to near zero and reverted to seven categories (ArPanEmo, 2023).
- Is there a better option than Ekman or Plutchik?
- Sometimes. When categories will not agree—subtle, mixed, or continuous affect—a dimensional model such as Russell's valence-arousal circumplex (Russell, 1980) sidesteps category boundaries entirely by rating position on continuous axes. Cowen and Keltner (2017) found self-reported emotion is best captured by 27 categories bridged by continuous gradients, which is neither strictly discrete nor strictly dimensional. Pick the representation your data and annotators can actually support.
Put this into practice
Tagaroo turns any rating scale or coding scheme into a guided annotation workflow — with inter-rater reliability computed as you go.