tagaroo

phenomena

Ekman's Basic Emotions: The Six, and How to Label Them

Ekman's basic emotions are six universal states: anger, disgust, fear, happiness, sadness, surprise. See why they anchor most emotion datasets.

Enrique Gutiérrez12 min readUpdated July 2026
Six minimal geometric face-shapes in a row, illustrating Ekman's six basic emotions as a label set.

Ekman’s basic emotions are six states he argued are recognized across every human culture from distinct facial expressions: anger, disgust, fear, happiness, sadness, and surprise (Ekman, 1992). That list is more than a psychology-class staple. It quietly became the default label set for emotion data in machine learning—when Google built the GoEmotions dataset in 2020, its authors noted that “the vast majority of existing datasets contain annotations for minor variations of the 6 basic emotion categories … proposed by Ekman” (Demszky et al., 2020). If you are labeling emotion in text, you are probably going to meet these six first.

What are Ekman’s basic emotions?

Ekman’s basic emotions are a set of six discrete affective states—anger, disgust, fear, happiness, sadness, and surprise—each proposed to have a universal, biologically based facial signature (Ekman, 1992). The claim is not just that these feelings exist everywhere, but that each is a distinct “family” with its own expression, physiology, and typical triggers.

EmotionTypical triggerNLP label convention
AngerA goal blocked; a norm violatedanger
DisgustSomething contaminating or offensivedisgust
FearA threat of harmfear
HappinessAchievement, connection, pleasurejoy
SadnessLosssadness
SurpriseSomething sudden and unexpectedsurprise
Ekman's six basic emotions, their prototypical triggers, and the label each maps to in emotion datasets (positive is almost always rendered 'joy'). Definitions on the Tagaroo Ekman scale page.

Two caveats belong right next to the list. First, terminology drifts: Ekman used “happiness,” “enjoyment,” and “joy” interchangeably, and NLP corpora almost always settle on joy.

Second, the six are not the whole of Ekman’s later thinking. He singled out contempt as the strongest candidate for a seventh universal expression, and by 1999 had proposed a broader set of around fifteen states (including amusement, pride, relief, and shame) that lack distinctive facial muscles. The strict, facially encoded core is six; the fuller affective set is larger. You can see the operational definitions on the Ekman’s Basic Emotions scale page.

Where did the six come from?

Ekman’s six come from cross-cultural studies of facial expression run in the late 1960s and early 1970s by Ekman and Wallace Friesen, most famously with the preliterate Fore people of Papua New Guinea (Ekman & Friesen, 1971). Because the Fore had minimal exposure to Western media, agreement on which face matched which emotional story was strong evidence that the expressions were not simply learned from culture.

The headline result was striking: about 92% of Fore choices for a “happy” story picked the happiness face (Ekman & Friesen, 1971). Fear and surprise were the main confusion, a wrinkle Ekman reported honestly at the time. Later work put the broader pattern on firmer statistical footing—a meta-analysis found emotions were recognized across cultures at a mean accuracy near 58%, far above the roughly 17% expected by chance in a six-way choice (Elfenbein & Ambady, 2002).

That same meta-analysis is also where the picture gets more honest. Recognition is above chance everywhere, but people read their own cultural group’s faces more accurately—an in-group advantage that a strong universality claim has to accommodate (Elfenbein & Ambady, 2002). Critics go further, arguing the classic forced-choice method (pick one of six labels) inflates agreement, and that free-labeling drops it. The defensible summary: the six are broadly recognizable across cultures, but “universal” is a contested word, not a closed case.

Why did the six become the default label set in NLP?

Ekman’s six became the default emotion taxonomy in natural language processing because they are compact, familiar, and balanced enough that annotators can apply them and models can learn them. In the first systematic survey of emotion-annotated text corpora, “9 [of] 14 are annotated with the six fundamental emotions defined by Ekman, with small variations” (Bostan & Klinger, 2018).

The through-line is easy to trace. The first emotion-recognition benchmark, SemEval-2007 Task 14, labeled news headlines on exactly the six Ekman emotions plus valence (Strapparava & Mihalcea, 2007). Google’s GoEmotions (2020) uses 27 fine-grained categories but ships an explicit “Ekman mapping” that collapses them back to the six plus neutral—making it the standard bridge between fine-grained and basic-emotion labels (Demszky et al., 2020). Even schemes that diverge define themselves against Ekman: the NRC Emotion Lexicon uses Plutchik’s eight precisely because they are a superset of Ekman’s six (Mohammad & Turney, 2013), and ISEAR keeps five of the six but swaps surprise for shame and guilt (Scherer & Wallbott, 1994).

Not every dataset is categorical, and that is worth knowing. EmoBank labels sentences on the dimensional valence-arousal-dominance model rather than Ekman’s categories, a reminder that “emotion labels” can mean a discrete set or a continuous space (Buechel & Hahn, 2017). But if you inherit an emotion dataset without asking, the odds are it speaks Ekman.

Ekman vs Plutchik: how many emotions should you label?

Ekman gives you six discrete categories; Plutchik gives you eight primaries with intensities and blends—so the choice is resolution against reliability. Ekman’s six are a strict subset of Plutchik’s eight, which add trust and anticipation and wrap the set in a structure of opposites and dyads (Ekman, 1992; Plutchik, 1980).

EkmanPlutchik
Categories6 basic emotions8 primaries (+ intensities, dyads)
BasisUniversal facial expressionsPsychoevolutionary function
Adds vs the otherTrust, anticipation
GranularityCoarserFiner
Inter-annotator agreementHigher (fewer boundaries)Lower (more boundaries)
Best whenYou need coders to agreeYou need mixed/graded emotion
Ekman vs Plutchik as label sets. Fewer categories are easier to agree on; more categories capture more, at a cost in reliability. Full head-to-head in the Plutchik guide.

The reliability cost of extra categories is real and measurable. Even within the six-label scheme, agreement varies sharply by emotion: in SemEval-2007, inter-annotator correlations ran from sadness at 0.68 down to surprise at 0.36 (Strapparava & Mihalcea, 2007). Add categories and you add boundaries for coders to disagree about. Whichever taxonomy you choose, quantify the agreement it produces with a coefficient like Cohen’s kappa rather than raw percent agreement, and if you need the richer scheme, the Plutchik’s wheel of emotions guide walks through its intensities and dyads.

Is basic emotion theory settled?

No—basic emotion theory is influential but actively contested, and using the six well means being candid about that. The leading challenge is Lisa Feldman Barrett’s theory of constructed emotion, which argues emotions are not fixed biological fingerprints but are built on the fly from core affect and learned concepts (Barrett, 2017). On that view there is no dedicated “anger circuit” waiting to fire; anger is constructed in context.

The empirical picture has also outgrown a list of six. Cowen and Keltner had more than 800 people rate 2,185 evocative videos and recovered 27 distinct categories of reported emotion, concluding that “the boundaries between categories of emotion are fuzzy rather than discrete” (Cowen & Keltner, 2017). Fuzzy boundaries are exactly why fine-grained labeling is hard: the line between fear and horror, or awe and admiration, is a gradient, not a fence.

So why still use six? Because for labeling work, coarser often wins. The same GoEmotions model scored an average F1 of 0.46 over the 27-way taxonomy but 0.64 when the labels were collapsed to Ekman’s six (Demszky et al., 2020)—fewer categories, more agreement, better performance.

The six are a defensible sweet spot: enough resolution to be useful, coarse enough to annotate reliably. Their known weakness is a single positive category (joy), which is precisely the gap richer schemes were built to fill.

How to label the six emotions in a transcript

To apply Ekman’s six to a transcript, tag the span where the speaker expresses one of the emotions in their own words and assign the label—rather than rating a whole passage with one global impression. The Ekman scheme is instance-mode: you mark the utterance that is the emotion, not a summary score for the interview.

Consider a short synthetic exchange:

Interviewer: How did you feel walking back into the office that first day?

Subject: My stomach turned. I honestly couldn’t stand the thought of seeing them again.

That reply supports tagging disgust on “my stomach turned … couldn’t stand the thought,” with the span attached as evidence. Coding the span rather than labeling the turn “negative” keeps the justification next to the label, so a second coder can check it—and only the speaker’s words are coded, never the emotion words in the interviewer’s question. That discipline is what makes the Ekman emotion taxonomy reproducible enough to compute agreement across coders.

On data handling: emotion labeling means working with personal, sometimes sensitive language, so the sane default is de-identified data and a privacy-first setup. Tagaroo supports a browser-side anonymous mode, so content can stay local rather than being uploaded—worth checking against your ethics approval before real data touches any tool. If you need finer resolution than six categories, the same evidence-first workflow carries over to Plutchik’s wheel of emotions, and the older lineage of coding affect from language is covered in Gottschalk-Gleser content analysis.

Common mistakes when labeling with the six

The most common mistake is treating the six as a complete, settled account of emotion rather than a compact labeling scheme, but a few others recur:

  • Forcing everything into six boxes. Real text carries contempt, pride, and relief that don’t fit cleanly. Add a neutral/other category and decide in advance where the borderline cases go, so coders don’t improvise.
  • Coding the interviewer’s emotion words. “Did that frighten you?” is not the subject expressing fear. Tag what the speaker feels, not what the prompt names.
  • Assuming the labels are equally easy. Surprise and disgust annotate far less reliably than sadness and fear (Strapparava & Mihalcea, 2007). Expect lower agreement on the hard categories and calibrate on them specifically.
  • Confusing categorical and dimensional labels. Ekman’s six are discrete categories; valence-arousal is a continuous space. Don’t average category labels as if they were numbers.

None of these are reasons to avoid the six—they are compact, well-known, and reliable enough to be the sensible default. They are reasons to use them with their limits in view: name the scheme, plan for the leftovers, and measure the agreement you actually get.

The practical upshot: Ekman’s basic emotions are the reference taxonomy for affect labeling because they hit a sweet spot between resolution and reliability, not because the science is closed. Use the six as a practical scheme, ground every label in the words that justify it, and reach for a richer taxonomy only when your coders can sustain the extra categories.

References

  • Ekman, P. (1992). An argument for basic emotions. Cognition and Emotion, 6(3–4), 169–200. doi:10.1080/02699939208411068
  • Ekman, P., & Friesen, W. V. (1971). Constants across cultures in the face and emotion. Journal of Personality and Social Psychology, 17(2), 124–129. doi:10.1037/h0030377
  • Elfenbein, H. A., & Ambady, N. (2002). On the universality and cultural specificity of emotion recognition: a meta-analysis. Psychological Bulletin, 128(2), 203–235. doi:10.1037/0033-2909.128.2.203
  • Barrett, L. F. (2017). The theory of constructed emotion: an active inference account of interoception and categorization. Social Cognitive and Affective Neuroscience, 12(1), 1–23. doi:10.1093/scan/nsw154
  • Cowen, A. S., & Keltner, D. (2017). Self-report captures 27 distinct categories of emotion bridged by continuous gradients. Proceedings of the National Academy of Sciences, 114(38), E7900–E7909. doi:10.1073/pnas.1702247114
  • Strapparava, C., & Mihalcea, R. (2007). SemEval-2007 Task 14: affective text. Proceedings of the 4th International Workshop on Semantic Evaluations, 70–74. aclanthology.org
  • Demszky, D., Movshovitz-Attias, D., Ko, J., Cowen, A., Nemade, G., & Ravi, S. (2020). GoEmotions: a dataset of fine-grained emotions. Proceedings of the 58th Annual Meeting of the ACL, 4040–4054. aclanthology.org
  • Bostan, L.-A.-M., & Klinger, R. (2018). An analysis of annotated corpora for emotion classification in text. Proceedings of COLING 2018, 2104–2119. aclanthology.org
  • Plutchik, R. (1980). A general psychoevolutionary theory of emotion. In R. Plutchik & H. Kellerman (Eds.), Emotion: Theory, Research, and Experience, Vol. 1 (pp. 3–33). Academic Press. doi:10.1016/B978-0-12-558701-3.50007-7
  • Mohammad, S. M., & Turney, P. D. (2013). Crowdsourcing a word–emotion association lexicon. Computational Intelligence, 29(3), 436–465. doi:10.1111/j.1467-8640.2012.00460.x
  • Scherer, K. R., & Wallbott, H. G. (1994). Evidence for universality and cultural variation of differential emotion response patterning. Journal of Personality and Social Psychology, 66(2), 310–328. doi:10.1037/0022-3514.66.2.310
  • Buechel, S., & Hahn, U. (2017). EmoBank: studying the impact of annotation perspective and representation format on dimensional emotion analysis. Proceedings of EACL 2017, 578–585. aclanthology.org

If you build emotion-labeled datasets from interviews or free text, Tagaroo turns Ekman’s basic emotions into a guided, evidence-anchored annotation workflow—with inter-rater reliability computed as your coders work.

Frequently asked questions

What are the six basic emotions according to Ekman?
Paul Ekman's six basic emotions are anger, disgust, fear, happiness, sadness, and surprise (Ekman, 1992). He argued these are recognized across cultures from distinct facial expressions. Ekman later described contempt as having strong evidence for a seventh universal expression, and in the 1990s proposed a broader list of around fifteen emotional states.
Are Ekman's basic emotions universal?
Ekman's cross-cultural studies, including work with the preliterate Fore of Papua New Guinea, found emotion faces are recognized well above chance everywhere (Ekman & Friesen, 1971). A meta-analysis put mean cross-cultural recognition near 58% against roughly 17% chance, but also found a reliable in-group advantage (Elfenbein & Ambady, 2002). Universality is well-supported but not uncontested—critics note forced-choice methods inflate agreement.
What is the difference between Ekman and Plutchik?
Ekman proposed six discrete basic emotions grounded in facial expressions; Plutchik proposed eight primaries on a psychoevolutionary basis, adding trust and anticipation plus intensities and blends (Ekman, 1992; Plutchik, 1980). Ekman's six are a subset of Plutchik's eight. Ekman's set is coarser and easier to agree on; Plutchik's is finer and represents mixed emotion.
Why do NLP emotion datasets use Ekman's six emotions?
The six give a compact, balanced-enough label set that annotators can apply reliably and models can learn. A survey of emotion-annotated text corpora found 9 of 14 used Ekman's six with minor variations (Bostan & Klinger, 2018). Benchmarks from SemEval-2007 to Google's GoEmotions (which ships an explicit Ekman mapping) use the six as a standard grouping.
Is basic emotion theory still accepted?
Basic emotion theory remains influential but is actively debated. Barrett's theory of constructed emotion argues emotions are built from core affect and learned concepts rather than fixed biological fingerprints (Barrett, 2017). Data-driven work found 27 emotion categories bridged by continuous gradients rather than a small discrete set (Cowen & Keltner, 2017). The six persist as a practical labeling scheme, not a settled account of how emotion works.

Put this into practice

Tagaroo turns any rating scale or coding scheme into a guided annotation workflow — with inter-rater reliability computed as you go.