Word Families, Lemmas, and Tokens: How Word Counting Actually Works

8/14/2026

Is “run,” “runs,” and “running” one word or three? It sounds like a pedantic question, but the answer changes what ends up on a “top 1,000 words” list, and it’s a real reason two frequency lists built from similar source material can disagree with each other more than you’d expect.

Four ways to count a word

Corpus linguists don’t actually have one definition of “word.” They have several, each one grouping things differently, and which one a frequency list uses changes the count.

A token is a single occurrence: every time a word shows up in a text, that’s one token, counted separately no matter how many times the same word repeats. A type is a distinct spelling: “habla,” “hablas,” and “hablo” are three different types, even though they’re all forms of the same verb. A lemma groups a type together with its inflected forms under one dictionary headword: “hablo,” “hablas,” “habla,” “hablamos,” “hablan,” and every other conjugation of “to speak” all collapse into the single lemma hablar, the way a dictionary lists one entry, not one per conjugation. A word family goes further still, grouping a lemma together with words derived from it too. Linguists Laurie Bauer and Paul Nation’s original 1993 framework, built specifically around English affixes, groups a base word with its regularly-formed derivatives: “encourage” pulls in “encouraged” and “encouragement” under one family, not just its own conjugated forms. The same idea, applied loosely to Spanish, would fold “hablante” (speaker) and “hablador” (talkative, chatty) in with “hablar,” though Bauer and Nation’s own affix-level criteria were built for English, so that’s an illustration of the concept carried over to Spanish, not a citable Spanish word-family count.

Four different counting units, four different vocabulary sizes for the exact same text. A frequency list built from raw token counts will rank “es” (is) extremely high, since it’s one of the most common surface forms in Spanish. A lemma-based list folds “es” into “ser” instead, and “es” never appears as its own entry at all.

What Word Quest 1000 actually counts

Word Quest 1000 counts by lemma, and you can confirm it directly from the list itself. Ser (rank 8) appears as the infinitive, “to be.” Its conjugations, “es,” “era,” “fue,” “son,” none of them show up anywhere else in the 1,000 ranked entries; they’re already folded into “ser.” (There’s a second, unrelated entry spelled “ser” further down the list, at rank 352, a noun meaning “a being,” as in “un ser humano.” That’s not a repeat, it’s a different lemma that happens to share a spelling with the verb, the same kind of distinction this whole post is about.) The same pattern holds for nouns and regular adjectives: año (rank 55) is on the list, “años” isn’t. Nuevo (rank 99) is on the list, “nueva” isn’t. Each entry is a dictionary headword, the citation form you’d look up, not a separate slot for every inflected form that word can take.

That’s not an arbitrary choice. It matches the counting unit of Word Quest 1000’s own underlying frequency source: Mark Davies’s 2005 corpus study, covered in why 1,000 words is enough, which explicitly counts lemmas rather than raw tokens or the broader word-family groupings. WQ1K’s rank list follows the same convention as the data it’s built on.

Why two “top 1,000” lists can disagree

This is also the honest mechanism behind why two different “1,000 most common Spanish words” lists can diverge even when both claim to be frequency-based. If one counts tokens, “es,” “fue,” and “son” each occupy their own slot near the top, crowding out lemmas that would otherwise make the cut. If one counts word families instead of lemmas, related derived words get folded together, which shifts rankings further down the list. Two lists can both be honestly built from real frequency data and still disagree, simply because “word” wasn’t the same unit of measurement in both.

It’s the same reason Nation’s word-family counts and Davies’s lemma counts aren’t directly comparable numbers, even when they’re describing the same underlying idea, that a small set of high-frequency words covers most of everyday language.

Why lemma counting is the right call for a learner

For a learner, lemma counting isn’t just a technical default, it’s the unit that actually matches how you use a dictionary. You look up “hablar,” not “hablamos.” Learning the lemma and the handful of regular conjugation patterns that go with it covers the token-level forms automatically; a list that burned separate slots on “hablo,” “hablas,” and “habla” would waste three of your first thousand words on one verb you’d have learned anyway. Counting by lemma keeps every entry doing genuinely new work.

Word Quest 1000 ranks 1,000 real lemmas by real frequency, not padded out with conjugations you’d learn for free. Check it out on Amazon (affiliate link).