Why 'Top 1,000 Words' Means Something Different in Every Language
8/15/2026
Word families, lemmas, and tokens covered one reason two frequency lists for the same language can disagree: they might not be counting the same unit. Cross the border into a different language entirely, and a second, deeper set of differences shows up, ones that have nothing to do with which counting convention a linguist picked, and everything to do with how that language actually works.
Some languages pack a lot more into one word
Languages differ enormously in how much grammatical information gets folded into a single word versus spread across separate words. Linguists sort this on a spectrum, from analytic (or isolating) languages, where each word tends to carry one unit of meaning and grammatical relationships are shown with separate words and word order, to agglutinative languages, where a single word can stack multiple meaningful pieces together, each one still individually identifiable.
Mandarin sits close to the analytic end: it has very little inflection, so a given root doesn’t multiply into many different surface forms the way a Spanish verb does across its conjugations. Turkish and Finnish sit at the agglutinative end, where a single root can take on a long chain of attached pieces, each one adding a distinct piece of meaning, producing far more surface forms per root than a language like Spanish or English ever generates. Spanish and English both fall somewhere in between, with Spanish carrying noticeably more verb inflection than English does.
That difference changes what a lemma-based list, the approach Word Quest 1000 itself uses, even means from one language to the next. Folding conjugations into one Spanish verb entry saves a moderate, predictable number of separate slots. Doing the same for a heavily agglutinative language would be folding in a much larger, more elaborate set of forms, since there’s simply more surface variation attached to each root to begin with.
Some languages don’t even mark where one word ends
Spanish, English, and most European languages separate words with spaces, so the first step of building a frequency list, finding the candidate words to count, is close to free. Mandarin text doesn’t work that way: it’s written as a continuous string of characters with no spaces between words at all. Before anyone can count how often a “word” appears in Mandarin, they first have to decide where the word boundaries even are, a genuinely hard problem in its own right, this process is called segmentation.
The same character string can often be split into words more than one defensible way, and different splits can carry different meanings, so segmentation isn’t a mechanical formality, it’s an active area of computational linguistics research with no single universally agreed-on solution. A Spanish or English frequency list starts counting from a text that already tells you where the words are. A Mandarin frequency list has to solve a real analytical problem before counting can start at all.
Corpora aren’t built the same way for every language
A frequency list is only as good as the text it’s counted from, and how much well-balanced text exists, spanning spoken conversation, fiction, news, and everyday writing, varies a lot by language. Some languages have large, carefully curated corpora spanning multiple registers, close to what Mark Davies built for Spanish in 2005. For others, the available text skews more heavily toward whatever’s easiest to collect at scale, formal writing and news being the most common bias, which pushes a frequency ranking toward vocabulary you’d read more than vocabulary you’d actually say out loud.
None of this is a flaw in any specific list. It just means two “1,000 most common words” lists from two different languages were never built from directly comparable raw material to begin with, on top of everything the previous post already covered about counting units within a single language.
Why this doesn’t undermine the whole idea
Zipf’s law, the pattern where a small number of words account for a large share of everyday use, shows up across languages regardless of how differently those languages package words together or how their text gets segmented. The mechanism behind a frequency list, rank vocabulary by real usage and front-load the highest-value words first, still works. What changes from language to language is the bookkeeping: what counts as one “word,” how many forms get folded together, and what kind of text the counting was done on. Those are real differences worth being upfront about, not reasons to distrust the approach itself.
If Word Quest 1000 ever expands past Spanish, this is exactly the kind of question that would need answering fresh for whatever language comes next, not assumed to carry over automatically from how Spanish was counted.
Word Quest 1000 ranks 1,000 real Spanish lemmas by real frequency, built from a corpus balanced across spoken conversation, fiction, and non-fiction. Check it out on Amazon (affiliate link).