Zipf's Law: The Math Behind Why Frequency Lists Work

8/8/2026

If you read the last post, you already know the headline number: the first 1,000 Spanish words cover about 88% of what people say in casual conversation, according to Mark Davies’s 2005 corpus study. What that post didn’t get into is why that’s even possible. Why would nearly all of a language’s daily use come from such a small slice of its vocabulary? There’s a name for the pattern behind it: Zipf’s law.

A word’s rank predicts its frequency

Zipf’s law says that in any body of natural language text, a word’s frequency is roughly inversely proportional to its rank. The most common word shows up about twice as often as the second most common, three times as often as the third, and so on down the list. Multiply a word’s rank by its frequency and you get close to the same number no matter where you are on the list, at least for the first few hundred words.

What that looks like in real text

The classic demonstration comes from the Brown Corpus, a one-million-word sample of American English text assembled in the 1960s and still used as a benchmark today. In it, “the” is the single most common word, accounting for close to 7% of all words in the corpus. “Of,” the second most common word, shows up at close to 3.5%, half of “the”’s share, right where rank 2 predicts.

Bar chart: "the" appears in about 7% of Brown Corpus text at rank 1; "of" appears in about 3.5% at rank 2, roughly half as often. ~7% "the" (rank 1) ~3.5% "of" (rank 2)

That’s the shape Zipf’s law describes: not a gentle decline, but one word towering over everything else, a second word at roughly half its share, and a long, thinning tail after that.

Where the name comes from

Linguist George Zipf described this pattern in 1932, and it now carries his name, but he wasn’t the first to notice it. Zipf himself credited a French stenographer, Jean-Baptiste Estoup, who documented the same rank-frequency relationship in French text back in 1916. The pattern shows up in essentially every natural language that’s been checked, and in plenty of things that aren’t language at all: city population sizes, income distributions, and website traffic all follow similar rank-based curves.

Why this is the actual mechanism behind “88%”

This is what’s really underneath Davies’s numbers. Because word usage concentrates this heavily at the top of the frequency list, ranking a vocabulary by how often it’s actually used front-loads nearly all its practical value into the first few hundred entries. That’s not a property of Spanish specifically, it’s Zipf’s law showing up in the data. It also explains why Davies found the second 1,000 words added only 4.9 percentage points of spoken coverage despite doubling the list size: by then you’re past the steep part of the curve and into the thinning tail.

What this doesn’t mean

Zipf’s law describes the shape of the frequency curve. It doesn’t tell you which specific 1,000 words matter for a given learner, and coverage still isn’t comprehension, a distinction the last post covered in more detail. What it does establish is that ranking a word list by frequency isn’t an arbitrary way to organize a dictionary, it’s tracking a real, measurable pattern in how the language gets used.

Word Quest 1000 ranks all 1,000 of its words by real frequency, the same kind of measurement behind the Brown Corpus numbers above, so you spend study time on the words actually doing the work. Check it out on Amazon (affiliate link).