Spoken vs. Written Frequency: Why Textbook Vocabulary Feels Useless in Conversation
8/16/2026
You studied vocabulary lists, read some articles, maybe worked through a textbook chapter by chapter. Then someone actually talks to you at normal speed and none of it seems to help. That gap is real, and it’s not just about needing more listening practice. Spoken and written language draw on different working vocabularies, and the difference is measurable, not just a feeling.
Speech reuses a smaller set of words, over and over
Corpus linguists measure how varied a stretch of language’s vocabulary is with something called lexical diversity, roughly, how many different words show up relative to how much total language you’re looking at. Research comparing spoken and written language consistently finds that speech has lower lexical diversity than writing: written language, given time to plan, revise, and reach for a more specific word, draws on a noticeably wider vocabulary. Speech, produced in real time without that luxury, leans much more heavily on repeating the same core set of common words.
That’s a different phenomenon from comprehension thresholds, it’s about production. When people talk, they don’t reach as often for the less common word sitting one step away from the frequent one; they default to whatever’s fastest and most automatic, which is disproportionately the highest-frequency vocabulary. The practical effect, holding across the languages this research covers: conversation leans harder on its most frequent words than written text does.
Why a lot of vocabulary study skews written anyway
Written text is easier to collect at scale than spoken conversation. News articles, books, and web pages already exist as text; casual conversation has to be recorded and painstakingly transcribed before it’s usable for anything, the same corpus-availability gap that makes balanced spoken corpora harder to build in the first place. A lot of vocabulary material, from textbooks to casual “top words” lists, ends up skewing toward whatever text was easiest to gather, which tilts toward written and formal registers.
That mismatch compounds the lexical-diversity gap instead of correcting for it. If the words you drilled lean written and speech itself concentrates even harder on a smaller core vocabulary, the vocabulary you studied and the vocabulary someone actually uses at conversational speed can end up further apart than either fact alone would suggest.
What this means in practice
None of this means studying vocabulary is pointless, it means the source of the frequency data matters more than it might seem. Word Quest 1000’s ranking comes from Mark Davies’s corpus, which was deliberately built to include a real spoken-conversation component alongside fiction and non-fiction, not assembled purely from written text. The headline “88% of spoken Spanish” figure specifically measures the spoken register, not writing standing in for it.
That doesn’t erase the underlying gap between reading fluency and following live conversation, no word list closes that entirely. But it does mean the first 1,000 words aren’t 1,000 words that happened to be common in books. They’re ranked against how people actually talk, which is the vocabulary that was falling short in the first place.
Word Quest 1000 ranks its words against real spoken Spanish, not text alone. Check it out on Amazon (affiliate link).