How many words do you need to speak Spanish?
Every answer to this question quotes the same numbers, and almost none of them say what was counted. This one is built on measurements of our own Spanish frequency list β 18,348 entries, ranked β so you can at least see what the claims rest on.
The short version: the usual figures are roughly right about coverage and badly wrong about effort, because a Spanish "word" and an English "word" are not the same size of thing.
The counting problem, measured
Here is something you can check on our list and will not find on other sites. Spanish frequency lists are lists of forms, not of words. Es, son, era, fue and sea are five entries, and they are all the verb ser.
We can measure how much of the list that accounts for, because our entries carry notes identifying forms of other words. Counting only the entries explicitly marked that way:
| Rank band | Entries that are a form of another entry |
|---|---|
| 1β100 | 14% |
| 101β250 | 28% |
| 251β500 | 28% |
| 501β1,000 | 37% |
| 1,001β2,000 | 41% |
| 2,001β5,000 | 44% |
By the time you are past the first two thousand, something approaching half of what looks like new vocabulary is conjugation of verbs already on the list.
These are floors, not exact figures. The annotation is not complete β about half the list carries a note at all β so the true proportion is higher than every number in that table. Read the direction, which is unambiguous, rather than the decimal.
What that means for the headline number
The coverage figures everyone quotes are measured in word families β a base word plus its inflections and derivations, counted once. Paul Nation's 2006 study, the source most of these numbers trace back to, uses Bauer and Nation's "Level 6" definition: inflections plus more than eighty derivational affixes. On that definition the family nation has twenty-one members, from national to internationalism.
A frequency list is not built that way. It is a list of forms, one row each.
So the two numbers are not comparable, and the mistake runs in the direction that flatters you. A thousand Spanish word families is a much larger achievement than a thousand rows of a frequency list, because those thousand rows contain, by the measurement above, several hundred entries that are the same verbs wearing different endings.
Nation says as much himself, and it is the sentence this whole post is really a footnote to: "If word lists were made for productive purposes, for speaking and writing, the lemma would be the largest sensible unit to use."
For the record, his actual figures are much larger than the ones that circulate. For unassisted comprehension at 98% coverage he puts the requirement at 8,000 to 9,000 word families for written text and 6,000 to 7,000 for spoken. Measured on a novel, 2,000 families gave 88% coverage, 4,000 gave 95%, and it took 9,000 to reach 98%.
The good news is the flip side: learning ser properly retires five entries at once. Effort spent on the conjugation system pays across the whole list rather than word by word, which is why a Spanish course that front-loads verbs is not being pedantic.
The shape of the list
Two more things our list shows directly.
Common words are short. Mean length runs 3.8 characters across the first hundred entries, 5.7 across ranks 101β500, and 7.7 by the 2,001β5,000 band. Frequency and brevity track each other closely, which is a real and well-known property of language rather than anything Spanish-specific.
Long words arrive late and are mostly transparent. Nothing in the first hundred entries is twelve characters or longer; by ranks 2,001β5,000, 7.2% are. Spanish builds those long words with a small set of productive endings β -ciΓ³n, -mente, -idad, -miento β so a great many of them are decodable the moment you know the pattern and the stem. A list counts them as separate vocabulary. Your brain, after a while, does not.
That is worth noticing because Spanish actually has more twelve-character-plus entries in that band than our German list does, at 7.2% against 5.5%. The stereotype about which European language has the long words does not survive contact with the data.
So what is the number?
There isn't one, and here is the honest shape of the answer instead.
- The first hundred entries are where the density is. Function words, the two verbs ser and estar, pronouns, prepositions. They are short, they are irregular, and they are in every sentence you will ever hear.
- The first thousand is the point at which listening stops being noise. Expect to recognise most of what is said and still lose the sentence, because the words you do not have are carrying the content.
- Two to three thousand is where reading becomes possible with a dictionary rather than through one.
- Past that, the returns flatten and what matters shifts from how many words you know to how well you know the ones you have.
The caveat that matters most
Coverage is not comprehension. Knowing 95% of the words in a text does not mean understanding 95% of it, because the 5% you are missing are not randomly distributed β they are the nouns and verbs carrying the specific meaning, while the words you know are mostly grammatical scaffolding.
Nation is blunt about this too, quoting Ronald Carver and then adding his own verdict: "even 98% coverage does not make comprehension easy." In one study of non-fiction, few learners reached adequate comprehension even at that level.
This is why "a thousand words and you can have a conversation" is true and misleading at once. You can have a conversation. You will not understand the answer.
What to do about it
Learn in frequency order, which is what the Spanish list is for and what the Spanish course teaches from. Then spend disproportionate time on the verbs, because that is where the list's own structure says the leverage is.
The general version of this argument, with the research behind the coverage figures, is in how many words do you need to learn a language. The same measurements for French and German tell noticeably different stories.
Sources
- Nation, I. S. P. (2006). How large a vocabulary is needed for reading and listening? The Canadian Modern Language Review, 63(1), 59β82. lextutor.ca
- Laufer, B. & Ravenhorst-Kalovski, G. C. (2010). Lexical threshold revisited. Reading in a Foreign Language, 22(1), 15β30. hawaii.edu
- Davies, M. A Frequency Dictionary of Spanish, built on the 20-million-word Corpus del EspaΓ±ol, about a third of it transcribed speech. corpusdelespanol.org
- The measurements of our own list in this post are reproducible from the published list.