How many words does the average person know?
The best-known careful estimate says that an average 20-year-old native speaker of American English knows 42,000 words β or, depending on how you count, 11,100. Both numbers come from the same study, and the gap between them is the most useful thing to understand about vocabulary size. It also explains why a learner needs far fewer words than a native speaker to get by.
The estimate, and where it comes from
In 2016 Marc Brysbaert, MichaΓ«l Stevens, PaweΕ Mandera and Emmanuel Keuleers, psychologists at Ghent University, combined a review of earlier research with a large crowdsourced vocabulary test, taken online by 221,268 people. Their summary, in Frontiers in Psychology:
Based on an analysis of the literature and a large scale crowdsourcing experiment, we estimate that an average 20-year-old native speaker of American English knows 42,000 lemmas and 4,200 non-transparent multiword expressions, derived from 11,100 word families.
The range is wide: 27,000 lemmas for the lowest 5% of people and 52,000 for the highest 5%. And the number keeps growing: between 20 and 60, the average person learns 6,000 more lemmas, or in the authors' words "about one new lemma every 2 days". One caution about who "average" is: the conclusion describes the typical participant as "an average 20-year-old student", and students are not everyone.
Why the number depends on what a "word" is
Earlier estimates, the authors note, "range from less than 10 thousand words known to over 200 thousand words mastered", and nearly all of that disagreement is about the unit being counted.
- A lemma is a dictionary headword with all its inflections: walk, walks, walked and walking are one lemma. This is the unit behind the 42,000.
- A word family also folds in the words derived from it: walker and walkable join walk. Counted that way, the same people know 11,100.
- A word form is every spelling separately β walked and walks are two. Frequency lists count forms, and counted as forms the numbers get much bigger again.
- Multiword expressions β give up, by and large β are meanings a word count misses entirely. The study adds 4,200 of the ones whose meaning cannot be worked out from their parts.
There is a second catch, which the authors state plainly: "The knowledge of the words can be as shallow as knowing that the word exists." Knowing 42,000 lemmas does not mean being able to use 42,000 lemmas. A native speaker's active vocabulary β the words they actually produce β is much smaller than the words they recognise.
How many words does a learner need?
Nowhere near a native speaker's total, because words are not equally useful. A small number of them do most of the work in any language. On film and television dialogue in the languages we teach, the most common word forms cover this much of what is said:
| Language | Forms for half of speech | For 90% | For 95% |
|---|---|---|---|
| French | 64 | 4,876 | 13,909 |
| Portuguese (Brazil) | 89 | 5,705 | 15,078 |
| German | 91 | 6,465 | 21,502 |
| Spanish | 95 | 7,577 | 20,854 |
| Italian | 123 | 7,181 | 19,471 |
| Mandarin | 161 | 14,835 | 41,853 |
Two things to take from that table. First, the start is astonishingly concentrated: somewhere between 64 and 161 word forms are half of everything anyone says, in every one of these languages. Second, the tail is long. These are word forms, not lemmas β so es, son and fue are three rows β and even so, the last few per cent of speech take thousands more.
Where a learner needs to be depends on what they want to do. Research on reading suggests that comfortable, unassisted comprehension starts at around 98% of the words in a text, and that figure, and how it was measured, is set out in how many words do you need to learn a language. For following a conversation, less will do, because speech is repetitive and you can ask.
So what is a realistic target?
- The first 1,000 words are the best return anywhere in language learning: between two thirds and four fifths of ordinary speech, depending on the language.
- 5,000 words cover between 82.7% (Mandarin) and 90.1% (French) of speech β roughly the range where a slow, clear speaker or a graded book becomes followable.
- Past about 10,000, the gains come less from new words and more from knowing the ones you have in more contexts β which is where most of a native speaker's 42,000 recognised lemmas live.
A native speaker's total is not the target. The words that make up most of what you will hear are, and they are in frequency order in each of our lists: Spanish, French, German, Italian, Portuguese and Mandarin. How the curves compare between languages is in which language has the steepest frequency curve, and how long the journey takes in hours in the hardest and the easiest languages for English speakers.
The Spanish, French and German lists can also be downloaded free: Spanish, French and German as printable PDFs, or Spanish, French and German as CSV files for Excel or Anki.
Sources
- Brysbaert, M., Stevens, M., Mandera, P. & Keuleers, E. (2016). How many words do we know? Practical estimates of vocabulary size dependent on word definition, the degree of language input and the participant's age. Frontiers in Psychology, 7, 1116. frontiersin.org
- Lison, P. & Tiedemann, J. (2016). OpenSubtitles2016: Extracting Large Parallel Corpora from Movie and TV Subtitles. Proceedings of LREC 2016. aclanthology.org
- Dave, H. FrequencyWords: frequency lists from OpenSubtitles 2018, content licensed CC BY-SA 4.0. github.com. Every coverage figure is computed by
scripts/frequency_curves.pyand printed from its output rather than typed.