How many words do you need to speak Italian?
Italian has something most languages do not: a basic vocabulary drawn up by hand, by one of its best-known linguists, and revised within the last decade. Tullio De Mauro's Nuovo vocabolario di base della lingua italiana sets out the words an Italian speaker knows best, in three layers. That makes Italian the one language where a frequency list can be checked against expert judgement rather than only against another frequency list.
We did that check with our own Italian list. It mostly agrees with De Mauro — and it also showed us a gap in ours: 315 of his fundamental words were missing. We fixed that on 24 September 2026, and this post reports the list both before and after, rather than quietly correcting it.
What De Mauro's list actually is
The figures usually quoted for it — about 2,000 fundamental words, 2,750 of high use, 2,300 of high availability, some 7,000 in all — are real, but they describe the first version, published in 1980. De Mauro's own introduction to the new version gives both.
- Fondamentale (fundamental): about 2,000 words. The most used words in the language.
- Alto uso (high use): 2,750 in 1980, about 3,000 in the new version. Much less used than the fundamental layer, but far more than everything else.
- Alta disponibilità (high availability): about 2,300 in 1980, about 2,500 now. Words that are not especially frequent in texts but that every speaker knows and would reach for — the example De Mauro gives from the 1980 reception is peperone, pepper, which a critic said belonged on the list.
The new version is built on a corpus of 18.8 million words in six kinds of text, from newspapers and textbooks to film scripts, online chat and recorded speech. Its headline figures: the fundamental words cover 86% of the words in those texts, and the high-use layer adds another 6%.
Our own count of the headwords in the published list, by typeface as its key describes, finds 1,978 fundamental, 2,975 high-use and 2,216 high-availability words. The first two match De Mauro's round figures. The third is lower than his "about 2,500", and the difference may well be our parsing rather than his list.
The coverage curve
De Mauro's 86% is measured on lemmas: essere, not è, sono and era separately. A frequency list of forms, as a learner meets them, gives a lower number. On subtitles, which are close to speech:
| Language | First 100 | First 500 | First 1,000 | First 2,000 | First 5,000 | First 10,000 |
|---|---|---|---|---|---|---|
| Italian | 47.0% | 67.7% | 74.7% | 80.8% | 87.7% | 91.9% |
| Spanish | 50.7% | 68.6% | 74.9% | 80.6% | 87.4% | 91.6% |
| French | 56.7% | 73.8% | 79.4% | 84.5% | 90.1% | 93.6% |
The first 1,000 Italian forms cover 74.7% of the dialogue and 2,000 cover 80.8%. Reaching 95% takes 19,471 forms. That puts Italian level with Spanish and well behind French, which the comparison of all seven explains.
The two ways of counting are not in conflict. A lemma list counts essere once; a form list counts every conjugated form of it separately, and Italian verbs have a great many forms. Both are right about what they count.
De Mauro against our list
Our Italian list is ranked by frequency and runs to 10,333 entries. Where do De Mauro's three layers land in it?
| De Mauro layer | Words in his list | Also in ours | In our first 1,000 | In our first 5,000 |
|---|---|---|---|---|
| Fundamental | 1,978 | 1,968 | 808 | 1,831 |
| High use | 2,975 | 2,163 | 111 | 1,483 |
| High availability | 2,216 | 659 | 3 | 145 |
Read the first column of results as the agreement, and it is strong. Of the first 1,000 words in our list, 808 are De Mauro fundamentals and 111 more are high-use words. Nearly all of the rest are forms a lemma list files elsewhere — è and sono under essere, della and nella under their prepositions — and the names of places.
The high-availability layer behaves exactly as De Mauro said it would. Only 3 of those words are in our first thousand, because being known by everyone and being frequent in text are different things. A frequency list, however good, cannot find them. That is the best argument there is for a hand-built list alongside a counted one.
The gap in ours, and the fix
The comparison turned up something we did not expect: 315 of De Mauro's fundamental words were not in our Italian list at all. They included cinque, dieci, caffè, cuore, albero, banca, cibo and cinema — and, just as surprising, the possessives mio, tuo and suo, niente, qualcosa, madre, gente and ieri.
Those are not rare words. Our list ranks what occurs most in its source texts, and those texts evidently favoured written, encyclopaedic Italian: the names of cities and countries ranked high, while everyday nouns and the numbers were thin. The course's lessons used many of the words anyway, but a learner working through the list's flashcards would never have been dealt them. The same check found a second fault: a few words stored without their final accent — qualita for qualità, finche for finché, giu for giù — some of them duplicates of the accented word already in the list.
What we changed. Each missing word was added by hand, with its meaning written out and its gender marked, and placed among the words the list already treats as equally basic: quattro and cinque beside tre, madre beside padre, marito beside moglie. The accentless rows were corrected or removed. Learners already past a new word's place in the list will not be dealt it; no word that was already in the list is skipped or dealt twice.
| Our Italian list | Before, 23 Sept | After, 24 Sept |
|---|---|---|
| De Mauro's fundamental words it contains | 1,663 | 1,968 |
| …of them in our first 1,000 | 775 | 808 |
| …of them in our first 5,000 | 1,552 | 1,831 |
| Fundamental words still missing | 315 | 10 |
The 10 still missing are left out on purpose: three vulgar words, two variant spellings (okay, tivù), two former currencies (marco, lira), two nouns the list already carries in the plural form people actually use (pantaloni, occhiali), and bisognare, which it now carries as bisogna — the form a learner meets. It is still the clearest case we know for checking a counted list against a considered one.
So what is the number?
- About 1,000 words — mostly De Mauro's fundamental layer — is where listening stops being noise.
- About 2,000, the whole fundamental layer, gets you to the figure De Mauro measured: 86% of the words in real texts, counted as lemmas.
- The high-availability words have to be learned deliberately. They will not come up often enough to be picked up from frequency order.
- Past that, as in every language, how well you know your words matters more than how many.
How many words do you need is the general version of this argument, and the Italian course teaches from the list in frequency order.
Sources
- De Mauro, T. (2016). Il Nuovo vocabolario di base della lingua italiana. Internazionale, 23 December 2016, with the full list as a PDF. internazionale.it
- Lison, P. & Tiedemann, J. (2016). OpenSubtitles2016: Extracting Large Parallel Corpora from Movie and TV Subtitles. Proceedings of LREC 2016. aclanthology.org
- Dave, H. FrequencyWords: frequency lists from OpenSubtitles 2018, content licensed CC BY-SA 4.0. github.com
- The coverage figures and the comparison are computed by
scripts/frequency_curves.pyin Verbavia's repository, and printed from its output rather than typed.