How many words do you need to speak Esperanto?
Esperanto was designed to be learned quickly, and the design is visible in its vocabulary. Every noun ends in -o, every adjective in -a, every present-tense verb in -as, and a small set of prefixes and suffixes turns one root into a family: lerni (to learn), lernanto (a learner), lernejo (a school), lernigi (to teach). The original vocabulary, in Zamenhof's Unua Libro of 1887, had around 900 roots.
So the natural prediction is that Esperanto should have the steepest frequency curve of any language: a few hundred roots, endlessly recombined, should cover more text than a few hundred words of French or Spanish. We tested it. As usually measured, the prediction is wrong — and the reason it is wrong is the interesting part.
The measurement
The counts come from the Esperanto section of the OpenSubtitles 2018 corpus, which is small: 403,882 words, compared with hundreds of millions for Spanish or Portuguese. To compare fairly, the comparison of all seven languages samples every other language down to that same size.
- Esperanto
- French
- Spanish
| Language | First 100 | First 500 | First 1,000 | First 2,000 | First 5,000 | First 10,000 |
|---|---|---|---|---|---|---|
| Esperanto | 51.2% | 67.5% | 73.3% | 79.0% | 86.4% | 91.6% |
| French | 56.6% | 73.9% | 79.5% | 84.7% | 90.7% | 94.5% |
| Spanish | 50.6% | 68.6% | 75.0% | 80.8% | 87.8% | 92.5% |
Counted in word forms — the way every other language in the comparison is counted — the first 1,000 Esperanto words cover 73.3% of the dialogue. That is less than Spanish at 75.0%, far less than French at 79.5%, and second from last of all seven. Reaching 95% takes 16,359 forms.
Why regularity flattens the curve
The regular endings that make Esperanto easy to learn are exactly what makes its curve look flat, because a frequency list counts forms. Esperanto marks the plural with -j and the object with -n, and an adjective agrees with its noun in both. So one adjective has four forms — bona, bonaj, bonan, bonajn — and each is a separate entry in the list. A noun has four more. A verb has six common endings. Every one of them splits the same word's frequency across several rows.
French and Spanish have irregular forms, but many of their commonest words never change at all: de, que, en, y, et. Esperanto's grammar marks more things, more consistently, in more words. For a learner that is a help — the forms are predictable — but a form-based frequency list cannot know that.
Counting stems instead
Esperanto is the one language here where the grammatical endings can be removed mechanically, because they are the same on every word. We did that — stripping -o, -a, -e, their -j and -n, and the verb endings — and counted stems instead of forms:
| Counted as | Different items | First 100 | First 1,000 | First 5,000 | Items for 95% |
|---|---|---|---|---|---|
| Word forms | 36,346 | 51.2% | 73.3% | 86.4% | 16,359 |
| Stems | 20,811 | 53.2% | 80.7% | 93.6% | 6,423 |
Counted by stem, 1,000 items cover 80.7% — higher than French forms. The regularity is real, and it pays.
But this is not a fair win, and it would be dishonest to call it one. The other languages were counted as forms. Counting them as dictionary words would raise their curves too, and for them that cannot be done by stripping endings: es, son and fue are all ser, and no rule finds that. What the table shows is how much of Esperanto's apparent vocabulary is grammar, not how it would rank if every language were counted the same way. No open corpus lets us answer that.
Two more limits on the stem count:
- It strips grammatical endings only. The derivational affixes — mal-, -ej-, -an-, -ig- — are left in, so lernejo and lernanto still count separately. Reducing words to roots would steepen the curve further, and would need a real morphological analyser rather than a rule.
- It is a rule, not a dictionary. It occasionally strips a final vowel that belongs to the root, and leaves short words alone. It is an estimate of the effect, not a precise count.
What this means for a learner
- Learn the endings once and the forms come free. The stem figures are the honest measure of what an Esperanto learner has to memorise: far less than the form count suggests.
- Learn the affixes early. They are where the "few roots, many words" design actually lives, and no frequency count shows them.
- Expect the usual shape anyway. Even by stem, the last few percent of any text are rare words, and those have to be learned one at a time, as in every language.
Our Esperanto list and the Esperanto course teach roots and affixes as they come, and how many words do you need is the general argument.
Caveats
- The corpus is small, 403,882 words, and probably drawn from a small number of films. A larger corpus of Esperanto writing might look quite different.
- Subtitles are one register, close to speech, and in Esperanto many will be translations.
Sources
- Esperanto vocabulary, including the roughly 900 roots of the Unua Libro. en.wikipedia.org
- Wennergren, B. Plena Manlibro de Esperanta Gramatiko (PMEG), on the endings and affixes. bertilow.com
- Lison, P. & Tiedemann, J. (2016). OpenSubtitles2016: Extracting Large Parallel Corpora from Movie and TV Subtitles. Proceedings of LREC 2016. aclanthology.org
- Dave, H. FrequencyWords: frequency lists from OpenSubtitles 2018 (
eo), content licensed CC BY-SA 4.0. github.com - The coverage figures and the stem count are computed by
scripts/frequency_curves.pyin Verbavia's repository and printed from its output rather than typed.