VERBAVIA

Which language has the steepest frequency curve?

Every language has a frequency curve: learn the commonest word, then the next, and track how much of real speech you now recognise. The curve always rises fast and then flattens. What nobody seems to publish is the comparison — the same measurement, on the same kind of text, for several languages side by side. So we made one.

Updated 24 September 2026 · 6 min read

The answer, for the languages Verbavia teaches: French has the steepest curve at every point we measured, and Mandarin, counted in words, the flattest. Esperanto, which is often expected to win, comes second from last for most of the curve. The reasons are more interesting than the ranking.

The chart

How much running text the most frequent words coverLine chart of text coverage against vocabulary rank on a logarithmic scale, for French, German, Portuguese (Brazil), Spanish, Italian, Esperanto, Mandarin (words), Mandarin (characters); every corpus sampled to the same 403,882 words. The same figures are in the table below.0%20%40%60%80%100%1101001,00010,000Words known, in frequency order (log scale)Share of running text covered
  • French
  • German
  • Portuguese (Brazil)
  • Spanish
  • Italian
  • Esperanto
  • Mandarin (words)
  • Mandarin (characters)
Coverage of OpenSubtitles 2018 text, every corpus sampled to the same 403,882 words. Source: FrequencyWords (CC BY-SA 4.0); computed by Verbavia.
LanguageFirst 100First 500First 1,000First 2,000First 5,000First 10,000
French56.6%73.9%79.5%84.7%90.7%94.5%
German51.4%71.1%77.6%83.0%89.1%92.9%
Portuguese (Brazil)51.6%70.2%76.8%82.8%89.7%94.0%
Spanish50.6%68.6%75.0%80.8%87.8%92.5%
Italian46.9%67.6%74.7%80.9%88.0%92.7%
Esperanto51.2%67.5%73.3%79.0%86.4%91.6%
Mandarin (words)44.9%62.7%69.5%75.7%83.3%88.6%
Mandarin (characters)54.2%81.3%91.1%97.4%100.0%100.0%

The figures are the share of running text — every word of every line of dialogue, counted each time it occurs — covered by the most frequent 100, 500, 1,000 and so on. The horizontal axis is logarithmic, because the interesting part happens in the first thousand words and would otherwise be squashed into the left edge.

How it was measured

Coverage needs counts, and our own word lists have rank only: they say which word comes first, not how often it occurs. So the counts come from an open source that has them for all seven languages: Hermit Dave's FrequencyWords lists, built from the OpenSubtitles 2018 corpus (Lison and Tiedemann, 2016). Subtitles are the largest openly available sample of something close to conversation, and every one of these languages has them.

Two things had to be controlled.

Corpus size. The corpora are wildly different sizes: 423 million words of Spanish, 85 million of Mandarin, and only 403,882 of Esperanto. A small corpus always looks steeper, because it has never met the rare words. So every language was measured on a random sample of exactly the same size — 403,882 words, the size of the Esperanto corpus — and the chart and the comparisons below use those equal-size figures only. For the six large corpora the equal-size figures at 1,000 words are within a fraction of a point of the full-corpus ones, so the sampling costs nothing where it matters.

What counts as a word. A word here is a form as it was written: es, son and era are three words, not one verb. That is deliberate — it is what a learner actually meets — but it matters a great deal for the ranking, as the German, Esperanto and Mandarin results show.

What the chart shows

French is steepest everywhere. The first 100 French words cover 56.6% of the dialogue, the first 1,000 cover 79.5%, and reaching 95% takes 11,020 words — fewer than any other language counted by words. Spoken French leans very heavily on a small set of pronouns and short function words, and subtitles are spoken French. The French post explains why written French would give a different answer.

German starts fast and then slows down. It is second at 500, 1,000 and 2,000 words, then falls behind Portuguese. On the full corpus, German needs 21,502 words to reach 95% against 13,909 for French: the longest tail of the five European languages. That is compounding, and the German post measures it.

Brazilian Portuguese is steeper than European Portuguese at every point, which is why the chart uses the Brazilian list — it is the variety the course teaches. The Portuguese post shows where the two lists part company.

Spanish and Italian sit together in the middle. At 1,000 words they are within half a point of each other: 75.0% and 74.7%. Spanish and Italian have posts of their own.

Esperanto is second from last from 500 words on. This is the result we did not expect, and it is explained in the Esperanto post: Esperanto's regular endings multiply the number of forms, and forms are what is being counted.

Mandarin by word is the flattest line; Mandarin by character is off the scale. The first 1,000 Mandarin words cover 69.5%; the first 1,000 characters cover 91.1%. Those are two different questions, and the Mandarin post is about why the difference matters so much.

Why "steepest" is not "easiest"

A steep curve means the commonest words carry more of the load. That is useful, and it is the whole argument for learning in frequency order. It does not make a language easy.

The steepest curve here belongs to French because spoken French recycles a small set of very frequent short words. Those are exactly the words that are hardest to hear, because they run together. German's early steepness is bought with a case system: the same few articles, der, die, das, den, dem, are near the top because they are in every sentence, and choosing between them is the hard part. None of that appears in a coverage figure.

And coverage is not comprehension. Research on reading puts comfortable understanding at about 98% of words known, and the words you are missing are never the easy ones. How many words do you need is the general version of this argument.

Caveats, all of them real

Sources and data

Start a course

Start the free Spanish course

8 languages, each taught from a frequency list, with the grammar explained properly and reading and listening built from words you already know. One lesson a day, free. See all 8.

More from the blog