VERBAVIA

How many words do you need to speak German?

German has a reputation for vocabulary: endless compounds, famously long nouns, a dictionary that never stops. It is the language people expect to need the most words for.

Updated 22 September 2026 · 5 min read

Measured against our own German frequency list โ€” 5,000 entries, ranked โ€” the reputation is not exactly wrong, but it points at the wrong problem. The thing that inflates German vocabulary counts is not compounding. It is that compounds break the counting itself.

Why counting German words does not work

A compound like Geschwindigkeitsbegrenzung is one entry in a word list. For a learner who knows Geschwindigkeit and Begrenzung it is not one new word; it is zero new words and a moment of assembly. For a learner who knows neither it is two.

So a German frequency list systematically overcounts what must be memorised and undercounts what can be decoded. Both errors at once, in the same column of numbers. No amount of care with the list fixes this, because the unit of measurement does not match the unit of learning.

This is worth stating plainly because it undermines every cross-language vocabulary comparison you will read. "German needs more words than Spanish" is not a finding; it is an artefact of what counts as a word in each.

What the list actually shows

Here is the part that surprised us. Our German list is not dominated by long words, and it has fewer of them than our Spanish list does.

Rank bandGerman entries โ‰ฅ12 charactersSpanish, same band
1โ€“1000%0%
101โ€“5000%0.8%
501โ€“1,0001.8%2.8%
1,001โ€“2,0002.6%5.3%
2,001โ€“5,0005.5%7.2%

Spanish has more long entries than German in every band where either has any. Spanish builds them with suffixes โ€” -ciรณn, -idad, -miento โ€” and German builds them by welding nouns together, but in the frequency ranges a learner actually works through, Spanish produces more of them.

The famous German compounds are real. They are just not common. They live far down the frequency curve, in technical and bureaucratic registers, which is exactly where a learner does not need them and exactly where journalists go looking for examples.

A real German coverage curve, from someone else's corpus

We cannot produce a coverage curve from our own list, because it carries rank and no frequency counts. Jones and Tschirner's frequency dictionary of German can, and does. It is built on a 4.2-million-word corpus divided evenly between speech, literature, newspapers and academic writing, and it publishes its own coverage figures:

Words knownShare of that corpus
the first 10about 27%
the first 20about 35%
the full 4,034-entry list80% to 90%, depending on register

Ten words for a quarter of everything. That is the frequency curve at its steepest, and it is the single best argument for learning in frequency order rather than by theme.

Note the last row carefully, though. Four thousand entries reaching 80โ€“90% is not the same claim as four thousand word families reaching 95%, which is roughly what the reading research finds. Different units, different corpora, different numbers โ€” which is the theme of this whole set of posts.

The inflection story is different from Spanish

Our German entries carry notes marking forms of other words, and in this list that note field is used for nothing else โ€” so the measurement is clean:

Rank bandEntries that are a form of another entry
1โ€“10014%
101โ€“25025%
251โ€“50023%
501โ€“1,00025%
1,001โ€“2,00026%
2,001โ€“5,00023%

About a quarter, and โ€” this is the interesting part โ€” flat. German's inflected forms sit at a steady proportion right across the list.

Spanish does something different: its equivalent figure climbs from 14% to 44% as you go down the ranks. Spanish verbs keep generating new high-frequency forms deep into the list; German's inflection is concentrated in a small number of very common words and then stops growing.

As with Spanish, treat these as floors. An unannotated entry is not counted, so the true proportion can only be higher.

What this means for effort

Three consequences, in the order they will hit you.

So what is the number?

The same honest non-answer as for every language, with a German-specific twist: whatever number you pick, it means less here than elsewhere, because the boundary between "a word you know" and "a word you can work out" is blurrier in German than in any other language we teach.

If you want a working target, the first thousand entries of the German list will carry you through ordinary conversation as far as vocabulary is concerned. What will stop you at that point is not vocabulary. It is the case system and the verb-final clause, which is why the German course teaches word order and cases before it teaches very much else.

Is the compounding claim actually true?

Partly, and the honest version is narrower than the folklore. Brysbaert and colleagues, reviewing how many words people know, cite research finding that German has more single-word compounds than English โ€” which partly explains why speakers appear to know more lemmas in German โ€” while noting that this "should not affect the number of monomorphemic lemmas known."

Which is the point exactly. German looks like it has more vocabulary because of how it writes things down, not because its speakers hold more distinct ideas in their heads. The compound is a spelling convention applied to something English would write as two words.

The caveat that applies to all of this

Coverage is not comprehension. Knowing 95% of the words in a German text does not mean understanding 95% of it โ€” and in German it can mean less than in Spanish, because the grammatical relationships you also need are carried by endings and word order rather than by the words themselves.

The general argument is in how many words do you need to learn a language, and the same measurements for Spanish and French show how differently three closely studied European languages behave.

Sources

Start a course

Start the free Spanish course

8 languages, each taught from a frequency list, with the grammar explained properly and reading and listening built from words you already know. One lesson a day, free. See all 8.

More from the blog