How many words do you need to speak Mandarin?
Ask how many words you need for Mandarin and you will usually be answered in characters. The two are not the same unit, and the difference is not a detail: on the same corpus of film and television dialogue, the first 1,000 characters cover 91.1% of what is said, and the first 1,000 words cover 69.2%.
Both numbers are correct. They answer different questions, and a learner needs to know which question is being answered before trusting either.
Characters and words are different units
A Chinese character is a written syllable. A Chinese word is one or more of them: 我 wǒ (I) is one character and one word, but 我们 wǒmen (we) is two characters and one word, and 知道 zhīdào (to know) is two characters whose meaning is not the sum of 知 and 道. A great many of the commonest words are two characters long.
So characters are few and reused endlessly, and words are many. Count characters and you are counting building blocks; count words and you are counting the things built from them.
The two curves
The same corpus, measured both ways:
- Mandarin (characters)
- Mandarin (words)
| Language | First 100 | First 500 | First 1,000 | First 2,000 | First 5,000 | First 10,000 |
|---|---|---|---|---|---|---|
| Mandarin (characters) | 54.3% | 81.3% | 91.1% | 97.3% | 99.9% | 100.0% |
| Mandarin (words) | 44.8% | 62.5% | 69.2% | 75.3% | 82.7% | 87.6% |
The character curve is almost vertical. 916 characters cover 90% of the dialogue and 1,458 cover 95%. The word curve is the flattest of any language we measured: 90% takes 14,835 words and 95% takes 41,853. The comparison of all seven languages sets this against the European languages, and Mandarin by word comes last at every point.
A check against the official list
China's 1988 List of Frequently Used Characters in Modern Chinese has two levels: 2,500 common characters and 1,000 less common ones. Its own figures, measured on written text, are that the first level covers 97.97% of characters in use and both levels together cover 99.48%.
Our subtitle measurement gives 98.4% for the first 2,500 characters and 99.4% for 3,500. A corpus of film dialogue and a survey of written Chinese thirty years apart agree to within half a percentage point. For characters, the numbers are solid.
Which number should a learner care about?
For reading, characters matter first. Knowing 1,000 characters means very few characters on a page will be new to you. But it does not mean you can read the page, because knowing 知 and 道 separately does not tell you that 知道 means to know.
For understanding, words matter. Speech arrives in words, and meaning lives in words. The honest figure for a learner is the word curve, and it is sobering: 1,000 words cover 69.2% of the dialogue, less than any European language at the same size.
The good news is in the gap. Because most words are built from a small set of characters, a learner who knows the characters can often guess a new word, or at least remember it quickly: 电 diàn (electricity) runs through 电话 (telephone), 电视 (television), 电脑 (computer) and 电影 (film). The characters are the leverage; the words are the goal.
Why the word curve is so flat
Some of it is real and some of it is measurement.
- Real: Chinese makes new words by compounding freely, and a great many of them are only moderately frequent. The long tail is genuinely long.
- Measurement: there are no spaces in written Chinese, so a word list depends on software that decides where one word ends. The corpus's own segmentation counts 766,612 different words, and where a name or a fixed phrase counts as one word or several is exactly the kind of decision another segmenter could make differently. A different tool would move the word curve, though not the character curve.
That is the reason every Mandarin figure you read needs its unit stated. A character figure and a word figure can differ by a wide margin at the same count and both be right.
What Verbavia does with this
Our Mandarin list counts words, not characters, because words are what you understand and say. Each word is taught with its characters, and the Mandarin course shows how characters combine into the words you learn, so the leverage of the character curve is not wasted.
How many words do you need is the general argument, and the same measurement for the other languages is in the comparison of all seven.
Caveats
- Subtitles are one register, close to conversation. Written Chinese has more formal vocabulary.
- Many subtitles are translations, of films made in other languages.
- The corpus is in simplified characters; a traditional-character corpus would differ in places.
Sources
- List of Frequently Used Characters in Modern Chinese (现代汉语常用字表), State Language Commission and State Education Commission, 1988; coverage figures as cited from Su Peicheng (2014), 现代汉字学纲要. en.wikipedia.org
- Lison, P. & Tiedemann, J. (2016). OpenSubtitles2016: Extracting Large Parallel Corpora from Movie and TV Subtitles. Proceedings of LREC 2016. aclanthology.org
- Dave, H. FrequencyWords: frequency lists from OpenSubtitles 2018 (
zh_cn), content licensed CC BY-SA 4.0. github.com - The coverage figures are computed by
scripts/frequency_curves.pyin Verbavia's repository — the character curve by splitting every word into its characters — and printed from its output rather than typed.