Which language says thank you most often?
Every language has a word for thank you, and every phrasebook puts it on the first page. But some cultures are said to thank constantly and others hardly at all, and that is something a word count can test. We counted the everyday word for thanks in the film and television subtitles of eight languages. Spanish says it most often, and Mandarin least: gracias 1,841 times in every million words, 谢谢 1,320. English sits in the middle, and the whole range is narrower than you might expect.
Thank you, per million words
| Language | Counted | Per million words | Rank of the commonest form | The count also includes |
|---|---|---|---|---|
| Spanish | gracias | 1,841.5 | 70 | also gracias a, thanks to |
| Italian | grazie | 1,765.1 | 74 | also grazie a, thanks to |
| Portuguese (Brazil) | obrigado, obrigada, obrigados, obrigadas, obrigadinho, obrigadinha, obrigadão, obrigadíssimo, obrigadíssima | 1,739.9 | 112 | also obrigado meaning obliged, forced |
| Portuguese (Portugal) | obrigado, obrigada, obrigados, obrigadas, obrigadinho, obrigadinha, obrigadão, obrigadíssimo, obrigadíssima | 1,639.4 | 118 | also obrigado meaning obliged, forced |
| English | thank, thanks, thankyou, thank-you | 1,595.1 | 131 | also thank God, and thanks to |
| French | merci | 1,481.4 | 95 | also merci de, and the rare à la merci de |
| German | danke, dankeschön | 1,472.1 | 98 | danke and dankeschön; vielen Dank is added below |
| Mandarin | 谢谢, 谢谢您, 谢谢你们, 多谢, 多谢您 | 1,320.6 | 92 | 谢谢 and 多谢, with their pronouns; not 谢 alone, which is also a surname |
Each language's word is counted with every spelling and form that a subtitler uses for it: obrigado and obrigada, danke and dankeschön, English thank and thanks. The figures are per million words, which puts a corpus of 85 million Mandarin words on the same scale as 735 million English ones.
Spanish, Italian and Brazilian Portuguese are close together at the top; French, German and Mandarin are at the bottom. But the top is only 1.39 times the bottom: every one of these languages says thank you somewhere between once and twice in every thousand words of dialogue.
What the counts include, and what they leave out
A word count has no context, so each count catches a little more than thanks, and the table above says what. Spanish gracias a and Italian grazie a mean "thanks to"; English thank includes "thank God"; Portuguese obrigado can also mean "obliged". A word count cannot separate those uses, so read every figure as slightly high.
German needs a second look. The list counts danke, but Germans also say vielen Dank, "many thanks", and a word list cannot tell that Dank from dank meaning "thanks to", or from Gott sei Dank, "thank God". So we read the full German subtitle text, with its capitals and word order intact, and counted only the Dank that comes after vielen, herzlichen, besten, schönen, tausend or recht. Those add 14.6% to danke, and lift German to 1,686 per million — ahead of European Portuguese, English and French. (Gott sei Dank turned up 9,822 times in the full text, and is not counted.)
Mandarin has a formal second word too. 感谢 (gǎnxiè, to thank, to be grateful), with 非常感谢, "thank you very much", adds about 176 per million, though some of it is 感谢上帝, "thank God", and it is not in the chart. 谢 on its own is left out entirely, because it is also a common surname, Xiè.
Slang is left out everywhere. Brazil's valeu ("cheers, thanks", 42 per million, and also "it was worth") and English cheers (65) would each add a little to their language.
Why "per million words" is only nearly fair
Counting per million words is the obvious way to compare corpora of different sizes, and it is not perfectly fair, for two reasons worth knowing.
Languages use different numbers of words to say the same thing. Spanish, Italian and Portuguese usually leave out the subject pronoun — gracias, lo tengo is "thanks, I've got it" with the "I" built into the verb — so the same conversation takes fewer words than in English or French, and every word in it, gracias included, gets a slightly larger share. Mandarin is counted in words that the corpus has already segmented, which is a choice about where one word ends rather than a fact about the language. There is no clean way to correct for this, so small differences between neighbouring languages in the chart should not be over-read.
Subtitles are not conversation. A large share of the films and series in the corpus were made in English and translated, so a translator's habits — whether to render "thanks, man" as a thank-you at all — are part of every number. And subtitles are condensed for reading, which tends to cut small talk. The honest claim is about dialogue as it is subtitled, which is the largest record of everyday speech there is for these languages, and the dialogue a learner will actually meet on screen.
Esperanto is not in the chart because its subtitle corpus is too small: 403,882 words, in which dankon appears 547 times.
What a count cannot tell you
A count says how often, not why. Mandarin's lower figure could be about politeness, about which films were subtitled, or about the many other ways Mandarin has of acknowledging a favour, and the word lists cannot separate those. The same goes for the top of the chart: Spanish dialogue saying gracias more often is a fact about the subtitles, not proof that Spanish speakers are more grateful. What the counts do settle is the size of the differences, and they are modest. They also put English, at 1,595, below Spanish, Italian and both kinds of Portuguese.
The replies are countable too. Mandarin's 不用谢, "no need to thank me", is common enough to have its own entry in the corpus, at 7 per million.
If you are learning one of these languages
Thank you is among the 150 commonest words of every one of them: gracias is number 70 in Spanish dialogue, merci 95 in French, 谢谢 92 in Mandarin. Every course's word list ranks words like these by how often they are actually said — Spanish, French, Italian, Portuguese, German and Mandarin — and the Spanish vocabulary test shows in two minutes how far down the list you already are. The same eight corpora are compared word by word in OK in other languages and love or money: what films say most, and the small words people fill their sentences with are counted in filler words in every language.
Sources
- Lison, P. & Tiedemann, J. (2016). OpenSubtitles2016: Extracting Large Parallel Corpora from Movie and TV Subtitles. Proceedings of LREC 2016. aclanthology.org
- Dave, H. FrequencyWords: frequency lists from OpenSubtitles 2018, content licensed CC BY-SA 4.0. github.com
- The full German text of the same release, from OPUS: OpenSubtitles v2018, German, read for vielen Dank and its kin.
- The words and forms counted, and what each count also includes, are listed in
scripts/frequency_extras.py(THANKS,DANK_THANKS).