Love or money, dogs or cats: what films say most
Film and television talk about love more than money and dogs more than cats in every one of the 8 subtitle corpora we measured — seven languages, with Portuguese counted separately for Brazil and Portugal — and coffee more than tea in all but possibly one. The margins are what differ. French says love 3.0 times as often as money; Mandarin barely more than once. Dogs beat cats by between 2.5 and 3.0 times. And English, for all its tea, still says coffee 1.5 times as often.
Love or money?
- love
- money
| Language | Love, per million | Money, per million | Love ÷ money |
|---|---|---|---|
| English | 1,345.1 | 641.9 | 2.1× |
| Spanish | 843.6 | 710.2 | 1.2× |
| French | 1,521.6 | 501.2 | 3.0× |
| Italian | 835.1 | 638.5 | 1.3× |
| Portuguese (Brazil) | 1,116.2 | 695.4 | 1.6× |
| Portuguese (Portugal) | 817.4 | 729.3 | 1.1× |
| German | 1,029.8 | 627.9 | 1.6× |
| Mandarin | 669.9 | 619.7 | 1.1× |
Love wins in all 8. It still wins in all 8 when the commonest slang word for money is added — fric in French, grana in Brazil, Kohle in German, cash in English.
French is the most romantic by a distance, and the reason is grammatical as much as sentimental: aimer is the everyday verb for both "love" and "like" (j'aime le chocolat), so French says it constantly. Spanish and Italian look cooler than they are, because the commonest way to say "I love you" in both avoids the love verb altogether: Spanish te quiero uses querer, "to want", and Italian ti voglio bene uses volere, "to want", and bene. Neither is counted here, because quiero and voglio mostly mean "I want". Mandarin is the closest race: 爱 ài (love) against 钱 qián (money), nearly level.
Portugal is a lesson in counting. European Portuguese puts the pronoun after the verb and joins it with a hyphen: amo-te, "I love you". The subtitle lists count amo-te as one word, and it is said 23,325 times. Count only the bare forms — amo, ama, amor — and Portugal says love 653.6 times per million words, less than money's 729.3. Count them properly, and love wins there too.
Dogs or cats?
- dog
- cat
| Language | Dog, per million | Cat, per million | Dog ÷ cat |
|---|---|---|---|
| English | 223.1 | 87.4 | 2.6× |
| Spanish | 206.1 | 73.7 | 2.8× |
| French | 185.1 | 68.4 | 2.7× |
| Italian | 190.1 | 67.5 | 2.8× |
| Portuguese (Brazil) | 206.9 | 80.2 | 2.6× |
| Portuguese (Portugal) | 218.3 | 71.9 | 3.0× |
| German | 182.4 | 64.9 | 2.8× |
| Mandarin | 187.7 | 73.9 | 2.5× |
Dogs win in all 8, and by a wide and steady margin: between 2.5 and 3.0 times as often. Part of that is the stories — police dogs, guard dogs, the family dog — and part is that "dog" is an insult in most of these languages and "cat" rarely is.
That is also why the feminine forms are left out of both columns. Spanish perra, French chienne and Italian cagna ("bitch") are mostly insults on screen, and chatte in French and gata in Brazilian Portuguese are mostly slang. Spanish perra and perras alone add 83.9 per million words to the dog column; leaving them out makes the dog's lead smaller, not larger.
Coffee or tea?
- coffee
- tea
| Language | Coffee, per million | Tea, per million | Coffee ÷ tea |
|---|---|---|---|
| English | 124.0 | 81.7 | 1.5× |
| Spanish | 155.6 | 75.7 | 2.1× |
| French | 143.8 | 63.2 | 2.3× |
| Italian | 71.6 | 36.0 | 2.0× |
| Portuguese (Brazil) | 207.7 | 79.1 | 2.6× |
| Portuguese (Portugal) | 159.1 | 67.7 | 2.4× |
| German | 128.1 | 60.1 | 2.1× |
| Mandarin | 103.3 | 44.7 | 2.3× |
As counted, coffee wins in all 8 — but only 7 of those wins are safe. Brazil is the most coffee-minded (2.6 to one), which fits the world's largest coffee producer; English is the least, at 1.5 to one — the only corpus where tea puts up a real fight, and probably the British share of it. Café also means a café in Spanish, French and Portuguese, which flatters coffee there.
Italy is the one we cannot call. Italian subtitles often spell tea the rather than tè, and the is also the English word left behind in subtitles, so it cannot simply be counted. It is said 174.3 times per million words of Italian. Scaling from other English words in the Italian list (of, and, you, to) suggests about 22,244 of its 43,070 occurrences are English; if every one of the rest were tea, Italian tea would reach 120.3 per million, ahead of coffee's 71.6. The truth is somewhere in between, and we would rather say so than print a winner.
The data
Every figure is how many times per million words of subtitles the words were said, with every form in the table below added together:
| Language | love | money | dog | cat | coffee | tea |
|---|---|---|---|---|---|---|
| English | love, loves, loved, loving | money | dog, dogs | cat, cats | coffee, coffees | tea, teas |
| Spanish | amor, amores, amo, amas, ama, amamos, aman, amar, amado, amada, amé, amó, amarte, amarla, amarlo | dinero | perro, perros | gato, gatos | café, cafés | té |
| French | amour, amours, aime, aimes, aiment, aimons, aimez, aimer, aimé, aimée | argent | chien, chiens | chat, chats | café, cafés | thé, thés |
| Italian | amore, amori, amo, ami, ama, amiamo, amate, amano, amare, amato, amata, amarti, amarla, amarlo | soldi, denaro | cane, cani | gatto, gatti | caffè, caffé | tè |
| Portuguese (Brazil) | amor, amores, amo, ama, amamos, amam, amar, amado, amada | dinheiro | cachorro, cachorros, cão, cães | gato, gatos | café, cafés | chá |
| Portuguese (Portugal) | amor, amores, amo, ama, amamos, amam, amar, amado, amada | dinheiro | cão, cães, cachorro, cachorros | gato, gatos | café, cafés | chá |
| German | liebe, lieben, liebst, liebt, geliebt, liebte | geld | hund, hunde, hunden, hundes | katze, katzen, kater | kaffee | tee |
| Mandarin | 爱, 爱情 | 钱, 金钱 | 狗, 小狗, 狗狗 | 猫, 小猫, 猫咪 | 咖啡 | 茶 |
Methodology, and using this data
- The corpus. OpenSubtitles 2018, the largest collection of film and television subtitles, via the frequency lists compiled by Hermit Dave (FrequencyWords). The corpora are large: 735 million words for English, 85 million for Mandarin, the smallest here. Mandarin subtitles are already split into words; the other languages are counted word by word as spelt.
- What counts as a word. Every inflected form of the word, singular and plural, and for love both the noun and the verb — because English love is both and cannot be split, the only fair comparison counts both everywhere. In Spanish, Italian, French and Portuguese, forms with a pronoun attached (amarte, amo-te) count too.
- Per million words, so that corpora of different sizes can be compared.
- What was left out, and why. The feminine forms of dog and cat (insults and slang); slang words for money in the main measurement (added in a second one, above); querer and volere in Spanish and Italian ("to want" first, "to love" second); Italian the. Each choice is stated where it could move a result. None of them changes a winner, except that Italian the could turn coffee against tea in Italy, as explained above.
- Esperanto is not included. Its subtitle corpus is only 403,882 words, and the six words are said between 35 and 388 times each in all of it — too few to put beside the others. Shanghainese has no subtitle corpus at all.
- What this measures. What is said in films and television, which is scripted, often translated, and fond of crime and romance. It is the closest large corpus to everyday speech, not a survey of what people like.
The figures and charts on this page may be reused with credit to Verbavia and a link to this page. Because the underlying counts come from FrequencyWords, which is licensed CC BY-SA 4.0, the figures here are shared on the same terms: credit "Verbavia, from OpenSubtitles 2018 via FrequencyWords".
More from the same data
The same corpus shows which language has the steepest frequency curve, and the words that don't exist in English and French words used in English come from it too. Every Verbavia course is built on a frequency list like these — the Spanish course starts from the 1,000 most common Spanish words, and the word lists for every course are free to download.
Sources
- Lison, P. & Tiedemann, J. (2016). OpenSubtitles2016: Extracting Large Parallel Corpora from Movie and TV Subtitles. Proceedings of LREC 2016. aclanthology.org
- Dave, H. FrequencyWords: frequency lists from OpenSubtitles 2018, content licensed CC BY-SA 4.0. github.com. The counts are computed by
film_wordsinscripts/frequency_extras.py, which lists every form counted.