Filler words in other languages: ouais, vale, allora, doch
Every language has words that carry no content and keep a conversation moving: English um, like and well, French euh and quoi, Italian beh and allora, German äh and doch. Textbooks mostly leave them out, and they are a large part of what makes a speaker sound native. We counted the main ones in the film and television subtitles of eight languages. The single most frequent pure filler in the whole set is Italian: beh, "well", said 1,680 times in every million words, number 82 among all the words of Italian dialogue.
The problem with counting fillers
Two things make this harder than it looks, and the tables below are built around them.
Most fillers are also ordinary words. Spanish tipo is the filler "like" and also "a type" or "a guy"; German doch softens a request and also means "yes" after a negative question; French quoi ends a sentence as "you know" and also means "what". The subtitle counts are of single words with no context, so the filler and the ordinary word are one number. We dealt with that by splitting every word into one of two groups:
- Pure fillers — words with no other job, so their whole count is filler. Only these are compared with each other.
- Ceilings — words with another meaning. Their count is the most the filler use could be, and the real figure is lower, sometimes much lower. They are listed with what else the word means, and are not ranked against the pure ones.
Subtitles tidy speech up. A subtitle has to be read in a couple of seconds, so hesitations are among the first things an editor cuts. Every hesitation sound below — um, euh, äh, 呃 — is therefore very likely rarer in subtitles than in real conversation. The counts are useful for comparing languages with each other, not as a measure of how much people really hesitate.
The pure fillers
| Language | Word | As a filler | Per million words | Rank in its language |
|---|---|---|---|---|
| Italian | beh | well | 1,680.0 | 82 |
| English | uh | a hesitation sound | 806.6 | 153 |
| Mandarin | 嗯 | a hesitation sound, or mm-hm | 806.2 | 137 |
| French | ouais | yeah | 758.6 | 151 |
| Spanish | eh | eh?, a hesitation or a tag | 456.7 | 242 |
| French | hein | eh?, right?, at the end of a sentence | 365.5 | 275 |
| English | um | a hesitation sound | 325.1 | 330 |
| Mandarin | 呃 | a hesitation sound | 248.0 | 417 |
| French | euh | a hesitation sound | 211.8 | 429 |
| German | äh | a hesitation sound | 178.5 | 567 |
| Italian | insomma | in short, well, come on | 132.4 | 753 |
| German | ähm | a hesitation sound | 97.8 | 876 |
| Italian | cioè | I mean, that is | 96.9 | 977 |
| Portuguese (Brazil) | né | right?, isn't it? (from não é) | 89.8 | 981 |
| German | naja | oh well (also written na ja) | 64.0 | 1,236 |
| Portuguese (Brazil) | hum | a hesitation sound | 49.8 | 1,644 |
| Italian | boh | dunno (with a shrug) | 3.0 | 14,904 |
Italian beh is far ahead, said five times as often as English um, and insomma ("in short", "well, come on") and cioè ("I mean") are both among the thousand commonest words of Italian dialogue. French ouais, plain "yeah", is number 151 in French; hein, the "eh?" at the end of a sentence, is number 275. Brazil's né — short for não é, "isn't it?" — tags the end of a sentence the way hein does.
English subtitles write uh more than twice as often as um. Together they are 1,131 per million words; German's äh and ähm together are 276. Mandarin 嗯 is as frequent as English uh, partly because it also means "mm-hm", yes, I'm listening.
At the bottom, Italian boh — the shrugging "dunno" — appears only 3.0 times per million. Whether that is because films rarely use it or because subtitlers leave it out, the count cannot say.
The fillers with a second job
| Language | Word | As a filler | The same word also means | Per million words | Rank in its language |
|---|---|---|---|---|---|
| English | like | “it was, like, huge” | the verb and the preposition | 4,059.8 at most | 43 |
| English | well | starts an answer, or a change of tack | the adverb: well done | 2,939.5 at most | 59 |
| Portuguese (Brazil) | então | so, well | then | 2,842.0 at most | 46 |
| French | quoi | you know, at the end of a sentence | what | 2,810.3 at most | 61 |
| German | doch | softens or insists: komm doch, do come | yes, after a negative question; but, after all | 2,442.6 at most | 68 |
| Spanish | bueno | well, OK | good | 2,392.4 at most | 50 |
| Mandarin | 就是 | I mean, that is | is exactly, precisely | 2,035.4 at most | 55 |
| Italian | allora | so, well then | then, at that time | 1,598.2 at most | 86 |
| French | bon | right, well then | good | 1,519.0 at most | 93 |
| Mandarin | 那个 | um, er (nèige) | that, that one | 1,136.7 at most | 97 |
| Portuguese (Portugal) | tipo | like, sort of | a type | 865.1 at most | 151 |
| Spanish | vale | OK, right (in Spain) | it is worth | 662.5 at most | 176 |
| Spanish | tipo | like, sort of | a type; a guy | 619.8 at most | 191 |
| Italian | tipo | like, sort of | a type; a guy | 562.9 at most | 225 |
| Portuguese (Brazil) | tipo | like, sort of | a type | 547.6 at most | 207 |
| Portuguese (Portugal) | pronto | right, OK, there you go (Portugal) | ready | 375.0 at most | 286 |
| German | halt | just, simply: das ist halt so | stop!; holds | 344.0 at most | 315 |
| Spanish | pues | well, so | since, because | 339.9 at most | 296 |
| French | genre | like, sort of | a kind, a sort; a genre; grammatical gender | 300.7 at most | 320 |
| German | eben | just, exactly | flat, level; just now | 193.9 at most | 536 |
| French | ben | well (from bien) | the name Ben | 185.5 at most | 478 |
| Portuguese (Portugal) | pá | mate, like (Portugal) | a shovel, a spade | 65.3 at most | 1,302 |
Read every number in this table as "at most". A few notes on how much of each is likely to be filler:
- English like is mostly the verb and the preposition; the filler is a small part of its 4,059 per million, and no word count can say how small.
- German doch (2,442 per million, number 68) is the hardest of these to separate, because even its other uses — "yes, it is", "but still" — sit close to its job as a softener. Halt and eben, "just, simply", are also ceilings, because halt is "stop!" and eben is "flat" and "just now".
- Tipo appears in four of the lists, as Spanish, Italian and both kinds of Portuguese, and is commonest in European Portuguese at 865 per million. But in every one of them it is also "a type", and in several "a guy", so the count cannot say which language uses the filler most.
- Spanish vale is "OK" in Spain, and "it is worth" everywhere; the subtitles mix Spain with Latin America. Vale against the borrowed OK is measured in OK in other languages.
- Portugal's pá ("mate", "like") is also "a shovel", and Portugal's pronto ("right, there you go") is also "ready" — and pá at least marks a country: it is said 7.3 times as often in Portugal's subtitles as in Brazil's, while Brazil's né is 10.6 times as common in Brazil's. Other differences between the two are in Brazilian vs European Portuguese.
- Mandarin 那个 (nèige), "that", doubles as "um, er" when a speaker is searching for a word, and 就是, "is exactly", as "I mean".
Some of the commonest fillers cannot be counted at all, because they are two words and the lists count one: Spanish o sea ("I mean"), English you know, and German na ja when it is written apart rather than as naja.
Why learners should care
Fillers do real work. They hold your turn while you think (euh, äh, 那个), soften what you are about to say (doch, halt, beh), check that the other person is with you (hein, né), and mark a change of subject (allora, bon, bueno). A learner who uses the English ones in another language — an English um in the middle of French — stands out at once. Swapping in euh costs nothing.
Each course's word list is ranked by how often words are really said, fillers included where they made the cut — the French, Italian, German, Spanish and Portuguese lists among them — and the French vocabulary test shows in two minutes how far down the list you already are. The same corpora are compared on other everyday words in which language says thank you most often, and words with no English equivalent at all are ranked in untranslatable words.
Sources
- Lison, P. & Tiedemann, J. (2016). OpenSubtitles2016: Extracting Large Parallel Corpora from Movie and TV Subtitles. Proceedings of LREC 2016. aclanthology.org
- Dave, H. FrequencyWords: frequency lists from OpenSubtitles 2018, content licensed CC BY-SA 4.0. github.com
- The words counted, what each does and what else it means are written out by hand in
scripts/frequency_extras.py(FILLERS); a word with any other meaning is reported only as a ceiling.