How many words do you need to speak Portuguese?
Portuguese is two frequency lists, not one. Brazil and Portugal share a written standard closely enough to read each other's newspapers, but the language people speak — and so the words a learner hears first — differs more than that suggests. This post measures how much, using the same kind of text for both.
The short answer: the two lists agree on the core, Brazilian Portuguese has the steeper curve at every point, and the differences that matter to a learner are concentrated in a few dozen very common words, starting with the word for you.
Two curves
Both come from film and television subtitles — the OpenSubtitles 2018 corpus, which has a Brazilian and a European Portuguese section — sampled to the same size so that neither gains from being bigger.
- Portuguese (Brazil)
- Portuguese (Portugal)
| Language | First 100 | First 500 | First 1,000 | First 2,000 | First 5,000 | First 10,000 |
|---|---|---|---|---|---|---|
| Portuguese (Brazil) | 51.6% | 70.2% | 76.8% | 82.8% | 89.7% | 94.0% |
| Portuguese (Portugal) | 49.8% | 68.7% | 75.1% | 81.1% | 88.1% | 92.7% |
The first 1,000 Brazilian words cover 76.8% of the dialogue; the first 1,000 European words cover 75.1%. The gap is widest further out: reaching 95% takes 12,091 Brazilian words and 14,656 European ones.
So Brazilian subtitles lean harder on their commonest words. A likely part of the reason is visible in the next section: Brazilian Portuguese uses one word, você, where European Portuguese spreads the same job across tu, its verb forms and você.
Where the lists agree
Far more than they differ. Of the 100 commonest words in each list, 89 are shared; of the first 1,000, 830. The articles, prepositions, pronouns and core verbs are the same language.
Where they part company
The rest is the interesting part, and a handful of words account for most of it. Each figure below is the word's rank in that country's list — lower means commoner.
| Word | Meaning | Rank in Brazil | Rank in Portugal |
|---|---|---|---|
| você | you | 7 | 59 |
| tu | you (familiar) | 972 | 54 |
| estás | you are (with tu) | 1,408 | 58 |
| pra | for, to (spoken para) | 101 | 859 |
| tá | is (spoken está) | 363 | 2,434 |
| oi | hi | 192 | 1,669 |
| cara | guy, mate | 105 | 365 |
| legal | cool, nice | 291 | 1,872 |
| fixe | cool, nice | 12,749 | 765 |
| celular / telemóvel | mobile phone | 995 / 19,832 | 8,162 / 873 |
| ônibus / autocarro | bus | 1,284 / 20,141 | 11,732 / 1,521 |
| trem / comboio | train | 1,123 / 7,044 | 7,140 / 1,307 |
| time / equipa | team | 976 / 7,398 | 6,467 / 425 |
Three patterns stand out.
The word for you is the biggest single difference. In Brazil você is one of the ten commonest words in the language. In Portugal the familiar tu is far more common than it is in Brazil, and with it come its own verb forms — estás, queres, podes — each a separate entry in a frequency list. That one choice reshapes a whole column of the list.
Brazilian speech shortens. Pra for para and tá for está are spelled out in Brazilian subtitles often enough to rank high; in the European list they are far down. Subtitles follow speech more closely than books do, which is why these forms show up at all.
The everyday nouns are simply different words. Bus, train, phone and team are each two words, and each is common in one country and rare in the other.
What this means for a learner
- Pick one variety and learn its list. For the core — the words the two share — it makes no difference. For you and the dozen words above, it does, and mixing them sounds odd in either country.
- Learning one does not lock you out of the other. The shared core is most of any text. What a frequency list cannot show at all is pronunciation, and that is where the two varieties differ most to the ear.
- Verbavia teaches Brazilian Portuguese. Our Portuguese list and the Portuguese course follow Brazilian usage, with você.
The comparison of all seven languages puts both curves in context, and how many words do you need is the general argument.
Caveats
- Subtitles are one register, close to speech, and many are translations of films made in other languages.
- Ranks here are ranks of forms, not of dictionary words: está and estás are separate entries.
- The two sections of the corpus are not matched film for film. Some of the difference may come from which films were subtitled in each country, not from the language.
Sources
- Lison, P. & Tiedemann, J. (2016). OpenSubtitles2016: Extracting Large Parallel Corpora from Movie and TV Subtitles. Proceedings of LREC 2016. aclanthology.org
- Dave, H. FrequencyWords: frequency lists from OpenSubtitles 2018 (
pt_brandpt), content licensed CC BY-SA 4.0. github.com - Every figure on this page is computed by
scripts/frequency_curves.pyin Verbavia's repository and printed from its output rather than typed.