How many Māori words do you need to know?
Fewer than in most languages you might compare it with. In the largest published count of spoken Māori, just 165 different words made up about 80% of everything said, where Spanish needs about 1,800 to cover the same share of film dialogue. In our own measurement, the first 1,000 words of Verbavia’s Māori list cover 76% of the running text of Māori Wikipedia and 94% of the course’s reading and listening passages.
Tap 🔊 to hear an example. Māori examples play Papa Reo, the te reo Māori voice made by Te Hiku Media from recordings of native speakers.
The short answer
A few hundred words take you a long way in te reo Māori, and a couple of thousand take you close to the point where you can follow ordinary speech. The reason is not that Māori has a small vocabulary. It is that Māori does its grammar with short, separate words that turn up in every sentence, and a coverage count counts every one of them.
This is how much of real Māori text the words on our Māori frequency list cover, learned in the list’s own order:
| Words learned | Māori Wikipedia | The course’s own passages |
|---|---|---|
| The first 100 | 50.5% | 68.1% |
| The first 500 | 70.2% | 89.6% |
| The first 1,000 | 76.4% | 94.1% |
| The whole list (2,002) | 80.9% | 96.2% |
Read the two columns as a range rather than as one answer. Wikipedia is an encyclopedia, full of place names, dates and the vocabulary of whatever an article is about, so it is a hard test. The passages are everyday Māori, but they were written for the course from the same word list, so they are an easy one. Ordinary conversation sits somewhere between. How both were measured, and everything that could bend the figures, is set out further down.
Why so few words go so far
Take one sentence from the course and look at what its words are doing:
I haere a Mere ki te toa i te ata me tōna hoa. ‘Mere went to the shop in the morning with her friend.’
Seven of its thirteen words are particles: i marks the past, a comes before a person’s name, ki means to, te (twice) means the, the second i means in or at, and me means with. Every phrase in a Māori sentence opens with a particle, and a few dozen of them do that work for the whole language. The words that carry the subject matter, here haere ‘go’, toa ‘shop’, ata ‘morning’ and hoa ‘friend’, are a minority of what you hear.
The second reason is that a Māori verb does not change its shape for tense or for who is doing it. The particle in front does that job:
| Māori | English | The marker |
|---|---|---|
| Kei te kai au. | I’m eating. | kei te: going on now |
| I kai au. | I ate. | i: it happened |
| Kua kai au. | I’ve eaten. | kua: a new state reached |
| Ka kai au. | I eat, or I’ll eat. | ka: it happens, or will |
One verb, kai ‘eat’, and one form of it. Spanish puts the same information inside the verb, so a learner meets como, comí, comeré and he comido, and a frequency list counts each as a separate word. Māori spends its words on the particles instead, and because the same particles serve every verb, they pile up at the very top of any count.
The main exception is the passive. A verb used in the passive takes an ending, and the ending has to be learned with the verb: kai becomes kainga, tuhi ‘write’ becomes tuhia, hoko ‘buy’ becomes hokona. Māori uses the passive far more than English does, and passive forms are among the commonest words that a list of base words misses.
The most common Māori words
The very top of any Māori count is the same short run of particles. This is how much of each text these words make up on their own, with their rank on our list and the meaning the list gives them:
| Word | What it does | Our rank | Wikipedia | Passages |
|---|---|---|---|---|
| te | the (one thing) | 18 | 11.5% | 11.7% |
| i | past-tense marker; at, in; marks the object | 21 | 5.4% | 7.7% |
| o | of, belonging to | 46 | 5.5% | 0.7% |
| ko | topic marker: it is, as for | 17 | 4.1% | 1.5% |
| ki | to, towards, at, into | 44 | 2.8% | 2.2% |
| ngā | the (more than one) | 26 | 2.7% | 2.1% |
| he | a, some | 19 | 2.5% | 3.1% |
| e | with ana, an action going on | 24 | 1.8% | 2.6% |
| a | before a person’s name; of | not on the list | 1.7% | 1.5% |
| ka | an action that happens, or will | 22 | 1.1% | 2.8% |
| kei | at (now); with te, going on now | 20 | 0.5% | 2.5% |
| kua | a new state reached: has, have | 23 | 0.2% | 1.0% |
te alone is about one word in nine in both texts, and with ngā it is one in seven. A set of 23 particles of this kind accounts for 45% of both, and the figure comes out the same in an encyclopedia and in everyday passages. That is the clearest sign that it belongs to the language rather than to either text.
The two columns part company where you would expect. Wikipedia describes things, so it is full of o ‘of’ and ko ‘it is’; the passages tell stories about people, so they lean on ka, kei te and kua, the markers of what is happening and what has happened.
Our list does not start with these words. It starts with kia ora ‘hello, thanks’, mōrena ‘good morning’ and the words for I and you, because on day one you need to say something; Māori greetings covers those first words properly. The particles follow within the first fifty.
How Māori compares with Spanish and French
The fairest comparison uses each language’s own counts: rank the words of a large text by how often they occur, and ask how many of the commonest you need to cover 80% of it. These figures come from Mary Boyce’s Māori Broadcast Corpus, about a million words transcribed from Māori-language radio and television in 1995 and 1996, and from the film and television subtitles behind our comparison of seven languages:
| Language and text | Words for 80% of the text |
|---|---|
| Māori, broadcast speech (Boyce, 2006) | 165 |
| Māori, Wikipedia (our measurement) | 381 |
| French, subtitles | 1,076 |
| German, subtitles | 1,358 |
| Spanish, subtitles | 1,845 |
Even Māori Wikipedia, where an encyclopedia’s vocabulary should make the count harder, needs fewer words to reach 80% than French dialogue does. The gap is too large to be an accident of the texts. The corpora do differ in size, register and date, but size is not what explains it: in the subtitle study, cutting each language down to a sample smaller than Boyce’s barely moved these numbers (Spanish went from 1,845 words to 1,808).
What the gap does not mean is that Māori can be learned in 165 words. It is largely a difference in what gets written as a word. Spanish fuses its grammar into its verbs and multiplies the number of forms; Māori writes its grammar as separate particles and keeps the number down. The work has not gone away. It has moved into knowing exactly what i, ka and kua are doing in a sentence, which is the subject of a good part of any course, and the reason is te reo Māori hard to learn? gives a more interesting answer than the sounds alone would suggest.
What 76% feels like, and the 98% line
Coverage counts running words, so 76% means about one word in four is one you do not know. Research on reading puts comfortable, unassisted understanding at around 98% of the words in a text, which is one unknown word in fifty; how many words do you need to learn a language sets out where that figure comes from and how firm it is.
For Māori the 98% line is closer than for most languages. In Boyce’s broadcast corpus the commonest 2,000 words covered 97.6% of the text. On our own measures the whole 2,002-word list covers 96% of the course’s passages, and what it misses there is mostly people’s and places’ names, the days of the week, and passive forms such as horoia ‘wash, in the passive’ and karangahia ‘call, in the passive’. On Wikipedia it stops at 81%, and the rest is the encyclopedia’s own world: iwi and place names, the names of countries, and words like taupori ‘population’ that a learner meets late or never.
Where our list falls short
Measuring a list against real text shows its gaps, and ours has three worth naming.
- The commonest word missing from it is a, which stands before a person’s name (as in a Mere above) and is also one of the words for of. It is between one word in sixty and one in seventy of both texts. The lessons teach it; the list has no entry of its own for it.
- The particle ai is missing too. It is the ai of e ai ki ‘according to’, and it is in the top fifty of the Ministry of Education’s list described below.
- Passive forms are counted as separate words, and only some of them have entries.
Our first 100 words cover 50.5% of Wikipedia; Wikipedia’s own commonest 100 cover 65.8% of it. Two points of that difference are a and ai. The rest is the order of the list, which puts greetings, pronouns and everyday nouns such as kurī ‘dog’ and ngeru ‘cat’ ahead of words an encyclopedia uses constantly, such as reo ‘language’, iwi ‘tribe’ and tāone ‘town’.
How we measured it
Everything above was computed by a script that anyone can re-run, scripts/mi_coverage.py in Verbavia’s repository, and the method is deliberately strict.
- A word counts only if it is spelt exactly as the list spells it, macrons included, ignoring capital letters. keke ‘cake’ and kēkē ‘armpit’ are different words, so an unmacronned spelling is not counted as known.
- Phrases on the list lend nothing to their words. A text is counted one word at a time, so an entry such as kia ora cannot match it whole. A word counts only if the list has it as an entry of its own. For the common phrases this costs nothing, since kia, ora, tēnā and koe are all entries. Crediting the words inside phrases would add two points to the whole list’s Wikipedia figure, almost all of it the word a, which the list holds only inside place names such as Te Whanganui-a-Tara ‘Wellington’. We did not count it.
- Māori Wikipedia is the article text of the whole of mi.wikipedia.org, 8,117 articles, from Wikimedia’s public dump. A large part of it was produced from templates: settlement stubs that repeat one sentence with a new place name, hundreds of plant-species stubs that differ only in the Latin name, and year and calendar pages. We dropped every sentence whose pattern, with names and numbers masked, appears in 20 or more articles. That removed 24,784 sentences and left 209,957 words, of which 183,597 are spelt like Māori words; English and other non-Māori words were left out by a spelling test. Left in, the templates would make the list look better than it is: its first 1,000 words would cover 80% rather than 76%.
- Wikipedia is not careful with macrons. It writes ngā without its macron about one time in seven, along with other common words. If macrons are ignored, the whole list covers 85% of it rather than 81%.
- The course’s passages are its 351 hand-written reading and listening texts, 15,960 words in all, up to the end of B1. They are not an independent test: they were written from the same list.
- Wikipedia’s own order was measured fairly: words were ranked on half of the articles and tested on the other half, both ways round, because ranking a small text on itself flatters it.
The Ministry of Education’s list
New Zealand’s Ministry of Education publishes a list of the 1,000 most frequent words of Māori for teachers, drawn from two corpora of Māori, one of them Boyce’s broadcast corpus. It is in frequency order, it leaves out proper nouns, and it notes that the first 360 or so words matter most to learners because they occur so often. Its first ten words are all particles, with te at the top. It is Crown copyright and its use is restricted to the New Zealand education sector, so we link to it rather than reproduce it; the link is in the sources below.
Our list was built for a different job. The Ministry’s is a reference for teachers; ours is the order in which a learner meets words in a course, which is why it opens with greetings and gets to the particles a few dozen words later. Both agree on what matters most.
What to do with this
- Learn the particles early and properly. Together they are close to half of all running Māori, and knowing what each one does is a large part of Māori grammar.
- Then go down the list in order. The first 500 words take you from half of Wikipedia to 70% of it, and the first 1,000 are on one public page, with no account needed.
- Learn each verb’s passive with the verb. Nothing tells you what it will be.
- Learn the sounds before the words, because a word you cannot say is hard to remember; the Māori pronunciation guide covers how the words are said, macrons included.
You also start with more than you think if you live in New Zealand. Many Māori words are part of everyday New Zealand English, though research has found that people who do not speak Māori can recognise far more Māori words than they can define; Māori words in New Zealand English covers the ones you probably know already. Numbers are among the first content words on the list, from tahi ‘one’ and rua ‘two’ up, and Māori numbers, days and months covers them.
Verbavia’s te reo Māori course follows the order of this list, with the grammar of the particles taught alongside it, and every lesson is written by hand from your first word to the end of B1.
Sources
- Ministry of Education (2010). 1000 frequent words of Māori, in frequency order. Te Whakaipurangi Rauemi. Crown copyright; copying restricted to the New Zealand education sector.
- Boyce, M. T. (2006). A Corpus of Modern Spoken Māori. PhD thesis, Victoria University of Wellington.
- Degani, M. (2012). Language Contact in New Zealand: A Focus on English Lexical Borrowings in Māori. Academic Journal of Modern Philology 1, 13–24 (describes the Māori Broadcast Corpus and quotes its coverage figures).
- Oh, Y. M., Todd, S., Beckner, C., Hay, J. and King, J. (2023). Assessing the size of non-Māori-speakers’ active Māori lexicon. PLOS ONE 18(8).
- Hu, M. and Nation, I. S. P. (2000). Unknown vocabulary density and reading comprehension. Reading in a Foreign Language 13(1).
- Wikimedia Foundation. Māori Wikipedia database dump (CC BY-SA), used for the coverage measurement.
- Te Aka Māori Dictionary: a (particle).
- Te Aka Māori Dictionary: e ai ki.