Nine accents, and how each one was checked
Every text-to-speech tool has an accent dropdown. Almost none of them can show you that the voice behind the label actually has the accent. This one measures it.
Common Voice asks each speaker where their English comes from, and the answer is their own. That answer is good enough to sort recordings into groups. It is not good enough to put on a page, because a voice built from a recording can lose the accent on the way through the machine, or pick up one that was never there.
So there is a second step, and it looks at one feature: the r at the end of car, farm and father. American and Canadian speakers curl the tongue for it, and that shape drops the third resonance of the vocal tract by hundreds of hertz. Nothing else in English does that, which makes it one of the few accent features that can be measured in a recording rather than argued about.
It measures r-colouring, not the presence of an r. Scotland is rhotic: a Scottish speaker does say the r in car, but taps or rolls it instead of curling it, and a tap leaves the third resonance where it was. So Scotland is expected to read as having no r-colouring, and was written down that way before anything was measured.
9 places where English is a first language
Every group here is a country or a nation within one, not a dialect. The line under each name is what the measurement found, not what the label claimed.
- United States/r is pronounced/
4 voices · 4 of 4 measured r-coloured
The r at the end of car and father is fully pronounced, and the t in butter and water is tapped so that it lands somewhere near a d. Those two habits do more to mark an American voice than any vowel does. This is the accent the American column of the transcription is written for.
Common Voice label: United States English
- Canada/r is pronounced/
4 voices · 4 of 4 measured r-coloured
Close enough to the United States that the transcription is the same, with one famous exception you can hear: the vowel in about and house starts higher, so the word lands nearer “aboot” than an American would say it. Listen for out and about in the same sentence.
Common Voice label: Canadian English
- England/r is dropped/
2 voices · 1 of 4 measured r-coloured
The r at the end of car is not said at all, and the vowel in bath is the long one of father rather than the short one of cat. This is the accent the British column of the transcription is written for. It is also the loosest group on the site: England holds several accents that are further apart from each other than Australia is from New Zealand, and the source recordings do not separate them.
Common Voice label: England English
- Scotland/r is pronounced/
4 voices · 0 of 3 measured r-coloured
Scotland keeps the r that England dropped, and it rolls or taps it rather than curling it the American way. Vowel length works differently too: the vowels in food and good are much closer together than in any other accent here.
Common Voice label: Scottish English
- Wales/r is dropped/
1 voice · 1 of 2 measured r-coloured
The clearest thing to listen for is not a sound but a shape: the pitch rises and falls across a sentence far more than in England, which is why Welsh English is so often called singing. It is also the smallest group on the site, and the page says how small rather than hiding it.
Common Voice label: Welsh English
- Ireland/r is pronounced/
1 voice · 1 of 2 measured r-coloured
The r is there, as in America, but the th in think and this often moves towards t and d, so think can land near “tink”. The vowel in face stays a single steady sound rather than sliding, which is the quickest way to tell an Irish voice from a Scottish one.
Common Voice label: Irish English
- Australia/r is dropped/
2 voices · 2 of 4 measured r-coloured
No r at the end of car, and the vowel in face and day slides much further than in England, so day drifts towards “die”. Word-final vowels are long and open: listen to the last syllable of Australia itself.
Common Voice label: Australian English
- New Zealand/r is dropped/
4 voices · 0 of 3 measured r-coloured
The short i of kit is the giveaway. In Australia it stays high and clear; in New Zealand it falls back towards the weak vowel this site is named after, so fish and chips is the standard demonstration. Say the two accents back to back and that one vowel does all the work.
Common Voice label: New Zealand English
- South Africa/r is dropped/
3 voices · 0 of 2 measured r-coloured
No r at the end of car, and the short i of kit splits in two depending on the consonants around it, which is a feature South African English shares with no other accent here. The Common Voice label covers Zimbabwe and Namibia as well as South Africa, so the group is named for the region rather than the country.
Common Voice label: Southern African (South Africa, Zimbabwe, Namibia)
The r at the end of car
Each of the 31 voices read two sentences twice. One holds 7 words with an r after a vowel and no r anywhere else. The other has no letter r in it at all, and was built from the same family of vowels so that the only difference between the two is the r. Both recordings were then tracked for the third resonance of the vocal tract, which drops sharply when an English r is curled.
- 13 r-coloured
- 15 not
- 23 of 28 match the label
13 voices came out r-coloured and 15 did not, matching the accent they were labelled with in 23 of 28 cases.
| Voice | Accent | F3 in r words | F3 elsewhere | Reading |
|---|---|---|---|---|
| en-us-1 | United States | 1702 Hz | 2356 Hz | r-coloured |
| en-us-2 | United States | 1878 Hz | 2682 Hz | r-coloured |
| en-us-3 | United States | 1599 Hz | 2223 Hz | r-coloured |
| en-us-4 | United States | 1273 Hz | 1697 Hz | r-coloured |
| en-ca-1 | Canada | 1891 Hz | 2717 Hz | r-coloured |
| en-ca-2 | Canada | 2018 Hz | 2688 Hz | r-coloured |
| en-ca-3 | Canada | 1587 Hz | 2474 Hz | r-coloured |
| en-ca-4 | Canada | 1716 Hz | 2082 Hz | r-coloured |
| en-eng-1 | England | 2751 Hz | 2624 Hz | not r-coloured |
| en-eng-2 | England | 2307 Hz | 2343 Hz | not r-coloured |
| en-eng-3 | England | 2680 Hz | 2619 Hz | not r-coloured |
| en-eng-4 | England | 1963 Hz | 2427 Hz | r-coloured · unlike its label |
| en-sco-1 | Scotland | 2182 Hz | 2520 Hz | no reading |
| en-sco-2 | Scotland | 2560 Hz | 2792 Hz | not r-coloured |
| en-sco-3 | Scotland | 2411 Hz | 2399 Hz | not r-coloured |
| en-sco-4 | Scotland | 2223 Hz | 2215 Hz | not r-coloured |
| en-wal-1 | Wales | 1704 Hz | 2690 Hz | r-coloured · unlike its label |
| en-wal-3 | Wales | 1945 Hz | 2052 Hz | not r-coloured |
| en-ie-3 | Ireland | 2501 Hz | 2443 Hz | not r-coloured · unlike its label |
| en-ie-4 | Ireland | 1788 Hz | 2320 Hz | r-coloured |
| en-au-1 | Australia | 2342 Hz | 2410 Hz | not r-coloured |
| en-au-2 | Australia | 1947 Hz | 2809 Hz | r-coloured · unlike its label |
| en-au-3 | Australia | 2217 Hz | 2191 Hz | not r-coloured |
| en-au-4 | Australia | 1861 Hz | 2441 Hz | r-coloured · unlike its label |
| en-nz-1 | New Zealand | 2704 Hz | 2755 Hz | not r-coloured |
| en-nz-2 | New Zealand | 2360 Hz | 2849 Hz | no reading |
| en-nz-3 | New Zealand | 2287 Hz | 2377 Hz | not r-coloured |
| en-nz-4 | New Zealand | 2466 Hz | 2405 Hz | not r-coloured |
| en-za-1 | South Africa | 2208 Hz | 2536 Hz | no reading |
| en-za-3 | South Africa | 1402 Hz | 1563 Hz | not r-coloured |
| en-za-4 | South Africa | 2131 Hz | 2210 Hz | not r-coloured |
The five voices that are not here
The interesting result is the disagreements. 5 of the 31 voices read against their own accent, and four of those five did it the same way: a voice built from an English, Welsh or Australian recording came out with an American r the speaker never had. The model that clones these voices has heard far more American English than anything else, and it leaks. Those five are not in the picker. Without the measurement they would be, labelled with accents they do not have.
How much of this is the measurement wobbling
Speech generation is not deterministic, so every sentence was generated twice and both takes measured. Two takes of the same voice land 118 Hz apart on average, which is the noise floor of the whole exercise. The line between r-coloured and not is drawn at 300 Hz, and a voice is only given a reading when both of its takes fall on the same side of that line. Three voices straddled it and are recorded as having no reading rather than being pushed to whichever side the average happened to land on.
This measures one feature, not a whole accent. A voice can be r-coloured and still sound nothing like the place it came from, and Scotland shows the reverse: rhotic speech that the measurement is blind to, because a tapped r and a curled r are different sounds. It is a floor, not a certificate.

Hear them
United States
- Kaylee
Woman 212 Hz · clear
- Cheyenne
Woman 184 Hz · full
- Wyatt
Man 122 Hz · full
- Dashiell
Man 112 Hz · full
Canada
- Alanis
Woman 201 Hz · middle
- Geneviève
Woman 184 Hz · full
- Gilles
Man 153 Hz · middle
- Grayson
Man 128 Hz · full
England
- Marjoriechosen
Woman 180 Hz · full
- Ralph
Man 121 Hz · full
Scotland
- Mhairi
Woman 197 Hz · middle
- Eilidh
Woman 191 Hz · middle
- Hamish
Man 122 Hz · full
- Alasdair
Man 110 Hz · full
Wales
- Rhys
Man 104 Hz · deep
Ireland
- Eoin
Man 124 Hz · full
Australia
- Bronte
Woman 216 Hz · clear
- Lachlan
Man 148 Hz · middle
New Zealand
- Aroha
Woman 189 Hz · middle
- Ngaire
Woman 181 Hz · full
- Manaia
Man 164 Hz · middle
- Tamati
Man 117 Hz · full
South Africa
- Thandi
Woman 165 Hz · full
- Sipho
Man 131 Hz · middle
- Jacques
Man 111 Hz · full
What this list is not
Nine groups is not nine dialects. England alone holds Received Pronunciation, Estuary, Scouse, Geordie and Brummie, and the distance between some of those is greater than the distance between Australia and New Zealand. Common Voice does not label accents below the level of a country in numbers that would support separate groups, so neither does this page. Saying so is cheaper than pretending otherwise and being found out by anyone from Liverpool.
Updated 2026-09-05