Skip to the tool

What Schwa is, and who makes it

Updated 5 September 2026

A free site that will not say what it lives on does not deserve trust. Here is the whole answer.

What Schwa does

It takes the text you paste and gives you two things: the same text written in the International Phonetic Alphabet, and an MP3 of it being read aloud. There are 25 English voices grouped into 9 accents, native-speaker voices for 64 more languages, and the option to build your own voice from a recording of three to eight seconds.

The transcription is written twice, in the two spellings dictionaries use: Received Pronunciation for British and General American for American. It is produced from the whole sentence rather than word by word, which is why words that are spelled one way and said two ways usually come out right.

A request can hold up to 500 words, with 5 seconds between two of them. There is no paid tier, nothing to upgrade to, and no watermark on the file.

Where the voices come from

From Common Voice, Mozilla’s open speech collection, where people read sentences aloud as volunteers. All of it is CC0 1.0, the licence that comes closest to giving up copyright entirely. No permission and no credit are required.

The English part of Common Voice is the largest of any language: 1,799,288 recordings from 60,649 people. That size is why this site can group voices by accent at all. Of the 22,879 speakers with enough clean recordings, 8,667 had stated where their English comes from.

From each chosen person the clearest recording was taken, measured as the distance between their voice and the room noise behind it. One person, one voice. None of these are the same voice at a different pitch.

How the accents were checked

Every speaker in Common Voice states their own accent. That is good enough to sort recordings, and not good enough to print on a page, because a voice rebuilt from a recording can lose the accent on the way through the machine.

So each voice reads a sentence packed with words that have an r after a vowel, alongside a control sentence with no r in it, and both are examined for the third resonance of the vocal tract, which drops when an English r is curled. 13 voices came out r-coloured and 15 did not.

5 of the 31 voices measured read against their own accent, most of them by acquiring an American r the speaker never had. Those are not on the site.

This measures one feature, not a whole accent, and it measures r-colouring rather than the presence of an r, so a tapped Scottish r does not show up. It is a floor rather than a certificate, and the page says so where the numbers are.

How the transcription is produced

The phonetic script comes from eSpeak NG, an open-source speech engine that has carried English pronunciation rules for over twenty years, plus a normalisation layer written for this site. eSpeak places the stress mark next to the vowel; dictionaries place it at the start of the syllable. Eight such differences are corrected, and each one is listed with its reason on the transcription page.

None of that is self-graded. The American transcription was compared against CMUdict, a pronunciation dictionary built at Carnegie Mellon and maintained by people with no connection to this site: they agree on 78.2% of the most common English words.

Why it is free

Because the two expensive parts are already paid for. The recordings are public domain, and the machine that produces the speech is a computer that already exists and does other work. What this site adds is electricity and time.

There are no advertisements, no tracking, no third-party scripts and nothing to sell. The limits on the page exist so that one person cannot use up the machine for everyone else, not to push anyone towards a paid version. There is no paid version.

What it does not do

It has no accent below the level of a country. England is one group here, although Received Pronunciation, Scouse and Geordie are not one accent. The source recordings do not separate them in usable numbers.

Indian and South Asian English is not on the site, and that is a decision rather than an omission: it is the third largest group in the corpus, ahead of six accents that are included. The line drawn here is communities where English is the first language of the community, because the two phonetic spellings on offer are British and American.

The transcription is a rule system, not a dictionary of every word. Rare names and new coinages are guessed from spelling, sometimes wrongly.

Under heavy load a request is refused rather than queued, because a browser that waits too long is cut off before it hears anything.