Text to Speech: How It Works
Text-to-speech turns written text into spoken audio using the voices already installed on your device. It runs locally through your browser's speech engine, which means nothing you type is uploaded — and it also means the available voices differ from one machine to another.
How browser speech works
The Web Speech API hands your text to the operating system's own speech engine. macOS, Windows, Android and iOS each ship their own voices, so the same page sounds different on different devices, and a voice available on one is often absent on another. Some browsers also offer higher-quality network voices that require a connection.
The practical consequence: a voice list is a property of your device, not of this page. If a language you need is missing, it can usually be added through your operating system's accessibility or language settings.
The controls
| Setting | Range | Notes |
|---|---|---|
| Rate | 0.1 – 10, default 1 | 0.8–0.9 suits comprehension; regular users often prefer 1.5+ |
| Pitch | 0 – 2, default 1 | Small changes sound natural; extremes sound artificial |
| Volume | 0 – 1 | Independent of system volume |
| Voice | Device-dependent | Determines language and accent |
Experienced screen reader users commonly listen at two to three times normal speed. It sounds incomprehensible at first and becomes natural within days — worth knowing if you intend to use speech for reading rather than proofreading.
Why it is excellent for proofreading
This is the use most people underrate. Reading your own writing silently, your brain supplies what you meant rather than what you wrote — missing words, doubled words and wrong homophones are invisible. A speech engine reads exactly what is on the page, so:
- Missing words leave audible gaps.
- Repeated words are unmistakable.
- Sentences that are too long become obvious when you run out of breath following them.
- Clumsy rhythm and unintended rhymes surface immediately.
It is the single most effective proofreading technique available, and it costs nothing.
Accessibility
Speech output supports people with visual impairment, dyslexia, ADHD and reading difficulties, and anyone whose circumstances make reading impractical. It is worth distinguishing this tool from a screen reader: a screen reader navigates an entire interface, announcing headings, links, form states and landmarks. This reads a block of text aloud. Both are useful; only the former makes an application usable without sight.
Pronunciation limits
Engines guess at unfamiliar words and get proper nouns, technical terms and non-English names wrong regularly. Homographs are the persistent problem — 'read', 'lead', 'live', 'wind' and 'bass' each have two pronunciations and no reliable way to choose between them from spelling alone. Some engines use context; most do not. Adjusting the spelling phonetically is the usual workaround when a specific word must be right.
Privacy
Synthesis happens on your device, so the text is not uploaded when a local voice is selected. Some browsers offer higher-quality voices that do process text remotely — if you are reading anything confidential aloud, choose a voice marked as local.