Your vocal folds make one rough, uninteresting buzz about a hundred times a second. Everything that turns it into words β every vowel, every consonant β is the tube above them changing shape. Type a sentence below and this page does the same job in software: letters to phonemes, phonemes to resonance targets, targets to a pressure wave you can hear.
A rule-based letter-to-sound converter, backed by a dictionary of irregular words. Dashed chips are unvoiced β no buzz at all, just noise. Click any chip to inspect it.
Every phoneme is a set of frequencies the tube has to resonate at. Two of them β F1 and F2 β are enough to place any vowel on the map below, which is why the vowels fall into the shape phoneticians call the vowel quadrilateral.
| Sound | F1 | F2 | F3 | ms |
|---|
Drag any of the nine handles to reshape the pitch contour, then press Play. The bars underneath are the phoneme durations on the same time axis β lit bars are voiced. The one control worth pressing is Flat pitch: same words, no melody.
A glottal pulse train β or white noise for the unvoiced sounds β is pushed through a cascade of three resonators tuned to F1, F2 and F3, sample by sample. Both pictures below are measured from the audio this page just generated, not drawn for effect.
What you are hearing is a formant synthesiser: the same family as Dennis Klatt's 1980 design and the DECtalk boxes built from it, one of which became Stephen Hawking's voice. It describes a human vocal tract with about a dozen numbers per five milliseconds. Real speech carries far more than that β breath noise, the exact shape of each glottal pulse, the way a nasal cavity leaks β so the result is intelligible but unmistakably a machine. Everything a modern phone does better, it does by throwing away this model and predicting the waveform directly. The gap you can hear is roughly the size of the last forty years of the field.
Press anywhere in the mouth below and drag the tongue. You will hear the vowel change, because that is genuinely all a vowel is: two resonances, set by where the highest point of your tongue happens to be. Nothing about the buzz underneath changes at all.
Text-to-speech has been rebuilt from the ground up three times. Each rebuild solved the previous one's characteristic failure and introduced a new one.