Lab 23

A voice is a buzz and a shape to push it through

Your vocal folds make one rough, uninteresting buzz about a hundred times a second. Everything that turns it into words β€” every vowel, every consonant β€” is the tube above them changing shape. Type a sentence below and this page does the same job in software: letters to phonemes, phonemes to resonance targets, targets to a pressure wave you can hear.

Try

1Text becomes phonemes

A rule-based letter-to-sound converter, backed by a dictionary of irregular words. Dashed chips are unvoiced β€” no buzz at all, just noise. Click any chip to inspect it.

2Phonemes become resonance targets

Every phoneme is a set of frequencies the tube has to resonate at. Two of them β€” F1 and F2 β€” are enough to place any vowel on the map below, which is why the vowels fall into the shape phoneticians call the vowel quadrilateral.

SoundF1F2F3ms

3Pitch and timing

Drag any of the nine handles to reshape the pitch contour, then press Play. The bars underneath are the phoneme durations on the same time axis β€” lit bars are voiced. The one control worth pressing is Flat pitch: same words, no melody.

4Buzz through filters becomes speech

A glottal pulse train β€” or white noise for the unvoiced sounds β€” is pushed through a cascade of three resonators tuned to F1, F2 and F3, sample by sample. Both pictures below are measured from the audio this page just generated, not drawn for effect.

Waveform
One 25-millisecond slice each spike is one closure of the vocal folds
Mel spectrogram 64 mel bands, 60 Hz to 8 kHz
Sounding now
β€”
Press Play. The panel follows the sentence sound by sound.
This utterance

Why this sounds like 1984, and why that is the point

What you are hearing is a formant synthesiser: the same family as Dennis Klatt's 1980 design and the DECtalk boxes built from it, one of which became Stephen Hawking's voice. It describes a human vocal tract with about a dozen numbers per five milliseconds. Real speech carries far more than that β€” breath noise, the exact shape of each glottal pulse, the way a nasal cavity leaks β€” so the result is intelligible but unmistakably a machine. Everything a modern phone does better, it does by throwing away this model and predicting the waveform directly. The gap you can hear is roughly the size of the last forty years of the field.

The human original

A source, and a filter, in your throat

Press anywhere in the mouth below and drag the tongue. You will hear the vowel change, because that is genuinely all a vowel is: two resonances, set by where the highest point of your tongue happens to be. Nothing about the buzz underneath changes at all.

Nearest vowel
Ι™
Press and drag inside the mouth.
Four ways to build a talking machine

Nobody does it like this any more

Text-to-speech has been rebuilt from the ground up three times. Each rebuild solved the previous one's characteristic failure and introduced a new one.