Skip to content

The source–filter model of speech

A vowel is a buzz through a tube: build one from a pulse train and three resonators, then recover the resonances by linear prediction.

Before this15.5 · 25.3 · 4 more
Chapter 30 · Lesson 1 of 2

First, the picture

A voice sets its note and its vowel separately. Watch a synthetic voice below sing three vowels on one note: its harmonics stay 125 Hz apart while the peaks of the curve over them move.

Same pitch, different vowels

A synthetic voice at 125 Hz, 8 kHz: an impulse train through a glottal pulse filter, lip radiation and three formant resonators (Peterson and Barney's values).

/a/: formants at 730, 1090 and 2440 Hz; harmonics every 125 Hz.

vowel
/a/
F1, F2, F3 (Hz)
730, 1090, 2440
pitch (Hz)
125
Vowel
0.00 / 13.00 s
Describe this picture

A synthetic voice at 125 Hz, 8 kHz: an impulse train through a glottal pulse filter, lip radiation and three formant resonators with Peterson and Barney’s values, 0.5 s of each vowel scaled to a peak of 1. Two panels. The first is the waveform, from −1.1 to 1.1 against time from 0 to 40 ms, with a tick every 8 ms: forty milliseconds hold five periods of the buzz. The second is the spectrum, from −60 to 5 dB against frequency from 0 to 4000 Hz, taken over the whole 0.5 s with a Hann window, from Window functions compared (15.2). The harmonics are a solid line, scaled so that the tallest sits at 0 dB, and anything lower than −60 dB is drawn on the floor. The envelope, dashed, includes the source’s tilt and the lips as well as the tract. It gets one common gain that makes it pass through the harmonics’ tops, so between harmonics it can rise a little above 0 dB, by up to 4.1 dB for /u/. Dotted vertical lines mark F1, F2 and F3. The readouts are the vowel, its three formants in hertz and the pitch. The 13 s clip opens on /a/: formants at 730, 1090 and 2440 Hz, harmonics every 125 Hz. From 3.5 s it morphs to /i/: the first formant drops to 270 Hz and the second rises to 2290 Hz, while the harmonics stay put. From 8 s it morphs to /u/: both low, 300 and 870 Hz, with the third at 2240 Hz. When the clip ends, three buttons, “/a/”, “/i/” and “/u/”, in a radio group named “Vowel”, switch the vowel; the vowel is kept in the link as vowel.v. A button “Hear it” plays the 0.5 s once.

A buzz shaped by a tube

Sing “ah” and then “ee” on the same note. The note stays and the vowel changes. So one part of your voice sets the note, and another part sets the vowel. This page pulls the two apart.

Air from the lungs passes the glottis, the gap between the vocal folds in the larynx. During a vowel the folds open and close many times a second, letting the air through in puffs. That stream of puffs is the source, a buzz.

The buzz then passes through the vocal tract, the tube of throat, mouth and lips from the larynx to the open air. The tract is a filter with a few resonances. They are called formants, numbered from the lowest up: F1F_1, F2F_2, F3F_3, in hertz.

Moving the tongue and lips changes the tube’s shape, and that moves the formants. A vowel is, to a first approximation, a set of formant positions.

So speech is a source through a filter: the source–filter model, set out by Gunnar Fant in 1960. The source sets the pitch and the filter sets the vowel. They are separate parts, so either can change while the other stays.

Every sound on this page is synthetic, built from the formulas below; nothing is recorded. The formant values come from Peterson and Barney’s 1952 study of American English vowels.

The source: a pulse train with soft puffs

I work at fs=8f_s=8 kHz, the telephone rate. Start with one unit impulse every 64 samples. That is a period of 64/800064/8000 s, which is 8 ms, so the buzz repeats 125 times a second.

In Signals as sums of sinusoids (7.1), “A second arrow, three times as fast” showed that a periodic signal is a sum of harmonics, at whole multiples of its fundamental. Here the fundamental is f0=125f_0=125 Hz, and it is the voice’s pitch. So the train’s spectrum is a comb of lines every 125 Hz.

A pitch of 125 Hz lies within the range of men’s voices: Peterson and Barney’s men averaged 124 to 141 Hz across their vowels. Finding the pitch from a recording is the job of Pitch detection (29.1).

Real vocal folds do not let out sharp clicks: each puff swells and fades. So I shape every impulse with two real poles at 0.95, the filter 1/(1−0.95z−1)21/(1-0.95z^{-1})^2. Its impulse response is (n+1) 0.95n(n+1)\,0.95^n. It rises for 18 samples, about 2.3 ms, and then dies away.

The glottal source xglot[n]x_\text{glot}[n] is the impulse train through this filter. A smooth puff has little fast detail, and the filter’s gain falls steadily with frequency. So the comb’s lines get shorter as they go up.

One more step happens at the lips. The pressure that reaches a listener follows how fast the airflow leaving the lips changes, not the flow itself. So the model adds lip radiation: the first difference 1−z−11-z^{-1}.

Oversampling and noise shaping (11.3), in “What a first difference does to slow and fast wiggles”, showed that its gain rises with frequency. It tilts the source back up a little. With both filters, the source is strongest near 65 Hz and about 26 dB weaker at 4 kHz.

The tract: three resonators

Each formant is a resonance, so I build it from the two-pole resonator of Transfer functions, poles & zeros (16.3), from “Where the pair sits is how h[n] moves”. Formant ii, for i=1,2,3i=1,2,3, gets

11−2ricos⁡θi z−1+ri2z−2.\frac{1}{1-2r_i\cos\theta_i\,z^{-1}+r_i^2z^{-2}}.

The pole angle θi\theta_i sets where it rings. 16.3’s rule f=θfs/2πf=\theta f_s/2\pi turns a formant frequency into an angle: θi=2πFi/fs\theta_i=2\pi F_i/f_s.

The radius rir_i sets the peak’s width. Each formant has a bandwidth BWi\mathrm{BW}_i, its 3 dB width in hertz. I take

ri=e−π BWi/fs.r_i=e^{-\pi\,\mathrm{BW}_i/f_s}.

Why this radius? In Resonators, notches and combs (17.4), “Poles behind the zeros narrow the notch” gave the width rule Δf≈(1−r)fs/π\Delta f\approx(1-r)f_s/\pi. When rir_i is near 1, e−πBWi/fs≈1−π BWi/fse^{-\pi\mathrm{BW}_i/f_s}\approx1-\pi\,\mathrm{BW}_i/f_s. So 1−ri≈π BWi/fs1-r_i\approx\pi\,\mathrm{BW}_i/f_s, and the rule gives a width of BWi\mathrm{BW}_i.

For /a/’s first formant, 730 Hz with a bandwidth of 60 Hz, that gives r1=0.9767r_1=0.9767 and θ1=0.5733\theta_1=0.5733 rad.

The tract is the three resonators one after the other, so their bottoms multiply:

Htract(z)=1A(z),A(z)=∏i=13(1−2ricos⁡θi z−1+ri2z−2).\begin{aligned} H_\text{tract}(z)&=\frac{1}{A(z)},\\ A(z)&=\prod_{i=1}^{3}\big(1-2r_i\cos\theta_i\,z^{-1}\\ &\qquad\qquad+r_i^2z^{-2}\big). \end{aligned}

Multiplied out, A(z)A(z) is one polynomial of degree 6. So the tract is an all-pole filter, the 1/A(z)1/A(z) of Parametric models and linear prediction (25.3). That will matter in the second half of the page.

Here are the three vowels, with Peterson and Barney’s averages for their 33 men. Each vowel is named by the word their speakers read. The columns F1, F2 and F3 give F1F_1, F2F_2 and F3F_3 in hertz.

VowelWordF1F2F3
/a/hod73010902440
/i/heed27022903010
/u/who’d3008702240

I give every vowel the same bandwidths: 60, 90 and 120 Hz for the first, second and third formant. Real bandwidths differ from vowel to vowel; holding them fixed lets only the formant frequencies change.

The whole chain is the train through the glottal filter, the lips and the tract. In z-transforms the factors multiply:

X(z)=Xglot(z) (1−z−1) Htract(z).X(z)=X_\text{glot}(z)\,\big(1-z^{-1}\big)\,H_\text{tract}(z).

On the unit circle the gains multiply too. The source gives the comb of harmonics; the rest gives a smooth envelope over it. So the spectrum of a vowel is harmonics under an envelope, and the envelope’s peaks are the formants.

Same pitch, different vowels

The picture at the top of the page sings “ah”, “ee” and “oo” on one note. Watch the harmonics’ spacing and the envelope’s peaks as the vowel changes.

Watch the harmonics first. They do not move: 125 Hz apart for all three vowels, because the source never changed. Only the envelope moves.

Now the formants. /i/ puts its first two far apart, with envelope peaks at 268.6 Hz and 2292.1 Hz. /u/ puts them close together, at 299.5 Hz and 864.9 Hz.

Those peaks are a few hertz from the set values, 270 and 2290, 300 and 870. In the chain, each resonance’s sloping sides tilt its neighbours’ peaks and pull them slightly. The largest pull is on /a/’s second formant: 1090 Hz shows up at 1083.8 Hz.

A formant need not sit on a harmonic. /a/’s first formant, 730 Hz, falls between the harmonics at 625 and 750 Hz. The comb only shows the envelope where a harmonic is. Here the harmonic at 750 Hz, the nearest to 730 Hz, is the tallest of all.

The waveform tells the same story in time. Each 8 ms period starts with a kick from a puff. Then the resonances ring and fade until the next puff. A higher pitch would bring the kicks closer together; a different vowel changes how each period rings.

The tract, recovered from the sound

Now turn the problem round. Suppose I am handed only the sound. Can I find the formants?

The tract is an all-pole filter, 1/A(z)1/A(z), driven by the source. That is the model 25.3 fits.

“One order at a time” there found A(z)A(z) from a signal’s autocorrelation by the Levinson–Durbin recursion. Used on speech, this is called LPC, for linear predictive coding. The plan has four steps.

Four steps from a frame to formants

1. Pre-emphasis. The source’s falling tilt would take up some of LPC’s poles. So first I lift the high frequencies of the whole vowel with 1−0.97z−11-0.97z^{-1}, a first difference that leaks a little. This is pre-emphasis.

Here it nearly undoes the source’s shape. The glottal filter, the lips and the pre-emphasis together are flat to within 1.7 dB from 125 Hz to 4 kHz. Without pre-emphasis, the source with the lips spans 24 dB over that range.

2. One frame. From the pre-emphasised vowel I take 256 samples, which is 32 ms, from sample 1800, 225 ms into the vowel. I multiply them by a Hamming window, from “Five windows, one trade” in 15.2. Thirty-two milliseconds hold four periods of the buzz.

Why 32 ms? Spectrograms & the STFT (15.5), “Short window or long window”, showed the trade. A longer frame sees finer frequency detail; a shorter one follows faster changes.

A real voice moves its tongue and lips within tens of milliseconds, so speech is analysed in frames of about 20 to 30 ms. My synthetic vowels never change, so here the frame length only has to hold a few periods.

3. Fit A(z)A(z). I estimate R^x[0]\hat R_x[0] to R^x[10]\hat R_x[10] from the windowed frame, as in 25.3, and run Levinson–Durbin to order p=10p=10. That gives A(z)=1+a1z−1+⋯+a10z−10A(z)=1+a_1z^{-1}+\dots+a_{10}z^{-10}.

Why 10? A man’s vocal tract has about one formant per kilohertz. Up to fs/2=4f_s/2=4 kHz that makes about four, and each takes a pole pair: eight poles. Two more cover what is left of the source’s tilt.

The usual rule is p≈fs/1 kHz+2p\approx f_s/1\,\text{kHz}+2, which is 10 at 8 kHz.

My vowels have only three formants, so LPC has poles to spare. They land well inside the circle, as we will see.

4. Read the roots. Each root pkp_k of A(z)A(z) is a pole, ∣pk∣ejθ\lvert p_k\rvert e^{j\theta}. Its angle gives a frequency by 16.3’s rule, f=θfs/2πf=\theta f_s/2\pi. Turning r=e−π BW/fsr=e^{-\pi\,\mathrm{BW}/f_s} round gives its bandwidth:

BW=−fsπln⁡∣pk∣.\mathrm{BW}=-\frac{f_s}{\pi}\ln\lvert p_k\rvert.

A formant’s poles sit close to the circle, so I keep the three roots with a positive angle that lie nearest to it. Those are the formant estimates. On all three vowels the choice is clear: those three lie at radius 0.947 or more, and every other root at 0.495 or less.

One change from the first picture: here the dashed curve is the tract alone, ∣Htract∣\lvert H_\text{tract}\rvert from the set formants, labelled “tract”. Pre-emphasis took the source’s tilt out before the fit, and an all-pole curve cannot follow the lips’ zero at 0 Hz. So the tract is what LPC can recover, and the tract is what it is compared with. The harmonics still carry the source’s shape, so at low frequencies they stand above the dashed curve.

The tract, recovered from the sound

LPC of order 10 on one 32 ms frame of each synthetic vowel.

The vowel /a/: its harmonics, and the tract's three formants.

vowel
/a/
largest error
—
LPC F1, F2, F3
—
Vowel
0.00 / 13.00 s
Describe this picture

LPC of order 10 on one 32 ms frame of each synthetic vowel, in two panels. The first is the spectrum, with the axes of the first picture and the same solid harmonics; its dashed curve is the tract alone, and the LPC envelope, 1/∣A∣1/\lvert A\rvert, is a thick solid line. Every curve is in dB re its own highest peak, and lower than −60 dB is drawn on the floor. The second, square, shows the unit circle, with the ten LPC poles as crosses and the poles of the three set formants as open rings. The readouts are the vowel, the LPC estimates of F1, F2 and F3 in hertz and the largest error in percent. The 13 s clip opens on /a/’s spectrum: its harmonics, and the tract’s three formants. From 3 s the LPC envelope and the poles draw in: LPC finds poles at 743 Hz, 1105 Hz and 2438 Hz, beside the set 730, 1090 and 2440 Hz, a largest error of 1.8 %. From 7.5 s it switches to /i/: 263 Hz, 2276 Hz and 3004 Hz against 270, 2290 and 3010 Hz, 2.5 %. When the clip ends, the three vowel buttons switch the vowel; for /u/ the estimates are 299 Hz, 872 Hz and 2242 Hz against 300, 870 and 2240 Hz, 0.4 %. The vowel is kept in the link as lpc.v.

Look at the /a/ frame. The thick LPC curve has its three peaks beside the dashed tract’s, and three crosses sit almost on the three rings. LPC never saw the formant values; it found them from 256 samples of sound.

The two curves stay close. The largest gap is 3.9 dB for /a/ and 3.4 dB for /i/. For /u/ it is 7.9 dB, at 4 kHz, where the tract is 67 dB down and drawn on the floor; on screen it is 5.6 dB. Within one bandwidth of each formant the gap is at most 5.6 dB.

Across all three vowels, the largest miss is 2.5 %, /i/’s first formant: 263 Hz for a set 270 Hz. Its neighbouring harmonics are at 250 and 375 Hz, and the estimate lies between the set value and the harmonic at 250 Hz.

The misses are small but not zero. One reason is that the model is only nearly right. 25.3 fitted 1/A(z)1/A(z) to white noise through the filter; here the input is a buzz, and the frame shows the envelope only at harmonics, every 125 Hz.

Now the spare poles. For /i/ they are two pairs, at 1139 and 3058 Hz. For /a/ and /u/ they are one pair and two real roots. None of them is near the circle, and none looks like a formant.

What is left after LPC? Filter the pre-emphasised vowel by A(z)A(z) and you get the prediction error of 25.3.

In this frame its four largest samples fall at the four puffs, 64 samples apart, for every vowel. Those four samples hold over 90 % of the error’s energy. What the predictor could not guess is the source’s pulse train.

That split is worth money. A coder can send A(z)A(z), the pitch and a loudness for each frame instead of the samples, and the receiver rebuilds a buzz through 1/A(z)1/A(z). This is an LPC vocoder.

CELP coders, used in phones, send A(z)A(z) plus a coded version of the prediction error. Both need only a few kilobits per second, against 64 kbit/s for plain 8-bit samples at 8 kHz.

The maths behind it · companion-matrix eigenvalues

The formants are eigenvalues. The roots of A(z)A(z) are the eigenvalues of its companion matrix, whose first row is −a1,…,−a10-a_1,\dots,-a_{10} with ones just below the diagonal. That is how the instrument finds them.

The maths behind it · model selection

LPC is the AR model of 25.3 fitted to one frame of speech. Its order is a model-selection choice, as there: too few poles merge formants, and too many fit bumps that are not there.

Worked example

1. A formant’s poles. Take /a/’s first formant, F1=730F_1=730 Hz with BW1=60\mathrm{BW}_1=60 Hz, at fs=8f_s=8 kHz. The radius is r1=e−π⋅60/8000=e−0.02356=0.9767r_1=e^{-\pi\cdot60/8000}=e^{-0.02356}=0.9767. The angle is θ1=2π⋅730/8000=0.5733\theta_1=2\pi\cdot730/8000=0.5733 rad, which is 32.85°.

So the resonator’s bottom is 1−2r1cos⁡θ1 z−1+r12z−21-2r_1\cos\theta_1\,z^{-1}+r_1^2z^{-2}, which is 1−1.6411z−1+0.9540z−21-1.6411z^{-1}+0.9540z^{-2}. Going back, 0.5733⋅8000/2π0.5733\cdot8000/2\pi gives 730 Hz, and −ln⁡(0.9767)⋅8000/π-\ln(0.9767)\cdot8000/\pi gives 60 Hz.

2. A root read back. For /a/, LPC returns a root pk=0.8198+0.5413jp_k=0.8198+0.5413j. Its size is 0.9824 and its angle 0.5836 rad. So the formant estimate is 0.5836⋅8000/2π=7430.5836\cdot8000/2\pi=743 Hz, with a bandwidth of −ln⁡(0.9824)⋅8000/π=45-\ln(0.9824)\cdot8000/\pi=45 Hz.

The frequency misses the set 730 Hz by 1.8 %. The bandwidth, 45 Hz against a set 60 Hz, is further off. On these vowels LPC places a peak more reliably than it measures its width.

3. The order. At fs=8f_s=8 kHz the rule gives p=8000/1000+2=10p=8000/1000+2=10: four formants below 4 kHz at two poles each, plus two for the source. At 16 kHz the same rule gives 18.

Where you’ll meet this

Mobile phones code speech with LPC. GSM’s full-rate coder fits order 8, and AMR and G.729, both CELP coders, fit order 10 to frames of 10 or 20 ms.

Phoneticians track formants through recorded speech with LPC frame by frame. The free program Praat, much used in phonetics, fits an LPC model to each frame to find them.

Formant synthesisers run this page backwards: a source through resonators, with the formants moved over time. Dennis Klatt’s 1980 synthesiser chains resonators like the ones above.

I left out much of speech. Nasal sounds add zeros as well as poles, and consonants such as “s” use noise as the source instead of a buzz. Detailed models of the glottal flow, such as Rosenberg’s and the LF model, and tube models of the tract, are left for the books.

Speech features (30.2) turns frames of speech into the features used to recognise words and speakers. For more, see L. R. Rabiner and R. W. Schafer, Theory and Applications of Digital Speech Processing (2011), chapters 5 and 9; G. Fant, Acoustic Theory of Speech Production (1960); and J. Makhoul, “Linear prediction: a tutorial review” (Proc. IEEE, 1975).

Reference card

QuantityFormulaNotes
Sourcepulses every fs/f0f_s/f_0 samples, through 1/(1−0.95z−1)21/(1-0.95z^{-1})^2sets the pitch f0f_0
Lip radiation1−z−11-z^{-1}gain rises with frequency
Formant1/(1−2ricos⁡θi z−1+ri2z−2)1/(1-2r_i\cos\theta_i\,z^{-1}+r_i^2z^{-2})ri=e−πBWi/fsr_i=e^{-\pi\mathrm{BW}_i/f_s}, θi=2πFi/fs\theta_i=2\pi F_i/f_s
TractHtract(z)=1/A(z)H_\text{tract}(z)=1/A(z)all-pole, sets the vowel
SpeechX=Xglot (1−z−1) HtractX=X_\text{glot}\,(1-z^{-1})\,H_\text{tract}harmonics under an envelope
Pre-emphasis1−0.97z−11-0.97z^{-1}flattens the source’s tilt
LPC orderp≈fs/1 kHz+2p\approx f_s/1\,\text{kHz}+210 at 8 kHz
Formants from LPCf=θfs/2πf=\theta f_s/2\pi of each root pkp_kthe three nearest the circle, angle above 0
Formant bandwidthBW=−fsπln⁡∣pk∣\mathrm{BW}=-\frac{f_s}{\pi}\ln\lvert p_k\rvertinverts r=e−πBW/fsr=e^{-\pi\mathrm{BW}/f_s}

End of lesson 30.1

Where to go next.

Phasorium
LibraryEvery lesson, in order

Parts

About Phasorium
Look