Skip to content

Filter banks

Split a signal into two bands at half the rate each and rebuild it exactly, because the aliases cancel; then DFT banks and the MDCT.

Before this15.6 · 22.4 · 7 more
Chapter 23 · Lesson 1 of 2

First, the picture

Here one tone, cos⁡(0.30πn)\cos(0.30\pi n), is split into a low band and a high band, each band keeps every 2nd sample, and the two are rebuilt and added. Watch the lines at 0.70π: each band has an alias there, but in the sum they cancel.

Two bands, two aliases that cancel

cos(0.30πn) through a two-channel db2 bank: split, keep every 2nd, rebuild, add.

The two db2 filters: at 0.30π the low one's gain is 0.94 and the high one's 0.33; they cross at 0.5π.

alias, low band
not yet
alias, high band
not yet
alias in the sum
not yet
0.00 / 14.00 s
Describe this picture

Four stacked strips against Ω from 0 to π rad/sample, for cos⁡(0.30πn)\cos(0.30\pi n) through a two-channel db2 bank. The first shows the two filters’ gains, from 0 to 1.1: ∣H0∣/2\lvert H_0\rvert/\sqrt2 as a solid curve labelled “low” and ∣H1∣/2\lvert H_1\rvert/\sqrt2 as a dashed one labelled “high”. At 0.30π they are 0.94 and 0.33, and they cross at 0.5π. The next three strips, “low band, rebuilt”, “high band, rebuilt” and “sum”, show amplitude from 0 to 1.1. In each the tone is a line at 0.30π with a dot at its top, and the alias a line at 0.70π with an open ring, marked ”+” or ”−” for its sign. The readouts are the alias in the low band, in the high band and in the sum, to 3 decimals. The strips fill in turn. The low band holds the tone at 0.890 and an alias of 0.313, and the high band a tone of 0.110 and an alias of 0.313 with the opposite sign. Last the two alias lines slide into the sum and wipe each other out: the tones add to 1.000, the alias reads 0.000, and the output is the input 3 samples late. After the clip a slider named “High-band gain” scales the high band by gg from 0 to 1, and the caption reads like “High band × 0.50: 0.156 of alias is left, and the tone is 0.945.”

Two bands, two aliases that cancel

An audio coder spends few bits on the bands you hardly hear and more on the ones you do. To do that, it first splits the sound into bands. This page is about splitting a signal into bands and putting it back together, without losing anything.

A bank of filters that splits a signal is an analysis bank. Here it has two filters: a low-pass H0(z)H_0(z) and a high-pass H1(z)H_1(z). Each band now holds about half the frequencies, so by Downsampling and decimation (22.1) it can keep every 2nd sample. I write that step as ↓2.

The two bands together then have as many samples per second as the input had. A bank that keeps the total rate the same is critically sampled. Nothing is wasted, and that matters when every sample costs bits.

The synthesis bank goes back. Each band is upsampled with a zero after every sample, as in Upsampling and interpolation (22.2); I write that as ↑2. Then each band goes through its own filter, F0(z)F_0(z) or F1(z)F_1(z), and the two are added.

x[n]H₀(z)H₁(z)↓2↓2lowhigh↑2↑2F₀(z)F₁(z)+y[n]analysissynthesis
Fig. A two-channel filter bank. Between ↓2 and ↑2 each band runs at half the rate, so the two bands together carry as many samples as x[n].

There is a catch. A band is not cut off cleanly at π/2\pi/2, so by 22.1 keeping every 2nd sample aliases. Let’s see exactly what ↓2 followed by ↑2 does. It keeps the even samples and puts zeros at the odd ones, so it multiplies the signal by

12(1+(−1)n),\tfrac12\big(1+(-1)^n\big),

which is 1 for even nn and 0 for odd nn.

Properties and the inverse z-transform (16.2) showed that multiplying by ana^n turns X(z)X(z) into X(z/a)X(z/a). With a=−1a=-1, (−1)nx[n](-1)^nx[n] has the transform X(−z)X(-z). So the pair ↓2, ↑2 turns V(z)V(z) into

12[V(z)+V(−z)].\tfrac12\big[V(z)+V(-z)\big].

On the unit circle, −z=ej(Ω+π)-z=e^{j(\Omega+\pi)}, so V(−z)V(-z) is the spectrum slid by π\pi. That is the frequency shift of the section “Slide the spectrum round the circle” in Properties of the DTFT (12.3).

Try it on one tone. Since (−1)ncos⁡(0.3πn)=cos⁡(0.7πn)(-1)^n\cos(0.3\pi n)=\cos(0.7\pi n), the pair turns cos⁡(0.3πn)\cos(0.3\pi n) into 12cos⁡(0.3πn)+12cos⁡(0.7πn)\tfrac12\cos(0.3\pi n)+\tfrac12\cos(0.7\pi n). Half the tone stays, and half of it appears at 0.7π0.7\pi, where there was nothing. That new line is the alias.

Now follow both bands through the bank. The low band is H0(z)X(z)H_0(z)X(z) before ↓2, and after ↑2 and F0F_0 it is 12F0(z)[H0(z)X(z)+H0(−z)X(−z)]\tfrac12F_0(z)\big[H_0(z)X(z)+H_0(-z)X(-z)\big]. The high band is the same with 1 in place of 0. Add the two and sort the terms:

Y(z)=Tdist(z) X(z)+Talias(z) X(−z),\begin{aligned} Y(z)&=T_\text{dist}(z)\,X(z)\\ &\quad+T_\text{alias}(z)\,X(-z), \end{aligned}

with

Tdist(z)=12[F0(z)H0(z)+F1(z)H1(z)],Talias(z)=12[F0(z)H0(−z)+F1(z)H1(−z)].\begin{aligned} T_\text{dist}(z)&=\tfrac12\big[F_0(z)H_0(z)\\ &\qquad+F_1(z)H_1(z)\big],\\ T_\text{alias}(z)&=\tfrac12\big[F_0(z)H_0(-z)\\ &\qquad+F_1(z)H_1(-z)\big]. \end{aligned}

TdistT_\text{dist} is the distortion transfer function. It acts on the signal itself, like the H(z)H(z) of Transfer functions, poles & zeros (16.3). TaliasT_\text{alias} is the alias transfer function, and it acts on the slid copy X(−z)X(-z).

So each band has an alias term of its own. But we choose F0F_0 and F1F_1, so we can make the two bands’ aliases cancel: Talias(z)=0T_\text{alias}(z)=0. That is aliasing cancellation.

If, as well, Tdist(z)=z−n0T_\text{dist}(z)=z^{-n_0}, the output is the input delayed by n0n_0 samples. That is perfect reconstruction.

The pair I’ll use is the Daubechies pair of 4 taps, called db2. Its low-pass taps are

h0[n]=142(1+3, 3+3,3−3, 1−3)=0.48296, 0.83652,0.22414, −0.12941.\begin{aligned} h_0[n]&=\tfrac1{4\sqrt2}\big(1+\sqrt3,\ 3+\sqrt3,\\ &\qquad3-\sqrt3,\ 1-\sqrt3\big)\\ &=0.48296,\ 0.83652,\\ &\qquad0.22414,\ -0.12941. \end{aligned}

The high-pass reverses them and flips every other sign, h1[n]=(−1)nh0[3−n]h_1[n]=(-1)^nh_0[3-n], which gives −0.12941, −0.22414, 0.83652 and −0.48296. The synthesis filters are the two reversed: f0[n]=h0[3−n]f_0[n]=h_0[3-n] and f1[n]=h1[3−n]f_1[n]=h_1[3-n].

The taps of h0h_0 add up to 2\sqrt2, so its gain at Ω=0\Omega=0 is 2\sqrt2, not 1. The picture at the top of the page divides both gains by 2\sqrt2, so that each pass band reads 1.

That picture sends one tone, cos⁡(0.30πn)\cos(0.30\pi n), through this bank. It shows each band rebuilt on its own, and then their sum.

Notice that the alias is not small in either band: 0.313 of a tone of size 1. It is only the sum that is 0.

It is like two clerks who copy the same column of figures. One writes a figure 5 too high and the other writes it 5 too low. Each copy is wrong, but their sum is right.

After the clip, use the slider to scale the high band by a gain gg from 0 to 1, as an equaliser would.

The tone in the sum is the low band’s 0.890 plus gg times the high band’s 0.110. The alias is (1−g)×0.3128(1-g)\times0.3128, because only the fraction gg of the high band’s alias is there to cancel the low band’s. Try 0: the alias is back to 0.313, and the tone is 0.890.

So the cancellation is a balance. Process one band, and part of the alias comes back. An audio coder that stores one band more coarsely than the other pays this price too.

Why the aliases cancel

Reversing the taps of a filter turns H(z)H(z) into H(z−1)H(z^{-1}): put m=−nm=-n in ∑nh[−n]z−n\sum_nh[-n]z^{-n} and it becomes ∑mh[m]zm\sum_mh[m]z^{m}. The reversed 4 taps then need a delay of 3 to start at n=0n=0 again, a factor z−3z^{-3}. So F0(z)=z−3H0(z−1)F_0(z)=z^{-3}H_0(z^{-1}) and F1(z)=z−3H1(z−1)F_1(z)=z^{-3}H_1(z^{-1}). Put m=3−nm=3-n in the transform of h1[n]=(−1)nh0[3−n]h_1[n]=(-1)^nh_0[3-n] and you get

H1(z)=−z−3H0(−z−1).H_1(z)=-z^{-3}H_0(-z^{-1}).

Put that into the two terms of TaliasT_\text{alias}:

F0(z)H0(−z)=z−3H0(z−1)H0(−z),F1(z)H1(−z)=−z−3H0(−z)H0(z−1).\begin{aligned} &F_0(z)H_0(-z)\\ &\quad=z^{-3}H_0(z^{-1})H_0(-z),\\ &F_1(z)H_1(-z)\\ &\quad=-z^{-3}H_0(-z)H_0(z^{-1}). \end{aligned}

They are equal and opposite for every zz, so Talias(z)=0T_\text{alias}(z)=0. That holds for every tone, not only 0.30π.

On the unit circle, Fk(ejΩ)F_k(e^{j\Omega}) is e−j3Ωe^{-j3\Omega} times the conjugate of Hk(ejΩ)H_k(e^{j\Omega}), because the taps are real (12.3, “The rules that carry over”). So TdistT_\text{dist} there is 12e−j3Ω(∣H0∣2+∣H1∣2)\tfrac12e^{-j3\Omega}\big(\lvert H_0\rvert^2+\lvert H_1\rvert^2\big). Daubechies chose h0h_0 so that ∣H0∣2+∣H1∣2=2\lvert H_0\rvert^2+\lvert H_1\rvert^2=2 at every Ω\Omega, and then Tdist(z)=z−3T_\text{dist}(z)=z^{-3}.

That is why the tones in the instrument, 0.890 and 0.110, add to 1.

The simplest such pair is the Haar pair, h0=(1,1)/2h_0=(1,1)/\sqrt2 and h1=(1,−1)/2h_1=(1,-1)/\sqrt2. It has 2 taps, so its delay is 1 sample. Its split is gentle: at 0.30π its low and high gains are 0.89 and 0.45. So each band’s alias is bigger, 0.405 at this tone, and the two still cancel.

The name QMF, quadrature mirror filter, comes from an older and simpler choice: H1(z)=H0(−z)H_1(z)=H_0(-z), so the high-pass is the low-pass slid by π\pi. With F0=H0F_0=H_0 and F1=−H1F_1=-H_1, the alias term is 12[H0(z)H0(−z)−H0(−z)H0(z)]=0\tfrac12[H_0(z)H_0(-z)-H_0(-z)H_0(z)]=0. The aliases cancel here too.

But perfect reconstruction fails. Split H0H_0 into its two polyphase components, as in Polyphase structures (22.4): H0(z)=E0(z2)+z−1E1(z2)H_0(z)=E_0(z^2)+z^{-1}E_1(z^2). Then

Tdist(z)=12[H0(z)2−H0(−z)2]=2z−1E0(z2)E1(z2).\begin{aligned} T_\text{dist}(z)&=\tfrac12\big[H_0(z)^2-H_0(-z)^2\big]\\ &=2z^{-1}E_0(z^2)E_1(z^2). \end{aligned}

For an FIR filter that is one delay only if E0E_0 and E1E_1 each have a single tap. So the classic QMF reconstructs perfectly only when H0H_0 has 2 taps, as Haar’s does. Longer QMF designs, such as Johnston’s, get close and accept a small ripple in TdistT_\text{dist}.

There is a cost to db2 as well. Its taps are not symmetric, so by the section “Why symmetric taps give a straight phase” of Linear-phase systems (17.3), its phase is not a straight line. Among pairs built the db2 way, only Haar has symmetric taps. Banks that need a straight phase use different filters for analysis and synthesis, called biorthogonal pairs.

The maths behind it · orthogonal matrices

An orthogonal two-channel bank such as db2 is an orthogonal matrix. Its rows are copies of h0h_0 and h1h_1 shifted by 2 samples, and each row has length 1 and is perpendicular to every other. Synthesis is the transpose, so reconstruction is Q⊤Q=I\mathbf{Q}^\top\mathbf{Q}=\mathbf{I}. The DFT bank of the next section is the DFT matrix of The DFT as a matrix (13.5), with a window in front.

One prototype, eight channels, a perfect sum

Two bands are often not enough. A uniform bank splits the band from 0 to 2π2\pi into MM equal channels. The easy way to make it uses one low-pass, the prototype p[n]p[n], and shifts it to each centre.

By 12.3, multiplying by ejΩ1ne^{j\Omega_1n} slides a spectrum by Ω1\Omega_1. So channel kk has the taps p[n] ej2πkn/Mp[n]\,e^{j2\pi kn/M}, centred at 2πk/M2\pi k/M. This is a DFT filter bank: the arrows ej2πkn/Me^{j2\pi kn/M} are the ones in the DFT.

My prototype is a Hann window of 16 samples, the periodic one of Inverting the STFT (15.6). It is scaled so its taps add to 1, so its gain at Ω=0\Omega=0 is 1: p[n]=w[n]/8p[n]=w[n]/8. With M=8M=8, the channels sit 2π/8=0.25π2\pi/8=0.25\pi apart.

These eight channels are the STFT of Spectrograms & the STFT (15.5) in disguise. Channel kk‘s output at sample nn weights the 16 samples ending at nn by the window and by an arrow turning at 2πk/82\pi k/8, and adds them. That is one cell of the STFT of the section “A row of short spectra”, at frequency 2πk/82\pi k/8. Its window is pp back to front, its frame ends at nn, and only an arrow of size 1 in front differs.

One prototype, eight channels, a perfect sum

A 16-sample Hann low-pass shifted to the centres 2πk/8, k = 0 to 7.

The prototype, channel 0: a 16-sample Hann window used as a low-pass, gain 1 at 0.

channels
1
their complex sum
not yet
0.00 / 13.00 s
Describe this picture

One panel: gain from 0 to 1.15 against Ω from 0 to 2π rad/sample, for a 16-sample Hann low-pass shifted to the centres 2πk/8, k = 0 to 7. Each channel’s gain is a thin curve labelled “k = 0” to “k = 7” at its peak. The line styles cycle through solid, dashed, dotted and dash-dot, so channels kk and k+4k+4 share a style. The readouts count the channels and give their complex sum. The clip draws the prototype, channel 0, with gain 1 at 0, then channels 1 to 7, each shifted by 2π/8 = 0.25π from the last. Each channel is 1 at its own centre, 0.5000 halfway to the next, and 0 at the next centre. Last a thick solid line labelled “all eight added” draws at gain 1 everywhere, and the sum reads “a delay of 8 samples”.

Watch the thick line at the end: added as complex numbers, the eight channels are exactly a delay of 8 samples, so nothing is lost.

Notice that the channels overlap: the magnitudes of the eight gains add to between 1.0000 and 1.0620, not to 1. Added as complex numbers, with their phases, they give exactly 1 at every Ω\Omega.

It is like eight spotlights on a stage. Their beams overlap, but added together they light the stage evenly.

Why exactly a delay? Add the taps of the eight channels:

∑k=07p[n] ej2πkn/8=p[n]∑k=07ej2πkn/8.\sum_{k=0}^{7}p[n]\,e^{j2\pi kn/8}=p[n]\sum_{k=0}^{7}e^{j2\pi kn/8}.

When 8 divides nn, every arrow is 1 and the sum is 8p[n]8p[n]. Otherwise the eight arrows are equal steps round the circle, and they add to 0, as the roots of unity did in Complex numbers for signals (3.3).

So the sum keeps only p[0]p[0] and p[8]p[8], times 8. The Hann window is 0 at n=0n=0 and 1 at n=8n=8, and 8p[8]=18p[8]=1. The eight channels add up to δ[n−8]\delta[n-8], a delay of 8 samples.

This is the arithmetic of the section “Windows that add up to a constant” in 15.6, with time and frequency swapped. There, copies of a window shifted in time added up to a constant. Here, copies of its spectrum shifted in frequency add up to a pure delay.

In the instrument every channel runs at the full rate, so there are 8 times as many samples as in the input. A practical bank keeps only every 8th output of each channel, which makes it critically sampled. Then it does not run 8 filters: it runs the 8 polyphase components of pp at the low rate and one 8-point FFT (22.4).

The MDCT

Audio codecs use a critically sampled bank called the MDCT, the modified discrete cosine transform. Each block makes NN new coefficients from NN new samples, so nothing is added to the rate. But each block’s window is 2N2N samples long, so neighbouring windows overlap by 50 %. It uses cosines in place of the DFT’s complex arrows.

Taken alone, one block cannot give back its 2N2N samples from NN coefficients: what it returns is aliased in time. Where two neighbouring blocks overlap, their aliases are equal and opposite, and overlap-add, as in 15.6, cancels them. This is time-domain aliasing cancellation, the twin of the first instrument with time in place of frequency.

AAC uses blocks of 1024 coefficients, with windows of 2048 samples. MP3 first splits the sound into 32 bands with a polyphase bank, then takes an MDCT with 18 coefficients in each band: 576 lines. Opus and Vorbis use the MDCT too.

Worked example

Let’s redo the page’s numbers by hand, from the filter gains.

1. The db2 bank at 0.30π. In the filter strip’s units the gains are 0.94344 (low) and 0.33156 (high) at 0.30π, and the other way round at 0.70π. The tone goes through HkH_k and then FkF_k at 0.30π, so the low band’s tone is 0.943442=0.8900.94344^2=0.890 and the high band’s 0.331562=0.1100.33156^2=0.110.

The alias goes through HkH_k at 0.30π and FkF_k at 0.70π. So both bands’ aliases are 0.94344×0.33156=0.3130.94344\times0.33156=0.313, with opposite signs. The tones add to 1.000, the aliases to 0, and the output is the input 3 samples late.

2. Processing breaks the cancellation. Scale the high band by g=0.5g=0.5. The alias left is (1−0.5)×0.3128=0.156(1-0.5)\times0.3128=0.156, and the tone is 0.890+0.5×0.110=0.9450.890+0.5\times0.110=0.945.

3. The Haar pair at 0.30π. Here ∣H0∣/2=cos⁡(0.15π)=0.891\lvert H_0\rvert/\sqrt2=\cos(0.15\pi)=0.891 and ∣H1∣/2=sin⁡(0.15π)=0.454\lvert H_1\rvert/\sqrt2=\sin(0.15\pi)=0.454. Each band’s alias is cos⁡(0.15π)cos⁡(0.35π)=0.405\cos(0.15\pi)\cos(0.35\pi)=0.405, and the tones are 0.794 and 0.206, adding to 1.

4. The DFT bank. The sum of the eight channels’ taps is 8p[n]8p[n] at n=0n=0 and n=8n=8, and 0 elsewhere. With p[0]=0p[0]=0 and p[8]=1/8p[8]=1/8 that is δ[n−8]\delta[n-8]. The magnitudes of the gains add to between 1.0000 and 1.0620.

5. In code. pywt.Wavelet('db2') in PyWavelets holds the four filters. Its rec_lo is h0h_0 as written here and rec_hi is h1h_1; dec_lo and dec_hi are the same taps reversed. scipy.signal.upfirdn(h, x, 1, 2) filters and keeps every 2nd sample in one call.

Where you’ll meet this

Filter banks are inside audio coders (MP3, AAC, Opus), equalisers with many bands, and hearing aids, which adjust the gain band by band. Some echo cancellers work band by band too, at a low rate in each. Software radios use DFT banks to pull many channels out of one wide band at once; they call them channelisers.

Split the low band again, and again, and you get the wavelet transform of Wavelets (23.2), built on the db2 pair of this page. How a codec decides how many bits each band gets is Perceptual audio coding (29.3).

The maths behind it · multiscale analysis

Splitting a time series into bands and studying each one is sub-band, or multiscale, analysis. Because an orthogonal bank keeps energy (Parseval), the variance of the series can be shared out among the bands.

Reference card

QuantityFormulaNotes
Two-channel bankY=TdistX+TaliasX(−z)Y=T_\text{dist}X+T_\text{alias}X(-z)analysis: HkH_k, ↓2; synthesis: ↑2, FkF_k, add
Distortion termTdist=12[F0H0+F1H1]T_\text{dist}=\tfrac12[F_0H_0+F_1H_1]acts on X(z)X(z)
Alias termTalias=12[F0(z)H0(−z)+F1(z)H1(−z)]T_\text{alias}=\tfrac12[F_0(z)H_0(-z)+F_1(z)H_1(-z)]acts on X(−z)X(-z), the spectrum slid by π\pi
Aliasing cancellationTalias(z)=0T_\text{alias}(z)=0the two aliases equal and opposite
Perfect reconstructionalso Tdist(z)=z−n0T_\text{dist}(z)=z^{-n_0}db2: n0=3n_0=3; Haar: n0=1n_0=1
Orthogonal pairh1[n]=(−1)nh0[Nh−1−n]h_1[n]=(-1)^nh_0[N_h-1-n], fk[n]=hk[Nh−1−n]f_k[n]=h_k[N_h-1-n]NhN_h taps; Daubechies; not linear phase beyond Haar
Classic QMFH1(z)=H0(−z)H_1(z)=H_0(-z)aliases cancel; perfect only with 2 taps
DFT bankchannel kk: p[n] ej2πkn/Mp[n]\,e^{j2\pi kn/M}an STFT; a 16-sample Hann with M=8M=8 adds to δ[n−8]\delta[n-8]
MDCTNN coefficients per NN samples, windows of 2N2N50 % overlap; AAC, MP3, Opus, Vorbis

End of lesson 23.1

Where to go next.

Phasorium
LibraryEvery lesson, in order

Parts

About Phasorium
Look