Here are two tones, at 0.10π and 0.35π rad/sample, and I keep only every -th sample, for = 1 to 4. Watch the second tone’s line: from on it passes π and folds back.
Keep every M-th sample: the spectrum stretches and folds
x[n] = cos(0.10πn) + 0.5cos(0.35πn); keep every M-th sample.
M = 1, every sample kept: two tones, at 0.10π and 0.35π rad/sample.
Describe this picture
Three stacked panels for . The first shows for samples 0 to 47, from −1.6 to 1.6: kept samples are stems with dot heads, dropped ones are open rings. The second is the spectrum of , amplitude from 0 to 1.1 against Ω from 0 to π rad/sample. Each tone is a vertical line as tall as its amplitude, labelled with its Ω. The third, the spectrum after keeping every -th sample, shows where the two lines land. A folded line has an open ring at its top and the label “folded”, and while a line moves a dotted arc shows its path, bouncing off π. The clip steps from 1 to 4, and the readouts give and where the 0.35π tone lands: 0.35π, 0.70π, then 0.95π and 0.60π, both folded. The caption names both tones’ places at each step and ends: “Keeping every M-th sample multiplies every frequency by M; what passes π folds back.” After the clip a slider named “Keep every M-th” sets from 1 to 8, and for above 4 the caption reads like “M = 6: the tones land at 0.60π and 0.10π (folded).”
Keep every M-th sample: the spectrum stretches and folds
In Shifting, reversing and scaling time (2.1) we sped a signal up by a whole number : . You keep every -th sample and throw the rest away. On this page I call that downsampling by , and I write the result as , with d for downsampled.
Why throw samples away? A converter that samples fast, as in Anti-aliasing and practical converters (10.4), hands over far more samples than the band you keep needs. Fewer samples means less to store and less to compute. But dropping samples does something to the frequencies, and that is what this page is about.
Let’s take two tones:
Keep every -th sample by putting in place of :
Read each cosine as . Every frequency is multiplied by . The amplitudes do not change: a cosine of amplitude 1 still has amplitude 1 when you skip some of its samples.
What happens when passes ? Frequency in discrete time (12.1) showed that and give the same samples. A cosine is even, so and give the same samples too. So a frequency past is the same as one below it: subtract , then drop the sign.
For the second tone goes to . Subtract to get , and drop the sign: . This is the fold of Sampling & aliasing (10.1), in rad/sample instead of hertz.
Notice that there is no clock and no analog part here. Aliasing happens inside the computer, from nothing more than dropping samples.
Think of a film projector that skips frames. A turning wheel seems to turn faster, and past some speed it seems to turn backwards. That backwards turn is the fold. The picture at the top of the page drops samples from these two tones for = 1, 2, 3 and 4, and moves each line to where it lands.
Notice that the 0.35π tone folds as soon as , because is past . The 0.10π tone moves up without folding, and neither line changes height.
After the clip, use the slider to set up to 8. Look at : the second tone lands at 0.10π, exactly where the first tone was in .
Then try . Both tones land on 0.80π, and the two lines stand side by side. They are now the same cosine, since , so . Two tones became one, and nothing in the samples can split them again.
The spectrum after downsampling
For two tones the rule was: multiply each by , then fold. For a whole spectrum it reads
Here is the DTFT of (The DTFT, 12.2), and is the DTFT of .
Read the term first. is stretched by : whatever had at now sits at . The other terms are the same stretched spectrum, shifted by . All terms are scaled by .
Why the shifted copies? A spectrum of samples repeats every (12.1), but the stretched one repeats only every . The copies fill the gaps, so the result repeats every again.
Where a copy overlaps the stretched original, the two add. That is aliasing: the overlapping copies of The sampling theorem (10.2).
And the , when the lines in the instrument kept their heights? A tone’s spectrum is made of arrows (12.2). Stretching the axis by makes each arrow’s area times larger, . That cancels the , so a sinusoid keeps its amplitude.
When does nothing alias? If is zero for , the stretched spectrum ends at . The copies, away, then do not overlap it. So nothing aliases when the signal has no content above .
Where the formula comes from
Start from the DTFT of :
Write . The sum then runs only over the that are multiples of , and . To let it run over every , I need a switch that is 1 when is a multiple of and 0 otherwise. The -th roots of unity of Complex numbers for signals (3.3) give one:
When is a multiple of , every term is 1 and . Otherwise the arrows are equal steps round the circle, and they add to 0, as in The DFT (13.2). Put in and swap the two sums:
The inner sum is the DTFT of at , and that gives the formula.
The maths behind it · selection matrices
Downsampling is a selection matrix: the rows of the identity, every -th one kept. It is wide and short, so it has a null space, and every signal in that null space is lost. Aliasing is two different signals with the same image under this matrix.
Filter first, then keep every M-th
The cure is the one of 10.4: remove what would fold before it can fold. Low-pass first, then keep every -th sample. The pair is called decimation, and the low-pass is the decimation filter.
How strong must the filter be, and where? With , let’s keep 0 to 0.2π, which becomes 0 to 0.8π after downsampling.
A component at lands at , and past it folds to . That is inside the kept band when , which is when .
So content from 0.2π up to 0.3π lands at or above, outside the kept band, and the stop band may start at 0.3π. With a kept band from 0 to , the same steps give
This is 10.4’s in the units of the new rate, which is on the old axis. The stop band does not have to start at .
For the filter I use a Kaiser design from Window-method FIR design (19.1). For 60 dB, . The transition runs from 0.2π to 0.3π, so , and Kaiser’s length formula asks for 74 taps.
I take 75, so that the delay is a whole number of samples, 37. The cutoff goes in the middle of the transition, 0.25π: in SciPy, firwin(75, 0.25, window=('kaiser', 5.6533)). I call it the 75-tap low-pass. It stays within ±0.00111 of 1 up to 0.2π, and from 0.3π it is at most −60.38 dB.
It is like blurring a photo slightly before you shrink it, so that fine stripes do not turn into false patterns. The instrument puts the 75-tap low-pass in front of the case. Its levels are in dB re 1: of the amplitude, so amplitude 1 is 0 dB and amplitude 0.5 is −6.0 dB.
Filter first, then keep every M-th
The same two tones, M = 4, with and without the 75-tap low-pass (cutoff 0.25π).
No filter, M = 4: the 0.35π tone lands at 0.60π, only 6 dB below the wanted tone.
Describe this picture
Two stacked panels for the same two tones and , with and without the 75-tap low-pass (cutoff 0.25π). The first, before dropping, shows level in dB re 1, from −90 to 5, against Ω from 0 to π rad/sample. The two tones are vertical lines rising to their levels, the filter’s gain is a dashed curve labelled “low-pass”, floored at −90 dB, and a dotted level marks −60 dB. The second, after keeping every 4th, has the same axes: the wanted line at 0.40π has a dot at its top and the folded line at 0.60π an open ring. The readouts are the filter, the wanted tone at 0.40π and the alias at 0.60π. The clip starts with the filter off: 0.0 dB and −6.0 dB. Then the filter’s gain draws, passing 0 to 0.2π and stopping everything from 0.3π. The 0.35π line drops by 68.14 dB while the 0.10π line stays at 0.00 dB. Last the second panel’s lines move, and the alias ends at −74.2 dB. After the clip two buttons in a group named “Anti-alias filter”, “filter off” and “filter on”, switch between the two states.
Watch the alias at 0.60π: without the filter it is only 6 dB below the wanted tone, and after it, −74.2 dB, below the 60 dB line.
Notice where the filter’s stop band starts: at 0.3π, not at . Its wide transition, from 0.2π to 0.3π, does no harm, because whatever lies there lands at 0.8π or above, outside the kept band.
Compute only what you keep
The filter makes one output for every input sample, and then three of every four are thrown away, so why compute them? Each output is a sum of 75 products, and I only need every 4th one. Computing only those costs 75 multiplies per output, not . Polyphase structures (22.4) builds this into the filter’s structure.
In SciPy, upfirdn(h, x, 1, 4) filters with h and keeps every 4th output, and it computes only those. decimate(x, 4, ftype='fir') does the same job with its own filter, an 81-tap Hamming design with cutoff 0.25π, and by default it removes the filter’s delay. Without ftype it uses an order-8 Chebyshev IIR filter (Analog prototype filters, 20.1), which this page leaves out.
One stage or two
A bigger drop costs more. Take 96 kHz down to 12 kHz, so , keeping 0 to 5 kHz with a 60 dB stop band. The stop-edge rule, in hertz, puts each stage’s stop edge at its output rate minus 5 kHz.
In one stage the stop edge is kHz. The transition is only 2 kHz wide at a 96 kHz rate, and Kaiser’s formula asks for 176 taps: 176 multiplies per output.
Now do it in two stages, 4 then 2. The first goes from 96 to 24 kHz, with its stop edge at kHz. Its transition, 5 to 19 kHz, is 14 kHz wide, so it needs only 26 taps. Whatever lies between 5 and 12 kHz after it is the second stage’s job.
The second stage goes from 24 to 12 kHz, with its stop edge at 7 kHz. Its transition is 2 kHz wide again, but at a 24 kHz rate that is four times wider in , so it needs 45 taps. The first stage runs twice for each final output, so the cost is multiplies per output, 55 % of one stage. This is multistage decimation.
The table adds a third design: three stages that each halve the rate.
| Design | Rates | Stop edges | Taps | Multiplies per output |
|---|---|---|---|---|
| one stage, 8 | 96 → 12 kHz | 7 kHz | 176 | 176 |
| two stages, 4 then 2 | 96 → 24 → 12 kHz | 19, 7 kHz | 26, 45 | 2 × 26 + 45 = 97 |
| three stages, 2, 2, 2 | 96 → 48 → 24 → 12 kHz | 43, 19, 7 kHz | 11, 14, 45 | 4 × 11 + 2 × 14 + 45 = 117 |
The early stages are short because their transition bands are wide: Kaiser’s length grows as (19.1). Three stages cost 117, more than two here: the two halving stages cost per output, against for the first stage of the two-stage design.
The table counts every tap. A halving stage can use a half-band filter (Special FIR filters, 19.4), in which every other tap is zero, and skip those multiplies.
Worked example
Let’s redo the page’s numbers by hand where we can, and with SciPy where we can’t.
1. Where the tones land. Multiply 0.10π and 0.35π by , then fold into 0 to π. To fold, subtract whole turns of until the frequency is within of 0, and drop the sign.
| Keep every M-th | 0.10π tone lands at | 0.35π tone lands at |
|---|---|---|
| 2 | 0.20π | 0.70π |
| 3 | 0.30π | 0.95π, folded |
| 4 | 0.40π | 0.60π, folded |
| 5 | 0.50π | 0.25π, folded |
| 6 | 0.60π | 0.10π, folded |
| 7 | 0.70π | 0.45π, folded |
| 8 | 0.80π | 0.80π, folded |
At both tones land on 0.80π.
2. The decimation filter. For 60 dB, . The transition is rad/sample, and the length formula gives . So and , which I make 75, with a delay of samples.
SciPy measures the 75 taps: within ±0.00111 of 1 up to 0.2π, at most −60.38 dB from 0.3π, and −68.14 dB at 0.35π. The alias after filtering is dB re 1. As in 19.1, every 4th tap from the centre is 0: 18 of the 75.
3. Stages. Kaiser’s formula gives 176 taps for one stage, and 26 and 45 for two. So the cost is 176 against multiplies per output. At 12 000 outputs a second that is 2.11 million against 1.16 million multiplies a second. On the 100 MHz processor of Real-time processing (21.4), that is 2.1 % against 1.2 %.
Kaiser’s formula is an estimate, as 19.1 found. Checked one tap at a time, the stop band first reaches −60 dB at 181 taps for one stage, and at 28 and 45 taps for two. That is 181 against 101, and two stages still cost 56 %.
Where you’ll meet this
A sigma-delta converter (Oversampling and noise shaping, 11.3) runs its 1-bit loop at megahertz rates. Its digital low-pass is a decimator, built in stages, that brings the rate down to 48 kHz or whatever the output needs.
Software-defined radios sample a wide band fast and decimate down to the rate one channel needs. Audio tools decimate when they turn a 96 kHz recording into 48 kHz. Data loggers sample a sensor fast and store fewer, filtered samples.
The reverse, raising the rate, is Upsampling and interpolation (22.2). Changing the rate by any ratio is Resampling by any factor (22.3). Computing only the kept outputs, built into the structure, is Polyphase structures (22.4), and banks of decimating filters are Filter banks (23.1).
The maths behind it · thinning a time series
Thinning a time series, keeping every -th observation, aliases its seasonal patterns the same way. Monthly data kept once a year cannot see the seasons at all: every kept value falls in the same month. Averaging before thinning is the low-pass.
Reference card
| Quantity | Formula | Notes |
|---|---|---|
| Downsampling | every frequency × M, then folded | |
| Spectrum | M stretched copies | |
| No aliasing | for | otherwise filter first |
| Decimation | low-pass, then keep every M-th | compute only the kept outputs |
| Stop edge | protects to | |
| Multistage | short early stages, wide transitions | 97 against 176 per output here |