To raise a sample rate, I start by putting zeros between the samples. Here a tone, , gets zeros after each sample, for from 1 to 4. Watch its spectrum: the tone’s line slides down, and new lines appear in the space it left.
Insert zeros: the spectrum squeezes and images appear
x[n] = cos(0.40πn), with L − 1 zeros put after each sample.
L = 1: the tone at 0.40π rad/sample, amplitude 1.
Describe this picture
Two stacked panels. The first is for samples 0 to 47, from −1.2 to 1.2: the original samples are stems with dot heads, and the inserted zeros are open rings on the axis. The second is the spectrum, amplitude from 0 to 1.1 against Ω from 0 to π rad/sample. The tone is a line with a dot at its top, labelled “tone”, and each image a line with an open ring, labelled “image”. The readouts are , where the tone is, and the height of each line. The clip steps from 1 to 4 as the zeros slide in and the lines move and split. At the tone is at 0.40π with amplitude 1. At it is at 0.20π, with an image at 0.80π, each 0.500. At the lines are at 0.13π, 0.53π and 0.80π, each 0.333. At the tone is at 0.10π and three images at 0.40π, 0.60π and 0.90π, each 0.250, and the caption ends: “Zeros divide every frequency by L and bring L − 1 images; a filter has to remove them.” After the clip a slider named “Upsample by L” sets from 1 to 6. At 5 the caption reads “L = 5: lines at 0.08π, 0.32π, 0.48π, 0.72π, 0.88π, each 0.200.” At 6 there are six lines, at 0.07π, 0.27π, 0.40π, 0.60π, 0.73π and 0.93π, each 0.167.
Insert zeros: the spectrum squeezes and images appear
In Downsampling and decimation (22.1) we lowered the sample rate by keeping every -th sample. Now I want to go the other way and raise it. Shifting, reversing and scaling time (2.1) already met the problem. Slowing down by 2 asks for samples halfway between the measured ones, at instants that were never measured.
So let’s leave the gaps empty first and fill them later. To raise the rate by a whole number , I put zeros after each sample. This is upsampling by , also called zero insertion or zero-stuffing. I write the result :
Nothing is lost. Every original sample is still there, now places from the next one.
What do the zeros do to a tone? Take and . At even , , and at odd it is 0. The factor is 1 at even and 0 at odd , so
The second line takes one step. is , and by Properties of the DTFT (12.3), multiplying by it slides every line of a spectrum by . The tone’s lines at move to and . On the circle of Frequency in discrete time (12.1), is , so the product is .
So the zeros did two things. The tone moved to half its frequency, . And a second tone appeared at , which was not in at all. I call it an image.
The tone and its image each have amplitude .
Think of a picket fence with a blank picket painted in after every real one. The old rhythm is still there, at half the pace, and a new, false rhythm sits beside it. The picture at the top of the page does this to the tone for from 1 to 4.
Watch two things. Every line’s height falls as : 1, 0.500, 0.333, 0.250. And the tone slides down to , while new image lines grow in the space the tone left. After the clip, the slider goes on to and 6.
Squeezed by L
Why does every frequency divide by ? Write the DTFT of (The DTFT, 12.2). Only the terms with are not zero, and there :
So the new spectrum is the old one, read at . What sat at now sits at : the spectrum is squeezed by . The old spectrum repeats every (12.1), so the new one repeats every . That fits copies into one turn, and the extra copies are the images.
Now follow a tone at . The old spectrum has its line at and at every copy, . The new one has a line wherever hits one of those, so at
Larger only repeat these a whole turn further on. Fold each one into 0 to , as in 22.1. For and that gives , , and , which fold to , , and .
Why is each line high? For the factor did it. For any the same job is done by
When divides , every arrow is 1 and the factor is 1. Otherwise the arrows are spread evenly round the circle and cancel, like the rows of The DFT as a matrix (13.5). So for the tone, is times this factor.
Each arrow slides the squeezed tone by (12.3) and carries the weight . That makes lines, each a cosine of amplitude .
The maths behind it · transposed selection matrices
Downsampling was a wide selection matrix: rows of the identity, every -th kept. Upsampling is its transpose: a tall matrix that places each sample and pads with zeros. Interpolation, the filter after upsampling, is a matrix whose columns are shifted copies of the interpolation pulse: a basis of in-between functions.
Remove the images, and multiply by L
The zeros moved the tone to the right place, , and added images. So the repair is a low-pass that keeps the squeezed spectrum and removes the rest. The old band, 0 to , now fills 0 to , so the cutoff is .
The tone also came out times too small. So the filter gets a gain of as well. A low-pass at with gain is the interpolation filter.
For the cutoff is , and 22.1 already has that filter. It is the 75-tap low-pass, a Kaiser design (Window-method FIR design, 19.1) with and cutoff . It passes 0 to within , stays at least 60.38 dB down from , and delays by 37 samples. Here I multiply its taps by 4.
Will it do? The tone sits at , in the pass band. The images sit at , and , all past , in the stop band.
In fact it serves any with nothing above . That content squeezes to at most , and its nearest image starts at .
Think of a flip-book with a blank page after every drawing. The filter draws the in-between pictures.
The second panel gives levels in dB re 1: a line of amplitude sits at . So the four lines of amplitude 0.250 sit at −12.0 dB, and the filter’s gain of 4 is +12.04 dB.
Remove the images, and multiply by L
L = 4: the zero-stuffed tone through the 75-tap low-pass of 22.1, times 4 (cutoff 0.25π).
L = 4: the tone at 0.10π and three images, all at −12.0 dB.
Describe this picture
Two stacked panels for : the zero-stuffed tone through the 75-tap low-pass of 22.1, times 4, with cutoff 0.25π. The first is the signal for samples 0 to 47, from −1.2 to 1.2. The zero-stuffed input is stems with dot heads and open rings for its zeros. The output is stems with square heads, moved back by its delay of 37 samples so that it lines up, and the sine is a thin dashed curve labelled “sine”. The second is the spectrum, level in dB re 1 from −90 to 15 against Ω from 0 to π rad/sample. The filter’s gain is a dashed curve labelled “low-pass × 4”, the tone and images are lines, and the images have open rings at their tops. The readouts are the tone’s amplitude and the largest image. The clip opens on the tone at 0.10π and three images, all at −12.0 dB, reading 0.250 and −12.0 dB. Then the filter’s gain draws: about 4 (12.03 to 12.05 dB) up to 0.2π, and at least 60 dB lower from 0.3π. Last the lines move to their filtered levels and the output stems grow into the gaps. The readouts end at 1.000 and −72.4 dB, and the samples lie on the sine.
Watch the images sink and the gaps fill. The images fell from −12.0 dB to −72.4, −74.3 and −79.6 dB. The tone, which was 0.250, came out at 1.000. And in time, the output samples lie on the sine to within 0.0006.
Why the gain is L, and why the gaps fill with the right values
The gain is because zero insertion divided the tone’s amplitude by . Only one sample in carries the signal, and the filter has to bring the level back.
Why does the output land on the sine? The filter’s taps are a windowed sinc, centred on tap 37. Taps 4, 8, 12, … away from the centre are 0, because the ideal low-pass at has when is a multiple of 4. The centre tap of is 0.99969.
Take an output at the place of an original sample. The input’s other non-zero samples are 4, 8, 12, … places away, so they meet only zero taps. The output there is 0.99969 times the original sample, so the original samples are kept almost exactly.
In between, every non-zero tap adds in a scaled pulse from each nearby original sample. The ideal taps times 4 are : a sinc pulse that is 0 at every old sample instant. That is the sum of Reconstruction (10.3), a sinc pulse at every sample, added, now done in discrete time and cut to 75 taps by the window.
The same 75 taps decimated by 4 in 22.1. Here, times 4, they interpolate by 4. In SciPy, upfirdn(4*h, x, 4) inserts the zeros and filters in one call. resample_poly(x, 4, 1) does the same with a Kaiser filter of its own, and it removes the delay for you.
Notice that three inputs in every four are zeros. For each output, only 18 or 19 of the 75 taps meet a non-zero input, and the rest multiply zeros. Polyphase structures (22.4) skips those multiplications.
Worked example
Let’s redo the page’s numbers, by hand where we can and with SciPy where we can’t.
1. The images of . The lines sit at , folded into 0 to . For : gives , gives , and gives , which folds to .
The same rule gives these lines, each of amplitude :
| Factor | Lines (rad/sample) | Each line |
|---|---|---|
| , | 0.500 | |
| , , | 0.333 | |
| , , , | 0.250 |
2. After the filter. The tone’s amplitude becomes . Each image becomes at its frequency: −72.4497 dB at , −74.3090 dB at and −79.6460 dB at .
3. The taps against the sinc. Next to the centre, is 0.89836, 0.63167 and 0.29499. The ideal for is 0.90032, 0.63662 and 0.30011. The window and SciPy’s scaling pull each tap down slightly.
4. One in-between value. Take an output one place after an original sample where the sine is at its peak. The sine there is . The filter gives 0.95096, which is 0.0001 low. The original sample beside it comes out at 0.99969 in place of 1.
Where you’ll meet this
Digital-to-analog converters upsample by 4 to 256 before the analog filter. The images then sit far above the audio band, so a gentle analog filter removes them. That is the easier filter of Anti-aliasing and practical converters (10.4), on the output side, and sigma-delta converters (Oversampling and noise shaping, 11.3) take it furthest.
Audio effects that would otherwise alias, such as distortion, run at a raised rate (Audio effects, 29.2). Zooming an image is interpolation along its rows and then its columns.
A ratio that is not a whole number is Resampling by any factor (22.3). In Filter banks (23.1), each band is upsampled and filtered like this, and the bands are added to rebuild the signal.
The maths behind it · kernel interpolation
Filling the missing observations of a regular time series with a smooth kernel is kernel interpolation. The sinc is the band-limited choice of kernel. A window trades accuracy for a short kernel, as in 19.1.
Reference card
| Quantity | Formula | Notes |
|---|---|---|
| Upsampling | if divides , else 0 | zeros inserted |
| Spectrum | squeezed by | |
| Lines of a tone at | , , folded | the tone, the rest images; amplitude each |
| Interpolation filter | low-pass at , gain | keeps the tone, removes the images |
| Its taps | windowed | 0 at every multiple of from the centre |
| Same filter | decimates by , or, times , interpolates by | 22.1 and 22.2 |