A known pulse, a gliding tone, is hidden somewhere in 512 samples of noise about as strong as the pulse. Slide a copy of the pulse along the record, and at each position multiply and add. Watch the output: where the copy lines up with the hidden pulse, it jumps.
Slide the template, find the pulse
512 samples: a 64-sample chirp of amplitude 1 starting at n = 300, in white noise of standard deviation 1 (seed 261; measured 0.985).
A chirp is hidden somewhere in these 512 noisy samples. Slide the template along and add the products at each lag.
Describe this picture
Two stacked panels for 512 samples: a 64-sample chirp of amplitude 1 starting at , in white noise of standard deviation 1 (seed 261; measured 0.985). The first, the received , runs from −4 to 4 against sample from 0 to 511. It draws the record faintly, and the template over it at the current lag, with a bracket marking the template’s 64 samples. The second, the matched-filter output, runs from −20 to 35 against lag ℓ from 0 to 448. The output grows as a line, one lag at a time, and once the peak is reached a filled diamond marks it, labelled “300”. The readouts are the lag and the output. There is no control. At lag 0 the template sits at the start of the record, and the output is −7.2. At lag 300 the template lines up with the hidden chirp, and the output jumps to 31.8. The template slides on to lag 448, where the output is −6.7. The end caption says that the peak, 31.8 at lag 300, is 5.8 noise standard deviations high, and that more than one pulse length away the output never passes 17.2.
Slide the template, find the pulse
You can pick out a friend’s voice in a noisy crowd, because you know exactly how it sounds. A radar receiver has the same advantage. It sent the pulse itself, so it knows the shape of the echo it listens for. What it does not know is when the echo comes back.
Here is the problem in symbols. A known pulse , the template, sits somewhere in a record, scaled by an amplitude and buried in noise:
The template is samples long and starts at , which I do not know. The noise is white, as in “White noise forgets, coloured noise remembers” of Random processes (24.2), with standard deviation .
My template is a chirp, the gliding tone of “Reading a chirp” in Spectrograms & the STFT (15.5). It has samples, for = 0 to 63. Its frequency, the slope of its phase, glides from to rad/sample.
Its energy, the sum of its squares from “Adding it all up” in How big is a signal (1.3), is . That is close to half of 64, since the squares of a cosine average about one half.
In the first instrument the record has 512 samples. The chirp has amplitude and starts at , and the noise has . It is a seeded draw (seed 261), so it is the same on every visit; this draw measures a standard deviation of 0.985.
To find , I do what Correlation (24.3) did in “Slide, multiply, add: no flip”. Slide the template along the record, multiply sample by sample, and add. At lag the template’s first sample sits on :
This is 24.3’s , with : the record stays still and the template slides. So the peak should land at , where the record lags the template by . The lags run from 0 to 448, the last place where all 64 samples of the template fit inside the record.
From statistics this page assumes only this: the noise is white and Gaussian, its power is known, and so is the pulse’s shape. Only the pulse’s position is unknown.
The picture at the top of the page slides the template along this record, one lag at a time. Notice the end frame: the peak, 31.8 at lag 300, stands 5.8 noise standard deviations high, and it is narrow. One lag to either side, at 299 and 301, the output is already down to 17.6 and 18.8.
Look back at the record panel, too. The chirp’s amplitude, 1, is about the noise’s standard deviation, so nothing there shows where it starts. The correlation finds it anyway.
Why the peak stands so tall
At the right lag, , each product of the pulse with the template is , never negative. So they all add up, to times the energy:
The noise’s products have random signs, so they partly cancel. Their sum adds 64 separate draws, each scaled by . Scaling a draw by multiplies its variance by , since the variance is an average of squares.
By “Sums of draws” in Random variables for signals (24.1), the variances of separate draws add. So the noise in the output has variance
and standard deviation , at every lag.
The peak stands noise standard deviations high. Square that, and you have a ratio of powers, the output SNR:
Only the pulse’s energy counts, not its shape. For the chirp, , so with and the peak should stand 5.71 standard deviations high.
In this record the peak is 31.8. The chirp alone would give 32.55, and the noise at that lag adds −0.77. Where the template no longer overlaps the pulse, 64 lags or more away, the output’s standard deviation measures 5.52, against 5.71 in theory. So the peak stands of them high.
With this draw’s measured 0.985 in place of , the theory says 5.62. The 322 lags are not separate draws, since neighbouring lags share 63 of their 64 samples, so 5.52 is a rough estimate.
In decibels, as in “Signal and noise, in decibels” of 1.3: during the pulse, the chirp’s power per sample is against the noise’s 1. That is an SNR of −2.9 dB. With this draw’s measured noise RMS, 0.987, it is −2.8 dB.
After the filter the SNR is dB. The filter has multiplied the SNR by , which is 18 dB. That holds for any pulse of samples, because is times the pulse’s power per sample.
No filter does better
Could other weights beat the template? Say a filter weights the 64 samples by some instead of . The same two steps give the pulse part , and the noise’s standard deviation . Their ratio is
The last fraction is the normalised correlation of “Between −1 and 1” in 24.3, of with at lag 0. It is at most 1, and it reaches 1 only when is a positive multiple of . Weights outside the pulse’s 64 samples collect only noise, so they cannot help either.
So in white noise no linear filter lifts the pulse higher above the noise than the template does: is the most it can reach. This is the Cauchy–Schwarz inequality, which linear algebra proves.
The maths behind it · the Cauchy–Schwarz inequality
Treat the 64 samples under the template as a vector , and the template as . The output at the right lag is the inner product . White noise spreads equally in every direction, so the direction that collects the most pulse against the noise is the pulse’s own, and the Cauchy–Schwarz inequality says no other unit-length direction does better.
The template, played backward
So far I have correlated. To do it with a filter, recall “A convolution with one signal reversed” in 24.3: convolving with a reversed signal is correlating with it. The reversal is “Playing it backward: reversal” of Shifting, reversing and scaling time (2.1). The flip in “Flip and slide: a mechanical way to compute it” of Discrete convolution (5.2) undoes it.
So the filter’s impulse response is the template played backward, moved to start at :
For the chirp, : it starts with the template’s last sample. The move by keeps at = 0 to 63, so the filter is causal and needs no future samples.
Its output, with , is the correlation again:
The output is the correlation, 63 samples late. Its peak comes at , the pulse’s last sample. The filter can add up the whole pulse only once the whole pulse has arrived.
This filter, matched to one pulse, is the matched filter. Built as a sliding correlation instead, the same receiver is often called a correlation receiver.
Everything here assumed white noise. If the noise is coloured, first pass the record through a filter that makes the noise white, such as the prediction-error filter of Parametric models and linear prediction (25.3). Then match the template as it looks after that filter.
Where to set the threshold
The peak tells me where the pulse is. A receiver must also decide whether there is a pulse at all. A weak echo may stand only a little above the noise, and noise alone sometimes rises high.
Think of a smoke alarm. Set it sensitive, and it rings for toast. Set it dull, and it stays silent through a small fire. A detector has the same dial.
First I make the output’s scale simple. Divide it by its noise standard deviation, . Without a pulse, the result at any lag has mean 0 and standard deviation 1. It is also Gaussian, because a weighted sum of Gaussian draws is Gaussian; I assume this; statistics proves it.
With the pulse at that lag, the same bell moves to
The noise is the same, so the bell keeps its width; only its centre moves. On this page is this distance between the two bells, in noise standard deviations.
For this section the chirp is weaker, with amplitude . Then , which the instrument shows as 2.00.
Now pick a threshold , in these units, and declare “pulse” whenever the output passes it.
Two things can go wrong. Noise alone can pass : that is a false alarm. A pulse can stay under it: that is a miss. The share of pulses that do pass is the detection rate.
The false-alarm rate is the area of the no-pulse bell beyond . For whole numbers, 24.1’s areas give it. Within 1, 2 and 3 standard deviations lie 68.27 %, 95.45 % and 99.73 % of the draws. The bell is symmetric, so half of the rest lies beyond : 15.87 %, 2.28 % and 0.13 % for = 1, 2 and 3.
For any , this tail area has a name. The complementary error function is a tabulated function that gives Gaussian tail areas; SciPy has it as scipy.special.erfc. The false-alarm rate is
The detection rate is the same area for the moved bell, .
Each threshold gives a pair of rates. Plot the detection rate against the false-alarm rate, one point per threshold, and the points trace the ROC curve, short for receiver operating characteristic.
Where to set the threshold
The matched-filter output ÷ its noise standard deviation, for the chirp at amplitude 0.35: bells at 0 (no pulse) and 2.00 (pulse).
Threshold 3: false alarms 0.13 %, but only 15.8 % of pulses detected.
Describe this picture
Two panels for the matched-filter output divided by its noise standard deviation, for the chirp at amplitude 0.35: bells at 0 (no pulse) and 2.00 (pulse). The first shows density, with no numbers, against the output from −4 to 6 noise standard deviations. The bell at 0, named “no pulse”, is dashed, and the bell at 2.00, named “pulse”, solid. Beyond η the area under the bell at 0 is shaded and hatched, and the area under the bell at 2.00 is shaded; a key names them “false alarms” and “detections”, and a vertical line labelled “η” marks the threshold. The second, the ROC, shows the detection rate against the false-alarm rate, each from 0 to 1, with a dotted diagonal labelled “chance”. The traced curve grows as the threshold moves, and a filled dot marks the current threshold. The readouts are the threshold η, the false alarms and the detections, in percent. There is no control during the clip. At η = 3 the false alarms are 0.13 %, but only 15.8 % of pulses are detected. At 2: 2.28 % false alarms, 49.9 % detected. At 1: 84.1 % detected, at the price of 15.87 % false alarms; the traced curve is the ROC, and every threshold is one point on it. After the clip a slider “Threshold η”, named “Detection threshold”, runs from −1 to 5 in steps of 0.1. At other settings the caption reads like “Threshold 2.5: false alarms 0.62 %, detections 30.7 %.”
Watch the dot trace the curve as the threshold slides from 3 to 1: the detections rise, and so do the false alarms.
After the clip, move the threshold yourself with the slider.
Notice the middle frame: at about half the pulses are caught, 49.9 %, at 2.28 % false alarms.
Reading the curve
Check that frame by hand with . The bell of the pulse is centred on , so half of it lies beyond: 50.0 %. The readout says 49.9 % because the instrument uses , slightly below 2. In the same way, gives 15.9 % at , against the readout’s 15.8 %; at both give 84.1 %.
Lowering moves the dot along the curve towards the corner where both rates are 1: more pulses caught, and more false alarms. Raising it moves the dot back towards the corner where both rates are 0. One rate improves only when the other gets worse.
Now the diagonal. Imagine a detector that ignores the record and declares “pulse” at random, a share of the time. It catches a share of the pulses and raises a share of false alarms, so it sits on the diagonal. Every point of the matched filter’s curve lies above it.
A stronger pulse moves the second bell further out. Then the curve bends further towards the corner with false-alarm rate 0 and detection rate 1, where every pulse is caught and noise never is.
Measured, not assumed
The bells are a model. To test it, I ran 2000 trials without the pulse and 2000 with it, each with 64 fresh draws through the matched filter (seed 2610). The outputs averaged 0.005 and 1.967, against the model’s 0 and 1.997. Their standard deviations were 0.974 and 1.010, against the model’s 1.
At = 3, 2 and 1 the false alarms were 0.10 %, 1.95 % and 15.05 %. The detections were 15.5 %, 48.8 % and 83.6 %.
How close should they be? Each trial either crosses or not: a draw that is 1 with probability and 0 otherwise. Its mean is , and its variance is , because its square is itself.
The share over 2000 separate trials adds 2000 such draws and divides by 2000. The variances add, and the division scales the variance by , as before. So the share has spread .
At that is 0.82 %, and the measured 15.05 % lies 0.82 below the bell’s 15.87 %, just under one spread. Each of the six measured rates lies less than one spread from the bells’ value.
These rates are for one decision. A search like the first instrument’s makes one decision at every lag, so the noise gets many chances. In that record, of the 322 lags at least 64 away from the pulse, one passed : lag 384, where the output was 17.14, just above .
The chirp of the first instrument, with , has , and catches 99.7 % of such pulses. So a search over many positions sets its threshold high, and pays for it with a strong pulse.
In practice you first decide how many false alarms you can afford, then set to give that rate. This is the Neyman–Pearson rule.
The maths behind it · hypothesis tests
Detection is a hypothesis test. “No pulse” is the null hypothesis, a false alarm is a type I error, a miss is a type II error, and the detection rate is the test’s power. The ROC curve is that power plotted against the size of the test.
Worked example
1. A matched filter by hand. Take the template = 1, 2, −1 at = 0 to 2, so and . Played backward, is −1, 2, 1. Let the record be the template starting at , with no noise: = 0, 0, 1, 2, −1 at = 0 to 4.
Convolving with gives 0, 0, −1, 0, 6, 0, −1 at = 0 to 6. At , for example, the products are . The peak, 6, is , at , the pulse’s last sample. The correlation itself peaks at lag 2, which is .
2. Output SNR. For the chirp, , with and , gives dB, against −2.9 dB per sample during the pulse. The peak should stand standard deviations high; this record’s stands 5.8.
3. One threshold. At , with : false alarms %. Detections are half the pulse bell, 50.0 % by hand with , and 49.9 % with the full .
4. How strong must the pulse be? Keep the false alarms at 0.13 %, so . To catch 84.1 % of pulses, the pulse bell must sit one standard deviation beyond , at , as 24.1’s 68.27 % shows. That needs , twice the 0.35 of the instrument.
Where you’ll meet this
Radar and sonar correlate each echo with the pulse they sent. They often send a chirp, as here: it lasts long, so it carries much energy, yet its peak is narrow. The chirp of this page lasts 64 samples, but its peak, alone, is only 3 lags wide at half its height. Squeezing a long pulse into a short peak is called pulse compression, the subject of Radar signal processing (33.1).
A GPS receiver correlates what it hears with each satellite’s known code, at many delays, and declares a satellite found when the peak passes a threshold. Image software slides a small template over a picture to find where it appears. Neuroscientists find nerve-cell spikes in noisy electrode recordings by matching a spike’s known shape.
The matched filter finds a pulse whose shape you know. When the signal itself is unknown and you want its whole waveform back from the noise, you need The Wiener filter (26.2).
For more, see A. V. Oppenheim and G. C. Verghese, Signals, Systems and Inference (2015), the chapters on hypothesis testing and signal detection, and H. L. Van Trees, Detection, Estimation, and Modulation Theory, part I. In SciPy, scipy.signal.correlate(x, q, mode='valid') gives the output at every lag, and scipy.special.erfc the tail areas.
Reference card
| Quantity | Formula | Notes |
|---|---|---|
| Matched filter | correlate with the template; the peak comes samples after the pulse starts | |
| Output SNR | the best any linear filter can do, in white noise | |
| SNR gain | over the SNR per sample during the pulse | |
| False alarms | in noise standard deviations; 15.87 %, 2.28 %, 0.13 % at 1, 2, 3 | |
| Detections | ; Gaussian noise | |
| ROC | detections against false alarms as moves | above the diagonal is better than chance |