A slow random signal is buried in white noise, and the noise grows. The best filter passes each frequency in proportion to the signal’s share of the power there. Watch the gain: as the noise rises, the point where it falls to ½ slides down.
Pass what is mostly signal
AR(1) signal (a = 0.9, power 1) in white noise: the two PSDs and the Wiener gain.
Noise at −10 dB: the gain is above ½ up to 0.516π; the error power falls from the noise's 0.1 to 0.059.
Describe this picture
Two stacked panels for an AR(1) signal (, power 1) in white noise, sharing the axis Ω from 0 to π rad/sample. The first shows the two PSDs, from −25 to 25 dB: the signal’s is a solid curve labelled “signal”, and the noise’s a flat dashed line labelled “noise”. The second draws the Wiener gain as a solid curve from 0 to 1; a dotted level marks 0.5, and a small tick marks where the gain crosses it. The readouts are the noise power in dB, where the gain is ½, in multiples of π, and the error power. With the noise at −10 dB the gain is above ½ up to 0.516π, and the error power falls from the noise’s 0.1 to 0.059. At 0 dB the gain is 0.950 at 0 and 0.050 at π, ½ at 0.144π, and the error power is 0.218 instead of 1. At 10 dB only the strongest low frequencies pass (gain 0.655 at 0, ½ at 0.032π), and the error power is 0.586 instead of 10. When the clip has finished, a slider “Noise power” sets the noise from −20 to 20 dB in steps of 1 dB. At 5 dB the caption should read “Noise at 5 dB: gain ½ at 0.075π, error power 0.375.” At −20 dB the gain stays above ½ everywhere, so the readout says “none, above ½”, and at 20 dB it says “none, below ½”.
The best guess of a signal in noise
In Simple smoothing filters (18.2), “A longer average: less noise, more delay” cleaned up a noisy scale with a moving average. A longer average removed more noise, but it also blurred the parcel’s landing more. So how long should the average be, and should every reading get the same weight?
This page answers both questions for a signal and a noise whose statistics we know. The answer is the Wiener filter, worked out by Norbert Wiener in the 1940s.
Here is the setting. The signal I want is , the desired signal. I never see it alone. I see it with white noise of power added, the white noise of “White noise forgets, coloured noise remembers” in Random processes (24.2):
Both have mean 0 and are stationary, and they are uncorrelated: for every and . The noise does not know what the signal is doing.
A filter turns into an output , my guess at the signal. The error is what the guess misses, . On this page is this error, not the rounding error of Quantization & noise (11.1).
One error sample can be large or small by chance, so I judge a filter by the error’s average size. The mean-square error is the error’s power:
Squaring makes a miss of 2 count four times as much as a miss of 1, and it leads to equations we can solve. That was the reason for least squares in “Least squares or least worst” of Optimal FIR design (19.3) too. The Wiener filter is the filter with the smallest , and that smallest value is the minimum mean-square error .
Two lazy filters set the bar. Passing everything, , leaves the error , so , the noise’s power. Passing nothing, , leaves the error , so is the signal’s power . A good filter beats both.
The best filter, frequency by frequency
First, let the filter use the whole record, the samples after as well as those before. That is a non-causal filter, fine for cleaning a recording afterwards. Then the best filter can be found one frequency at a time.
A filter with impulse response and frequency response gives the output . So the signal reaches the error through the system , and the noise through with a minus sign. A sign does not change a power.
In Power spectral density (24.4), “White, pink, brown” ended on the rule : a system multiplies a PSD by its gain squared. The two parts of the error are uncorrelated, so their autocorrelations add, and by Wiener–Khinchin so do their PSDs. Using the rule twice, the error’s PSD is
where and are the PSDs of the signal and of the noise. I leave out the argument to keep it short. The noise is white, so , the same at every (24.4).
is the error’s power, the area under that PSD divided by (24.4). Each frequency adds its own share, and the filter can choose its gain at each frequency separately. So I make the error’s PSD as small as possible at every .
At one frequency this is a tug of war. A gain near 1 lets the noise through; a gain near 0 throws the signal away. An imaginary part of adds to both terms, so the best is real. For real the error’s PSD is a parabola in , and its lowest point is where its slope, , is 0:
In words: at each frequency, the gain is the signal’s share of the power there. Where the signal is ten times the noise, the gain is . Where they are equal it is ½, and where the noise is ten times the signal it is .
Put back in, and the error’s PSD becomes . That is , and also , so it is smaller than both and . Its area gives the least error:
So the Wiener filter beats both lazy filters at every frequency, and therefore in total. is real and positive, so it adds no delay, like the zero-phase response of Linear-phase systems (17.3). It is also even in , so its impulse response is symmetric about : it leans on the samples after as much as on those before. That is why it needs the whole record.
One signal for the whole page
The signal on this page is the AR(1) noise of 24.2 with . Each sample is 0.9 times the one before plus fresh white noise of power 0.19, so its power is and its autocorrelation is . In “Memory in time, narrowness in frequency”, 24.4 found its PSD:
It is 12.8 dB at and dB at , and its area divided by , the power, is 1.000. It is a slow signal, with its power gathered at low frequencies. The white noise spreads its power evenly, so the Wiener filter will be a low-pass filter, and the noise level decides how far up it reaches.
Pass what is mostly signal
The picture at the top of the page draws the two PSDs and the gain while the noise grows. Think of listening to a friend in a noisy room: you attend to the pitch range where their voice is louder than the room.
Its noise power is in dB, , so −10 dB is a noise power of 0.1 and 10 dB is 10. Its error power is .
Notice where the gain is ½. As the noise rises a hundredfold, from −10 dB to 10 dB, that point moves from 0.516π down to 0.032π. More noise pulls the gain down everywhere, and only the frequencies where the signal is strongest keep a gain above ½.
At 0 dB the gain is ½ at 0.144π. That is where , the frequency where “Memory in time, narrowness in frequency” in 24.4 saw the AR(1) PSD cross the white level, at .
Now compare the error with the two lazy filters. At −10 dB the error, 0.059, is 59 % of the noise’s 0.1. At 10 dB it is 0.586, under 6 % of the noise’s 10, and also below the signal’s power, 1, which passing nothing would leave.
When the clip has finished, drag the slider to 5 dB: the gain is ½ at 0.075π, and the error power is 0.375. At −20 dB the gain stays above ½ everywhere, and at 20 dB below it.
A causal filter with a few taps
needs the future, and its impulse response never quite ends. A filter running live must be causal, and I would like a short one. So let’s look for the best FIR filter of taps to :
The error is linear in the taps, so is a quadratic in them, a bowl, like the mean square of 19.3’s least squares. Its lowest point comes from linear equations.
Parametric models and linear prediction (25.3) showed the quickest way to them. At the best taps, the error must be uncorrelated with every input sample the filter uses. If it were still correlated with , adding a little more of would shrink it. So for ,
In calculus terms, this sets the slope of along each tap to 0: that slope is .
Put in . Averages add, and each is the autocorrelation of 24.2. So
The right side is a cross-correlation, as in Correlation (24.3), but averaged over the process: what I want against what I see, samples earlier. These are the Wiener–Hopf equations.
In matrix form they read . Here holds the best taps, and is the matrix with in row , column : the Toeplitz matrix of 25.3. The vector holds the right sides. In 25.3 it held autocorrelations; here its entries are the cross-correlations .
From linear algebra this page needs one fact. When is positive definite, as 25.3 defined it, and it is on this page, the system has exactly one solution. SciPy’s solve_toeplitz finds it with a Levinson-type recursion that uses the Toeplitz pattern, as 25.3 did.
For this page’s signal both sides are easy. The signal and the noise are uncorrelated, so their autocorrelations add. The noise’s is at lag 0 and 0 elsewhere (24.2), so
On the right, is plus , so the entries of are .
With noise power 1 and three taps, the Wiener–Hopf equations are
and their solution is , and .
How small is the error at the best taps? There the error is uncorrelated with every , and so with , which is made from them. So , and multiplying out gives
where is the sum of the products of matching entries. For the three taps above it is .
The best FIR filter, from correlations
The second instrument solves the Wiener–Hopf equations on the page with noise power 1, and runs the filter over a seeded record. Picture a weighted average of recent readings, whose weights come from how alike neighbouring readings are.
The record is 4000 samples of this page’s signal. I took 4500 draws of the site’s generator (seed 262), scaled them by and passed them through the AR(1) recursion, then dropped the first 500 outputs. The noise is 4000 more draws (seed 2620). Measured over the record, the signal’s power is 0.875 and the noise’s 0.986.
The best FIR filter, from correlations
The AR(1) signal (seed 262; measured power 0.875) in white noise of power 1 (seed 2620; measured 0.986), samples 100 to 299.
One tap: the best single weight is 0.500, which halves the error power to 0.500.
Describe this picture
Two panels for the AR(1) signal (seed 262; measured power 0.875) in white noise of power 1 (seed 2620; measured 0.986). The wide one shows samples 100 to 299, from −5 to 5: the noisy is small dots, the desired a dashed line, and the Wiener output a solid line; a key names each by its mark. The narrower one shows the taps as stems, from 0 to 0.6 against from 0 to 31. The readouts are the number of taps , the error power in theory, , and the error power measured on the record. With one tap the best single weight is 0.500, which halves the error power to 0.500; measured, 0.464. With four taps, 0.314, 0.204, 0.139 and 0.105, the error power is 0.314; measured, 0.303. With sixteen it is 0.304, the same as eight to three decimals, and the taps beyond the eighth are at most 0.0073; measured, 0.299. When the clip has finished, a slider “Taps” sets from 1 to 32.
Watch the solid output follow the dashed signal through the dots, and the error power level off as taps are added. The measured value is the mean of over samples 100 to 3999.
Notice how the error levels off. One, two, four and eight taps give 0.500, 0.373, 0.314 and 0.304, which is −3.01, −4.28, −5.03 and −5.17 dB. From eight taps on, the error stays at 0.304, and even 400 taps reach no lower.
Look at the taps in the end frame. Apart from the last few, each is 0.63 times the one before. The weights fade with distance, like the signal’s own memory, but faster than its 0.9 per step. An older sample brings as much noise as a newer one, but tells less about the present. From the ninth tap on, every tap is below 0.008, so more taps add almost nothing.
There is a neat check in the captions: the first tap equals the error power, 0.500 with 0.500 and 0.314 with 0.314. That holds whenever the noise is white. The error is uncorrelated with , so , and only the tap ever meets , giving .
The measured errors sit a little below the theory: 0.464 against 0.500, 0.303 against 0.314 and 0.299 against 0.304. The theory is an average over every possible record, and this one is a single record. One reason is visible in the subtitle: its signal is weaker than nominal, with power 0.875 instead of 1.
When the clip has finished, use the slider to try 2 taps: the error power drops from one tap’s 0.500 to 0.373, and the measured error to 0.351.
Looking ahead, and what it assumes
This FIR filter only looks back. Allowed to look ahead too, the error at noise power 1 falls to the first instrument’s 0.218. That is below 0.304, the least error any causal filter can reach here.
The best causal filter of unlimited length, the causal IIR Wiener filter, comes from splitting into a causal and an anti-causal factor, a step called spectral factorisation. I leave it out here; for this signal its error is the same 0.304 that long FIR filters reach.
The Wiener filter makes three assumptions. The signal and the noise are stationary, they are uncorrelated with each other, and their correlations are known. When the correlations are not known, Adaptive filters: LMS (26.3) learns the filter from the data instead.
The derivation never used the noise being white, only uncorrelated with the signal. With coloured noise, and change, and the method does not. If the signal also passes through a known blur first, the same equations give the filter that undoes the blur as far as the noise allows: Wiener deconvolution.
The maths behind it · orthogonal projection
The FIR Wiener filter is a projection. The error is orthogonal to every input sample the filter uses, : this is the orthogonality principle. The Wiener–Hopf equations are the normal equations of that projection, with averages playing the part of inner products.
The maths behind it · the conditional mean
In statistics, the Wiener filter is the best linear predictor of from the observations. When everything is Gaussian it is also the best predictor of any kind, the conditional mean . Its error is the variance that the observations leave unexplained.
Worked example
1. The gain by hand. At noise power 1 and , , so . At , , so . These are the gains in the first instrument’s middle caption.
2. One tap by hand. With one tap the Wiener–Hopf equations shrink to , that is . So , and , half the noise power. On the record the measured error is 0.464.
3. Two taps by hand. The equations are and . Taking 0.45 times the first from the second gives , so and . Then , which equals , as the noise is white.
4. The non-causal limit. At noise power 1 the non-causal filter reaches , which is −6.6 dB. The best causal filter reaches 0.304, or −5.18 dB.
Where you’ll meet this
Many speech enhancers apply a Wiener gain to each bin of the short-time spectrum of Spectrograms & the STFT (15.5). They measure the noise’s PSD in the pauses between words. Images blurred by a known lens or motion are restored by Wiener deconvolution.
In a receiver, the channel smears each symbol into its neighbours. An equaliser built from these same equations undoes the smear, in Channels and equalisation (33.2). The Wiener filter is also the target that the adaptive filters of Adaptive filters: LMS (26.3) learn. The Kalman filter of RLS and the Kalman filter (26.4) extends it to signals whose statistics change.
Compare it with Matched filters and detection (26.1). The matched filter asks whether a known pulse is there, and its output need not look like the pulse. The Wiener filter wants the waveform itself, as close as it can get.
For more, see N. Wiener, Extrapolation, Interpolation, and Smoothing of Stationary Time Series (1949); S. Haykin, Adaptive Filter Theory (5th ed., 2014), ch. 2; and J. G. Proakis and D. G. Manolakis, Digital Signal Processing, ch. 12. In SciPy, solve_toeplitz(c, r) gives the FIR Wiener taps, with c holding to and r the cross-correlations.
Reference card
| Quantity | Formula | Notes |
|---|---|---|
| Observation | signal and noise uncorrelated | |
| Error | , | mean square |
| Non-causal Wiener | signal’s share of the power | |
| Its error | below both lazy filters | |
| FIR Wiener | Wiener–Hopf, Toeplitz | |
| Entries | ; | matrix; right side |
| Its error | for white noise | |
| Orthogonality | at the best taps |