The obvious way to get a spectrum from a recording is to square its DFT and divide by the number of samples. Here that is tried on white noise, whose true spectrum is flat at 0 dB, with 64, 256 and then 1024 samples. Watch the dots: they get denser, but no tighter.
More data, same scatter
The periodogram of white Gaussian noise of variance 1 (seed 251; measured 1.026 over all 1024 draws), the first N samples.
64 samples, 31 bins: the dots scatter from −16.5 dB to 7.3 dB about the true 0 dB.
Describe this picture
One panel: the periodogram of white Gaussian noise of variance 1 (seed 251; measured 1.026 over all 1024 draws), for the first samples, in dB from −40 to 10 against Ω from 0 to π rad/sample. Each bin is one dot. A solid level at 0 dB is the true PSD, and a shaded band from −12.9 dB to 4.8 dB holds 90 % of values; a key above the panel names both. The readouts are , the mean and spread ÷ mean. There is no control. At , 31 bins, the dots scatter from −16.5 dB to 7.3 dB about the true 0 dB: mean 1.144, spread ÷ mean 1.099. At 256 there are four times the points, and they scatter at least as widely, from −25.6 dB to 7.8 dB: 1.080 and 0.944. At 1024, 90.4 % of the 511 dots lie in the 90 % band, and spread ÷ mean is 0.963, with mean 1.025. The end caption says: “More data has not steadied the estimate: the periodogram is not consistent.”
A spectrum from one record
Power spectral density (24.4) gave a random process a spectrum, the PSD . In “Memory in time, narrowness in frequency” it came from a formula, because the process was known. In practice you have a recording, not a formula, and you want its spectrum anyway.
A recording is one realisation, in the words of Random processes (24.2). As in “Stationary, and ergodic”, I assume the process is ergodic, so that one long recording can stand for the whole ensemble. This page asks how well it does.
Take the samples to and their DTFT, the sum of turned samples of “Samples, turned and added” in The DTFT (12.2):
Square its size and divide by . That is the periodogram of the samples:
Arthur Schuster gave it that name in 1898, when he used it to hunt for hidden periods in meteorological records. The letter is the average power of How big is a signal (1.3); the subscript and the argument mark the periodogram.
Why divide by ? At the bins, , the DTFT is the DFT , as Sampling the spectrum (13.1) showed. By “Energy in time and in bins” in The DFT (13.2), . Divide both sides by : the values of average to the mean square of the samples, their power.
So the periodogram spreads the record’s power over frequency, as a PSD does. Its area works out too: is the same mean square, just as 24.4’s PSD integrates to .
To read it in volts squared per hertz, one-sided, as in “Why the floor depends on N” of Reading a spectrum: scaling and units (15.4), multiply it by . Bins 0 and are not doubled. That is 15.4’s PSD formula with no window, where every is 1 and .
Bias, variance and consistency
The periodogram is a guess, and a guess needs words. An estimator turns data into a guess at something fixed. I write any estimate of the PSD as , with the hat of Random variables for signals (24.1). The periodogram is one such estimate.
A new record gives a new guess, so an estimate is itself a random variable, with a mean and a spread. Its bias is how far its average lies from the truth. Its variance is how widely it scatters about its own average. It is consistent if both shrink to 0 as grows: then a long enough record pins the truth down.
Think of a bathroom scale. One that always reads 2 kg heavy is biased. One whose needle wobbles has variance. A good scale has neither, and standing on it longer should only help.
I measure scatter by the spread, the standard deviation of “Draws pile up into a density” in 24.1. Dividing it by the mean gives the readout “spread ÷ mean”: 0.1 means the guesses wander by a tenth of their size.
Let’s test the periodogram where the truth is known. White Gaussian noise of variance 1 has a flat PSD (24.4): , which is 0 dB as a power (, 1.3), at every . Its periodogram averages exactly to the truth, as the next section shows. So whatever you see is scatter, and you would hope it shrinks as grows.
More data, same scatter
The picture at the top of the page draws the periodogram of white noise for the first 64, 256 and 1024 samples of one record, one dot per bin. Its shaded band holds 90 % of values; the next section says where the band’s edges come from. Spread ÷ mean reads 1.099, 0.944 and 0.963 for the three lengths.
Its readouts are taken over bins 1 to N/2 − 1, strictly between 0 and π. For a real signal the bins above mirror those below (“Bins in hertz, and the mirror”, 13.2), and bins 0 and are real numbers that follow a different law.
Notice that the dots get denser but no tighter. Spread ÷ mean stays near 1 at every : each dot is wrong by about the size of the truth itself. It is like asking more people a question that each of them answers at random. You get more answers, not better ones.
The mean readout stays near 1. It sits close to each record’s own mean square, 1.128 for the first 64 samples and 1.026 for all 1024, because by 13.2’s energy rule all values average to exactly that.
Why the scatter stays
In 12.2 each sample became an arrow, turned back by , and the arrows were added nose to tail. With noise, the arrows have random lengths and signs. Their sum is the end point of a random walk of steps.
How far does it get on average? Multiply out and take the average term by term, since averages add (24.2). Each product averages to , the autocorrelation of 24.2. For white noise that is when and 0 otherwise:
So the squared length grows like on average, and dividing by makes the average right: at every and every . For white noise the periodogram has no bias.
But one walk’s end is random. Sometimes the steps nearly cancel and sometimes they line up. The squared length scatters about as widely as its own average, and both grow like . Dividing by scales them together, so spread ÷ mean stays near 1 however long the record is.
For Gaussian noise the law of that scatter is known. I assume it here; statistics derives it. At every bin strictly between 0 and , follows the exponential law: the share of values above times the truth is ,
Its spread equals its mean, so the law’s spread ÷ mean is 1. The measured 1.099, 0.944 and 0.963 scatter about it, because 31 to 511 dots only estimate the law.
Now read the band from it. A share 0.05 lies above times the truth, which is 4.8 dB. A share 0.05 lies below times the truth, where , which is dB. So 90 % of values lie between dB and 4.8 dB of the truth, at any , and the 511 dots put 90.4 % there.
The band is lopsided: it reaches 12.9 dB down but only 4.8 dB up. On a dB axis the low tail is long, and more dots reach further into it: the lowest is dB at and dB at 1024.
Even the middle sits low. Half of all values lie below of the truth, dB, and 63.2 % below the truth itself, since . Of the 511 dots, 313 do, 61.3 %.
So the periodogram fails the last test. Its bias is 0 for white noise, but its spread stays at the size of itself whatever is, so its variance never shrinks: it is not consistent. A longer record buys more bins, closer together, each as unreliable as before.
Short records blur the peak
White noise is the easy case, because its PSD is flat. A PSD with a sharp peak shows the second fault. First I need the periodogram’s average for any process, not only white noise.
Multiply out as before, but now group the products by their lag . There are of them at lag .
Added up, they are times 24.2’s estimate from “White noise forgets, coloured noise remembers”, which divides by the number of products. A negative lag uses the same products as the positive one, so . Dividing by ,
The periodogram is the DTFT of the estimated autocorrelation, with its lags weighted by a triangle.
Each product at lag averages to the true , so does too, and averages add:
Without the triangle this sum is the DTFT of , which 24.4 showed is the true PSD. With it the far lags are cut down, so the average is the truth seen through a window. For white noise only is left, and the triangle is 1 there: no bias, as found above.
Which window? Put the record window itself into the lag form: ones, the cut of “Cut a tone short, and its arrows spread” in Properties of the DTFT (12.3). Every product is 1, so every is 1, and the right side is the DTFT of the triangle alone. The left side is , with 12.3’s :
A product of two sequences in the lag domain is a periodic convolution of their transforms, by “Multiplication in time is periodic convolution in frequency” (12.3). So the average periodogram is the true PSD smoothed by this kernel:
The kernel is never negative, and its own area, , is 1, the triangle’s value at . So the average at each is a weighted average of the true PSD nearby. An average can never exceed the largest value it averages, or fall below the smallest, so a peak can only drop and a valley can only fill.
How wide is the smoothing? has the zeros of , so its main lobe is 12.3’s wide. Its side lobes are the leakage of “A tone between bins leaks into all of them” in Windowing & spectral leakage (15.1).
A noise with a sharp peak
The test process is the ringing recursion of “Where the pair sits is how h[n] moves” in Transfer functions, poles & zeros (16.3), driven by white noise of variance :
It is 16.3’s pair with radius at the angle : and , so its poles are . A sample depends on its last two, so this is an AR(2) process, 24.2’s AR(1) with one more past value.
Its PSD follows from of 24.4, with white noise going in. Here , 16.3’s transfer function on the unit circle, so
It peaks at 22.1 dB at , just below the poles’ angle, and falls to dB at . The peak is narrow. The rule of “Poles behind the zeros narrow the notch” in Resonators, notches and combs (17.4), a width of about rad/sample, holds well, and the dB points are 0.103 rad/sample apart. Its power is , and 90 % of it lies between and on either side of 0.
Think of a photograph taken through a small aperture. The bright spot spreads out, and its light reaches the dark corners.
Short records blur the peak
AR(2) noise, poles 0.95∠±0.3π: the true PSD and the periodogram's average for N samples.
32 samples: on average the peak reaches only 19.1 dB (true 22.1 dB), and leakage lifts the valley at π to −5.8 dB (true −9.6 dB).
Describe this picture
One panel for AR(2) noise with poles 0.95∠±0.3π: PSD in dB from −15 to 25 against Ω from 0 to π rad/sample. A key names the dashed curve “true PSD” and the solid curve “average periodogram” for samples. The readouts are and the average periodogram’s peak and value at Ω = π. There is no control. At the peak reaches only 19.1 dB (true 22.1 dB), and leakage lifts the valley at π to −5.8 dB (true −9.6 dB). The curve morphs to : 21.3 dB and −8.2 dB. Then to : 22.0 dB and −9.4 dB, within 0.3 dB of the truth, and the end caption says that the bias shrinks as grows.
Watch the solid curve close in on the dashed one as grows. It is the exact average of the sum above, not one noisy record.
Notice the first frame: at 32 samples the valley fills by 3.8 dB, more than the peak drops, 2.9 dB.
Both follow from the kernel’s width. At its main lobe is rad/sample wide, nearly four times the peak’s 0.103, so the peak is smeared over its neighbours. At 128 the main lobe, 0.098 rad/sample, is about as wide as the peak, and the peak falls only 0.7 dB short. At 1024 it is 0.012 rad/sample, and the peak is within 0.1 dB.
The valley at sits 31.7 dB below the peak, so even the faint side lobes of the kernel carry enough of the peak’s power there to matter. At , 52 % of the average at comes from the band around the peak, to on either side of 0. That is leakage, the same as in 15.1, only now it leaks noise power instead of a tone.
One root, two cures
Both faults come from the lag form. One record gives one estimate per lag, and the far lags rest on few products, down to a single one at . Every one of them enters the sum, noise and all: that is the variance. The triangle tapers the far lags away: that is the bias.
The cures follow. A smoother window than the plain cut has lower side lobes, as in Window functions compared (15.2), and so leaks less. Averaging the periodograms of many pieces of the record cuts the variance; that is the subject of Averaged periodograms: Bartlett and Welch (25.2).
The maths behind it · the chi-square law
Bias, variance and consistency are the properties statistics asks of an estimator. Each periodogram bin of Gaussian noise strictly between 0 and is a scaled chi-square variable with 2 degrees of freedom, the exponential law behind the first instrument’s band. Bins 0 and have one degree of freedom.
The maths behind it · orthonormal coordinates
The unitary DFT of The DFT as a matrix (13.5) gives the record’s coordinates along orthogonal directions, and is the squared size of coordinate . A longer record adds directions, not data per direction. Each bin is one random coordinate, so it scatters like one draw, not like an average.
Worked example
1. A periodogram by hand. Take = 1, 2, 1, 0, so . The DFT is = 4, , 0, , and the periodogram at the bins is = 4, 1, 0, 1. These average to 1.5, the mean square .
Now by lags. 24.2’s estimates are , , and . With the triangle’s weights 1, 3/4, 2/4 and 1/4, the lag form gives
At = 0, and this is 4, 1 and 0, as before. Between bins it gives new values: at it is 2.25.
Scale it as a one-sided PSD for Hz, with in volts. The bins are 0, 250 and 500 Hz. Bin 0 is V²/Hz, bin 1 is doubled, V²/Hz, and bin 2 is 0. Times the bin width, 250 Hz, they add to V², the mean square again.
SciPy’s periodogram(x, fs=1000, detrend=False) returns exactly these numbers. By default it first subtracts the mean, and then bin 0 reads 0.
2. White noise. For = 64, 256 and 1024 seeded samples, spread ÷ mean is 1.099, 0.944 and 0.963. The exponential law says 1 at every , and puts 90 % of the values between dB and 4.8 dB of the truth.
3. AR(2) bias. At 32 samples the average periodogram’s peak is 19.1 dB against a true 22.1 dB, 2.9 dB short. At it is dB against a true dB, 3.8 dB high. At 1024 samples both are within 0.3 dB.
Where you’ll meet this
The periodogram is the raw material of almost every spectral estimate. Averaged periodograms: Bartlett and Welch (25.2) averages the periodograms of segments. Each column of a spectrogram in Spectrograms & the STFT (15.5) is, up to a scale factor, the periodogram of one windowed frame.
Alone, it serves best for strong tones in weak noise. A tone of amplitude on a bin gives (13.2), so its periodogram value grows with , while the noise bins only scatter about their truth. At a tone of amplitude 1 in white noise of variance 1 reads 24.1 dB, 19.3 dB above the top of the 90 % band.
For more, see A. V. Oppenheim and R. W. Schafer, Discrete-Time Signal Processing, ch. 10; J. G. Proakis and D. G. Manolakis, Digital Signal Processing, ch. 14; and P. Stoica and R. Moses, Spectral Analysis of Signals (2005), ch. 2. In NumPy, np.abs(np.fft.fft(x))**2 / len(x) gives at the bins, and SciPy’s scipy.signal.periodogram gives the one-sided density. Models that sidestep the periodogram’s faults come in Parametric models and linear prediction (25.3).
Reference card
| Quantity | Formula | Notes |
|---|---|---|
| Periodogram | a PSD estimate from samples | |
| One-sided density | in V²/Hz | bins 0 and not doubled (15.4) |
| Lag form | 24.2’s estimates, tapered by a triangle | |
| Its average | true PSD smoothed by : bias | |
| Its scatter | spread ≈ the PSD itself, at any | not consistent |
| White Gaussian noise | 90 % of bins between −12.9 dB and 4.8 dB of the truth | exponential law |
| Estimator | bias: average minus truth; variance: scatter | consistent: both shrink to 0 as grows |