Skip to content

Random processes

A random signal is one draw from an ensemble. Average across the ensemble or along one recording, and watch white noise forget and coloured noise remember.

Before this24.1 · 3 more
Chapter 24 · Lesson 2 of 4

First, the picture

Here are eight recordings of the same kind of noise, drawn one above the other. You can average across them at one instant, or along one recording. Watch the two averages: for plain noise both sit near 0, but once each recording gets a level of its own, the average along one recording follows its level.

Across the ensemble, along a realisation

8 of 200 realisations, 256 samples each (seed 242). A: white noise. B: the same noise plus a level drawn once per realisation.

Process A, white noise. Across all 200 realisations at n = 30 the average is −0.121; along realisation 8 it is 0.065: both near 0, the process's mean.

process
A: white noise
across 200 at n = 30
−0.121
along realisation 8
0.065
0.00 / 14.00 s
Describe this picture

Eight stacked strips, realisations 1 to 8 of 200, each a thin trace of 256 samples (seed 242), named “r = 1” to “r = 8”. They share the sample axis, 0 to 255, and each runs from −5 to 5. Realisation 8 is picked out, because the second readout follows it, and a dashed vertical line labelled “n = 30” marks the instant where the first readout looks down the ensemble. The readouts are the process, the average across 200 at n=30n=30 and the average along realisation 8. There is no control. Process A, white noise, comes first. Dots light up where the dashed line crosses the strips, a line marks realisation 8’s own average, and a short bar beside each strip shows that realisation’s average, drawn 2.5 times taller than the strip; in process A all eight bars sit at 0. Across the ensemble the average is −0.121, and along realisation 8 it is 0.065: both near 0, the process’s mean. Then each realisation slides up or down by its own level, which turns it into process B. Each strip gets a dotted line at its level, with the band between 0 and the level shaded, and the bars now sit at different heights; a key names the lines “time average” and “level”. Across the ensemble the average is still near 0 (−0.082), but along realisation 8 it is −1.357, close to that realisation’s own level, −1.422. The end caption says: “One recording cannot reveal the ensemble mean: B is not ergodic.”

A whole signal at random

In Random variables for signals (24.1) each draw gave one number. Noise on a wire is more than one number, though. It is a whole signal, with a new value at every sample, and every recording of it comes out different.

So let’s draw a whole signal at once. A random process hands out whole signals x[n]x[n] at random. Each signal it hands out is a realisation, and I number them: x(r)[n]x^{(r)}[n] is realisation rr. The collection of all the realisations it could hand out is the ensemble.

Picture 200 identical microphones in 200 identical quiet rooms, all recording at once. Each recording is a realisation. Line them up one above the other, and you are looking at a slice of the ensemble.

From 24.1 I assume what “Draws pile up into a density” and “A cloud that leans: correlation” teach: a density, the mean μx\mu_x, the variance σx2\sigma_x^2, estimates such as μ^x\hat\mu_x, independence and the correlation coefficient. I add one fact: averages add, so the mean of a sum is the sum of the means. Statistics proves these; this page only uses them.

Fix one instant nn and look down the ensemble. Every realisation has a value there, so x[n]x[n] is a random variable of 24.1, with its own density and its own mean μx[n]\mu_x[n]. The mean carries the index because it may change with nn.

That gives two ways to average. Across the ensemble means many realisations at one instant. With 200 realisations it estimates μx[n]\mu_x[n]:

μ^x[n]=1200∑r=1200x(r)[n].\hat\mu_x[n]=\frac1{200}\sum_{r=1}^{200}x^{(r)}[n].

Along a realisation means one realisation at many instants. That is the mean xˉ\bar x of “Four everyday sizes” in How big is a signal (1.3), the running average of “Average per sample” without the squaring:

xˉ=1N∑n=0N−1x(r)[n].\bar x=\frac1N\sum_{n=0}^{N-1}x^{(r)}[n].

When do the two agree? A process whose averages along one long realisation equal its averages across the ensemble is called ergodic. This matters, because in practice you have one recording, not 200.

The picture at the top of the page compares two processes. Process A is white noise of variance 1: every sample is a fresh Gaussian draw of 24.1, independent of all the others. The name gets its exact meaning later on this page.

Process B is the same noise plus a level: one more Gaussian draw of variance 1, made once per realisation and added to every sample of it. Within one realisation the level never changes; from one realisation to the next it does.

Across the ensemble, along a realisation

In process A, the picture’s two averages came out at −0.121 across all 200 realisations at n=30n=30, and 0.065 along realisation 8. Both are near 0, the process’s mean. In process B they are −0.082 across the ensemble, still near 0, but −1.357 along realisation 8, close to that realisation’s own level, −1.422. One recording cannot reveal the ensemble mean: B is not ergodic.

Let’s check what “near 0” means. At n=30n=30 the 200 values are independent draws of variance 1. By “A longer average: less noise, more delay” in Simple smoothing filters (18.2), averaging 200 of them leaves a spread of 1/200=0.0711/\sqrt{200}=0.071.

So −0.121-0.121 is 1.7 spreads from 0. Along realisation 8, 256 samples leave a spread of 1/256=0.06251/\sqrt{256}=0.0625, and 0.065 is about one spread from 0. Both are as near to 0 as so few draws allow.

In B each sample is noise plus a level, two independent draws of variance 1. Their variances add (24.1), so each sample has variance 2, and an average of 200 has spread 2/200=0.1\sqrt{2/200}=0.1. B’s −0.082-0.082 is well within it. It differs from A’s value only by the average of the 200 levels, 0.040.

Along realisation 8, the noise averages away as it did in A, but the level stays. B’s time average is A’s, 0.065, plus the level, −1.422-1.422, which gives −1.357-1.357. A longer recording would not help: the noise part would shrink towards 0, and the time average would close in on −1.422-1.422, not on 0.

The other realisations agree. Over all 200, A’s time averages scatter with a spread of 0.065, and B’s with a spread of 0.925, close to the levels’ own spread, 0.929. Each B recording reports its own level, not the process’s mean.

It is like the average height of everyone in a town today, against one person’s height averaged over a week. The first is the town’s; the second is that person’s, however long you measure.

Stationary, and ergodic

Neither process drifts as nn goes on. To say that properly, I need a second average across the ensemble: the mean of the product of the values at two instants, the correlation

Rx[n1,n2]=E{x[n1] x[n2]}.R_x[n_1,n_2]=\mathbb{E}\{x[n_1]\,x[n_2]\}.

When both values have mean 0 and the same variance, dividing by that variance gives 24.1’s correlation coefficient of x[n1]x[n_1] and x[n2]x[n_2].

A process is wide-sense stationary, WSS for short, when two things hold. Its mean μx[n]\mu_x[n] is the same at every nn, and Rx[n1,n2]R_x[n_1,n_2] depends only on the lag ℓ=n1−n2\ell=n_1-n_2, how far apart the instants are. Then I write it Rx[ℓ]R_x[\ell]. In words, its averages do not care when you look, only how far apart.

Process A has mean 0 at every nn. Two different samples are independent, so their product averages to 0; a sample times itself averages to its variance, 1. So Rx[ℓ]R_x[\ell] is 1 at ℓ=0\ell=0 and 0 at every other lag.

Process B has mean 0+0=00+0=0 at every nn. Its product at two instants is (noise + level) times (noise + level). Averages add, so take the four products one at a time. The noise parts are independent of each other and of the level, so only noise times noise at the same instant, and level times level, survive.

So for B, Rx[ℓ]R_x[\ell] is 2 at ℓ=0\ell=0 and 1 at every other lag. The level is shared by every pair of instants, however far apart.

Both depend only on the lag, so both processes are WSS. Yet only A is ergodic. B’s correlation never dies away: samples 1000 apart still share their level. Each realisation never forgets its level, so it never sees the rest of the ensemble.

So stationary does not mean ergodic. In practice there is often only one recording, and analysing it alone means assuming ergodicity: its time averages are taken to stand for the ensemble’s. From here on I assume it too.

The maths behind it · stationary time series

This is what statistics calls a time series: a stationary series, its autocorrelation function (ACF), and the AR(1) model of the next section. For a process with mean 0, that track writes XtX_t for x[n]x[n] and ρ(h)\rho(h) for Rx[ℓ]/Rx[0]R_x[\ell]/R_x[0].

White noise forgets, coloured noise remembers

For a WSS process, the autocorrelation

Rx[ℓ]=E{x[n] x[n−ℓ]}R_x[\ell]=\mathbb{E}\{x[n]\,x[n-\ell]\}

says how alike, on average, two samples ℓ\ell apart are. At lag 0 it is the mean square, which is the variance when the mean is 0. It is symmetric, Rx[−ℓ]=Rx[ℓ]R_x[-\ell]=R_x[\ell], because a pair read the other way round is the same pair.

If the process is ergodic, one long realisation of NN samples is enough. Average the products along it:

R^x[ℓ]=1N−ℓ∑n=ℓN−1x[n] x[n−ℓ].\hat R_x[\ell]=\frac{1}{N-\ell}\sum_{n=\ell}^{N-1}x[n]\,x[n-\ell].

There are N−ℓN-\ell products, from n=ℓn=\ell to N−1N-1, and I divide by their number. The hat means “estimated”, as in 24.1. Some texts divide by NN instead; for the lags on this page that changes the values by less than 0.001.

White noise has Rx[ℓ]=0R_x[\ell]=0 at every ℓ≠0\ell\ne0: no two different samples are correlated. Process A is white. Power spectral density (24.4) shows why “white” fits.

I write white noise as v[n]v[n], with variance σv2\sigma_v^2. Dither (11.2) used the same letter for dither, which is white noise too. So Rv[ℓ]R_v[\ell] is σv2\sigma_v^2 at ℓ=0\ell=0 and 0 elsewhere, which is σv2 δ[ℓ]\sigma_v^2\,\delta[\ell] with the unit impulse of Impulse, step and ramp (3.1).

Coloured noise is noise whose neighbouring samples resemble each other. A simple recipe feeds white noise into a loop: the leaky integrator of “With a loop and without one” in Difference equations (6.1). In “A 9-point average and its one-multiply twin”, 18.2 calls the same loop the EMA. With white noise as its input it reads

x[n]=a x[n−1]+v[n],∣a∣<1.x[n]=a\,x[n-1]+v[n],\qquad\lvert a\rvert<1.

The output depends on its own last value, so this is an autoregressive process of order 1, AR(1) for short. The EMA of 18.2 is the same loop with the input scaled by 1−a1-a.

How big is x[n]x[n]? The value x[n−1]{x[n-1]} was built from noise up to n−1n-1 only, so it is independent of the fresh v[n]v[n]. Their variances add, and multiplying by aa multiplies a variance by a2a^2. In a stationary process x[n]x[n] and x[n−1]{x[n-1]} have the same variance, so

σx2=a2σx2+σv2,σx2=σv21−a2.\sigma_x^2=a^2\sigma_x^2+\sigma_v^2,\qquad\sigma_x^2=\frac{\sigma_v^2}{1-a^2}.

Now the likeness. Multiply both sides of the recursion by x[n−ℓ]{x[n-\ell]} for some ℓ≥1\ell\ge1, and average. The term with v[n]v[n] averages to 0, because v[n]v[n] has mean 0 and is independent of the past. What is left is Rx[ℓ]=a Rx[ℓ−1]R_x[\ell]=a\,{R_x[\ell-1]}: each step of lag multiplies the likeness by aa. With the symmetry,

Rx[ℓ]=σv21−a2 a∣ℓ∣.R_x[\ell]=\frac{\sigma_v^2}{1-a^2}\,a^{\lvert\ell\rvert}.

The instrument takes a=0.9a=0.9. With white noise of variance 1 as input, the output’s variance would be 1/(1−0.81)=5.2631/(1-0.81)=5.263. So the input is scaled by 0.19=0.4359\sqrt{0.19}=0.4359, which scales its variance to 0.19 and brings the output’s variance to 1. Then Rx[ℓ]=0.9∣ℓ∣R_x[\ell]=0.9^{\lvert\ell\rvert}, and white and coloured noise are compared at equal power.

Divided by Rx[0]R_x[0], the autocorrelation is 24.1’s correlation coefficient of two samples ℓ\ell apart. Next-door samples of this noise have a correlation coefficient of 0.9: in 24.1’s picture, a cloud of pairs leaning hard along a line. Think of coin tosses against the daily temperature: yesterday’s toss says nothing about today’s, but yesterday’s temperature says a lot.

White noise forgets, coloured noise remembers

Variance-1 white noise and AR(1) noise x[n] = 0.9x[n−1] + √0.19·v[n]; autocorrelation estimated from 4000 samples.

White noise: next-door samples are unrelated. Measured R_x[1] = 0.005, R_x[5] = −0.006: near zero apart from lag 0 (1.021).

R_x[1]
0.005
R_x[5]
−0.006
0.00 / 12.00 s
Describe this picture

Two stacked panels for variance-1 white noise and AR(1) noise x[n]=0.9x[n−1]+0.19 v[n]x[n]=0.9x[n-1]+\sqrt{0.19}\,v[n], with the autocorrelation estimated from 4000 samples. The first shows the first 64 samples as stems, from −3 to 3, and names the noise it shows, “white noise” or “AR(1) noise”. The second shows Rx[ℓ]R_x[\ell] from −0.2 to 1.1 against lag ℓ from 0 to 10: the estimates are stems with square heads labelled “measured”, and the formula’s values open rings labelled “formula”. The readouts are the estimates Rx[1]R_x[1] and Rx[5]R_x[5]. There is no control. White noise comes first: next-door samples are unrelated, and Rx[1]=0.005R_x[1]=0.005 and Rx[5]=−0.006R_x[5]=-0.006, near zero apart from lag 0 (1.021). Then the samples and stems change to the same white noise through the AR(1) loop, and the formula’s rings appear. Neighbours now resemble each other: Rx[1]=0.914R_x[1]=0.914 and Rx[5]=0.593R_x[5]=0.593, beside the formula’s 0.9 and 0.95=0.5900.9^5=0.590.

Watch the stems at lags 1 to 10: near 0 for white noise, fading slowly for coloured noise. Look at the white noise first. Its lag-0 value, 1.021, is the measured variance of these 4000 draws, close to the nominal 1. At lags 1 to 10 the estimates are not quite 0: the largest is 0.021.

That is the scatter you should expect. Each estimate averages about 4000 products of unrelated draws, and by 18.2’s rule that leaves a spread of about 1/4000=0.0161/\sqrt{4000}=0.016.

Now the coloured noise. Its estimates sit close to the formula: 0.914 against 0.9 at lag 1, and 0.593 against 0.590 at lag 5. At lag 10 the noise is still 0.366 alike, against the formula’s 0.910=0.3490.9^{10}=0.349. The likeness fades step by step, and it first drops below a half at lag 7, where the formula gives 0.478.

The coloured estimates stray further from their formula than the white ones do, because neighbouring products are alike too, and alike errors do not average away as fast. The periodogram (25.1) comes back to how far an estimate can stray.

The maths behind it · Toeplitz matrices

The autocorrelations of a stationary process fill a Toeplitz matrix R\mathbf{R}, one whose entries are constant along every diagonal: entry (p,q)(p,q) is Rx[p−q]R_x[p-q]. It is symmetric and positive semidefinite. It comes back in the Yule–Walker equations of Parametric models and linear prediction (25.3) and the Wiener–Hopf equations of The Wiener filter (26.2).

Stationary for a while

Real signals drift. Speech changes from one sound to the next, and a radio channel fades as a car drives along. So a recording is treated as stationary over a window short enough that its statistics do not drift.

For speech that window is about 20 to 30 ms. That is why “Reading a chirp” in Spectrograms & the STFT (15.5) used a 25 ms window for speech.

Worked example

1. Two averages. At n=30n=30 the 200 realisations of process A average −0.121-0.121, and realisation 8 averages 0.065 along its 256 samples. Their expected spreads are 1/200=0.0711/\sqrt{200}=0.071 and 1/256=0.06251/\sqrt{256}=0.0625, so both are near the mean, 0. In process B, realisation 8 averages 0.065+(−1.422)=−1.3570.065+(-1.422)=-1.357, its own level plus A’s average, while the ensemble still gives −0.082-0.082.

2. An estimate by hand. Take xx = 2, 1, −1, −2, so N=4N=4. At lag 0 there are four products, 4+1+1+4=104+1+1+4=10, so R^x[0]=10/4=2.5\hat R_x[0]=10/4=2.5. At lag 1 there are three, 1⋅2+(−1)⋅1+(−2)⋅(−1)=31\cdot2+(-1)\cdot1+(-2)\cdot(-1)=3, so R^x[1]=3/3=1\hat R_x[1]=3/3=1. Next-door samples here are 1/2.5=0.41/2.5=0.4 alike.

3. AR(1) with a=0.9a=0.9. With input variance 1 the output’s variance would be 1/(1−0.81)=5.2631/(1-0.81)=5.263. Scaling the input by 0.19=0.4359\sqrt{0.19}=0.4359 brings it to 1, and then Rx[5]=0.95=0.590R_x[5]=0.9^5=0.590. The 4000 seeded samples measure 0.593.

Where you’ll meet this

A noise model in a receiver or a sensor is a random process. Thermal noise in a resistor or an amplifier is modelled as white Gaussian noise, and its power sets the noise floor. A slowly wandering sensor offset is modelled as coloured noise, often AR(1).

The next pages build on this one. Correlation (24.3) compares two signals by the same products, Power spectral density (24.4) turns Rx[ℓ]R_x[\ell] into a spectrum, and The periodogram (25.1) estimates that spectrum from one recording.

AR models describe speech in Parametric models and linear prediction (25.3). The Wiener filter of The Wiener filter (26.2) and the Kalman filter of RLS and the Kalman filter (26.4) start from models of the signal and the noise as random processes.

For more, see A. V. Oppenheim and R. W. Schafer, Discrete-Time Signal Processing, appendix A; A. Papoulis and S. U. Pillai, Probability, Random Variables and Stochastic Processes (4th ed., 2002), ch. 9–11; and J. G. Proakis and D. G. Manolakis, Digital Signal Processing, ch. 12. In NumPy, np.dot(x[l:], x[:N-l]) / (N - l) gives this page’s estimate at lag l.

Reference card

QuantityFormulaNotes
Ensemble meanμx[n]=E{x[n]}\mu_x[n]=\mathbb{E}\{x[n]\}across realisations, at one nn
Time averagexˉ=1N∑nx(r)[n]\bar x=\frac1N\sum_nx^{(r)}[n]along one realisation
AutocorrelationRx[ℓ]=E{x[n] x[n−ℓ]}R_x[\ell]=\mathbb{E}\{x[n]\,x[n-\ell]\}WSS: depends on ℓ\ell only; Rx[−ℓ]=Rx[ℓ]R_x[-\ell]=R_x[\ell]
Ergodictime averages = ensemble averagesone recording stands for the process
White noiseRv[ℓ]=σv2 δ[ℓ]R_v[\ell]=\sigma_v^2\,\delta[\ell]no memory
AR(1)x[n]=a x[n−1]+v[n]x[n]=a\,x[n-1]+v[n]coloured; ∣a∣<1\lvert a\rvert<1
AR(1) autocorrelationRx[ℓ]=σv21−a2a∣ℓ∣R_x[\ell]=\frac{\sigma_v^2}{1-a^2}a^{\lvert\ell\rvert}each lag step multiplies by aa
EstimateR^x[ℓ]=1N−ℓ∑n=ℓN−1x[n] x[n−ℓ]\hat R_x[\ell]=\frac1{N-\ell}\sum_{n=\ell}^{N-1}x[n]\,x[n-\ell]from one realisation

End of lesson 24.2

Where to go next.

Phasorium
LibraryEvery lesson, in order

Parts

About Phasorium
Look