Skip to content

Dither

Add a little noise before rounding and see why the error stops following the signal, why triangular noise keeps the hiss steady, and what it costs.

Before thisQuantization & noise (11.1)

3 more before it

What a signal is (1.1), How big is a signal (1.3), Signals as sums of sinusoids (7.1)

Before this11.1 · 3 more
Chapter 11 · Lesson 2 of 3

First, the picture

Take a smooth gradient from dark to light, 256 pixels wide, and store it with only 8 shades. Watch hard bands appear where the shade jumps, and then watch them dissolve into fine grain.

Bands in a smooth sky

A gradient from dark to light, 256 pixels wide.

A smooth sky, from dark to light.

shades available
any
0.00 / 9.00 s
Describe this picture

The picture has no controls and plays once. It shows a smooth gradient, then the same gradient stored with 8 shades, in hard bands, then the 8-shade gradient with dither, where the bands become fine grain. Under it a strip plots brightness against position. The readout “shades available” first shows “any”, then “8”.

Bands in a smooth sky

Quantization and noise (11.1) ended on a quiet sine. The error of a quiet or simple signal repeats with the signal, so it does not sound like hiss. It sounds like extra harmonics, a distortion that follows the music. This page is about a cure that sounds strange at first: add a little noise before you round.

I started with a picture, because a picture shows the problem fastest. In What a signal is (1.1) you saw that one row of a picture is a signal: brightness at each pixel position. Store a smooth gradient with only 8 shades and the rounding error is large, and it is not random. It is the same small sawtooth each time the brightness crosses a level, so your eye sees it as banding: hard edges where the shade jumps.

In the last step of the picture at the top I add dither: a small random value, spread evenly over plus or minus half a shade, added to each pixel before rounding. The picture still has only 8 shades, but the bands dissolve into fine grain.

Look at the strip of brightness against position under the picture at the top. With dither, the rounded values flicker between neighbouring shades instead of switching once. That flicker is what your eye averages into a smooth gradient. A newspaper photo works the same way: each dot is black or white, and the grey you see is the average over a small area.

Right on average

The picture suggests an idea about averages, and I want you to see it with numbers. Dither adds noise, so you would expect the output to get worse. Before you run the next instrument, decide: I will add noise to a quiet signal before rounding it. Will the result get better or worse?

Here is the setup. The input is a sine 1.4 steps tall, with 32 samples per period. A step is the distance Δ\Delta between neighbouring levels, and the levels sit at whole steps: …,−1,0,1,…\ldots,-1,0,1,\ldots Without dither, the quantizer Q{⋅}Q\{\cdot\} rounds each sample to the nearest level. Every period gives the same output, a fixed staircase.

With dither, we round x[n]+v[n]x[n]+v[n] instead, where v[n]v[n] is a fresh random value for each sample, spread evenly between −12-\tfrac12 and +12+\tfrac12 step. This is rectangular dither. The question is what one run does to a sample, and what many runs do on average.

Right on average

A sine 1.4 steps tall, 32 samples per period, rounded to whole steps.

One run each. Without dither: a fixed staircase. With dither: a noisier staircase.

runs averaged
1

largest gap from the input (steps)

without dither
0.464
with dither
0.778
0.00 / 12.00 s
Describe this picture

Two stacked panels, “without dither” and “with dither”, each drawing the input as a dashed line labelled “input” and the average of the runs as stems. The picture plays by itself, and the readout “runs averaged” climbs from 1 to 100. Under the heading “largest gap from the input (steps)”, the readouts “without dither” and “with dither” give each panel’s largest gap. Two buttons, “Hear without dither” and “Hear with dither”, play sound and do not change the picture.

With one run, the gaps read 0.464 without dither and 0.778 with it, so dither is worse for a single run. Watch the stems as the count rises. At 10 runs the gap with dither has fallen to 0.273, and at 100 runs to 0.102. Without dither the stems never move: every run is the same, so the gap stays at 0.464.

Here is why. Suppose the input sits a share pp of the way from one level to the next, so it is pp steps above the lower level, with 0≤p<10\le p<1. The sum x+vx+v spreads over ±12\pm\tfrac12 step around the input. The output rounds up when v≥12−pv\ge\tfrac12-p, a stretch of values of length pp out of a total length 1, so it rounds up in a share pp of the runs. So the output is the upper level in a share pp of the runs and the lower level in the rest.

Average those outputs and you get pp steps above the lower level, which is the input. The average of many runs closes in on the input. That is what “right on average” means.

Press Hear without dither and Hear with dither. Each plays that panel’s output for a single run, for 2 seconds: a 1.5 kHz tone at 48 kHz, with a step of 0.05 of full scale. The second has a steady hiss. I think of it this way: the error without dither repeats every period, so it is made of harmonics (see Signals as sums of sinusoids, 7.1), and the third harmonic of the error is 0.273 of a step. The error with dither never repeats, so it is hiss, which the ear tolerates far better than a tone that follows the music.

Dither has a cost. The rounding error’s mean square is Δ2/12\Delta^2/12, as in 11.1. The dither adds its own mean square, which for a value spread evenly over ±12\pm\tfrac12 step is also Δ2/12\Delta^2/12. Averaged over the input position, the total is Δ2/6\Delta^2/6: twice as much, which is 10log⁡102=3.0110\log_{10}2=3.01 dB more noise.

Rectangular or triangular dither

Rectangular dither fixes the average, but one problem remains. How much hiss you get still depends on where the input sits between two levels. Put the input exactly on a level and the dither never pushes it across the halfway point, so the output is exact and there is no error. Put it halfway between two levels and the error is as large as it gets.

Here is the mean square of the error for an input pp steps above a level. The output is up by 1−p1-p steps from the input with probability pp, and down by pp steps with probability 1−p1-p, so

e2‾=p(1−p) Δ2.\overline{e^2} = p(1-p)\,\Delta^2 .

It is zero at p=0p=0 and at p=1p=1, and largest at p=12p=\tfrac12, where it is 0.25 Δ20.25\,\Delta^2.

When the input is a slow fade, the hiss swells and shrinks with it. That is noise modulation, hiss that breathes with the music. A fading piano note is the classic case: the background hiss pulses along with the note.

The fix is triangular dither: the sum of two independent rectangular values. It spreads over ±1\pm1 step and is most often near 0. For this shape the mean square of the error is Δ2/4\Delta^2/4 for every input. Watch both kinds of dither as a steady input moves from one level to the next.

Rectangular or triangular dither

A steady input moving from one level to the next. Each value is averaged over all dither values.

Input exactly on a level. Rectangular dither leaves no error at all here; triangular dither adds its hiss anyway.

input level
0.00 steps
rectangular
0.000
triangular
0.250
0.00 / 10.00 s
Describe this picture

The mean square of the error, divided by Δ2\Delta^2, from 0 to 0.3, against the input level in steps above a level, from 0 to 1; the picture has no controls. Each value is averaged over all dither values, so these are exact expectations, not random draws. A solid curve is labelled “rectangular” and a dashed line “triangular”; a dot on the curve and a diamond on the line follow the current input. The readouts “input level”, “rectangular” and “triangular” start at 0.00, 0.000 and 0.250, both read 0.250 at 0.50, and at 1.00 the rectangular one is back at 0.000 while the triangular one still reads 0.250.

The rectangular curve is an arch from 0 up to 0.25 and back, and the triangular line stays flat at 0.25. A steady hiss does not breathe, and that is the reason triangular dither is the usual choice.

Its cost is Δ2/4\Delta^2/4 against the undithered Δ2/12\Delta^2/12: three times the noise, or 10log⁡103=4.7710\log_{10}3=4.77 dB. Compare rectangular dither’s 3.01 dB.

There is one way to avoid most of that cost. In subtractive dither (Roberts’ method, in MIT 6.003 lecture 22) you add known noise, round, and subtract the same noise afterwards. The error is then Δ2/12\Delta^2/12 and does not depend on the input. But both ends must share the noise, so most audio uses triangular dither instead. When a 24-bit master is reduced to 16 bits, triangular dither is the usual choice.

Worked example

  1. No dither. Take an input 0.30.3 steps above a level. It always rounds down, so the error is −0.3-0.3 steps every time, and the mean square is 0.32=0.090.3^2=0.09 in units of Δ2\Delta^2.
  2. Rectangular dither. The output is the upper level in a share 0.30.3 of the runs, so the average output is 0.30.3 steps above the lower level, the input. The mean square of the error is p(1−p)=0.3×0.7=0.21p(1-p)=0.3\times0.7=0.21.
  3. Triangular dither. The mean square is 0.250.25 here, as for every input. The rectangular value moves with the input, and the triangular one does not.
  4. Averaged over the input. Undithered, 112=0.0833\tfrac1{12}=0.0833. Rectangular, 16=0.1667\tfrac16=0.1667, a factor 2, so +3.01+3.01 dB. Triangular, 14=0.25\tfrac14=0.25, a factor 3, so +4.77+4.77 dB.

Where you’ll meet this

Every time audio is reduced to fewer bits, dither is applied first, usually triangular. The same idea prints photographs with a few ink colours, and it keeps a slowly fading screen gradient from showing bands.

Dither makes the error steady, but it does not make it smaller. The next page, Oversampling and noise shaping (11.3), shows how to move that hiss to frequencies where it matters less.

Reference card

QuantityFormulaNotes
Undithered outputQ{x[n]}Q\{x[n]\}error repeats for a quiet or simple signal
Dithered outputQ{x[n]+v[n]}Q\{x[n]+v[n]\}v[n]v[n] is a fresh random value for each sample
Rectangular dithervv spread evenly over ±12\pm\tfrac12 stepthe dither’s own mean square is Δ2/12\Delta^2/12
Triangular dithersum of two rectangular values; ±1\pm1 stepthe dither’s own mean square is Δ2/6\Delta^2/6
Round-up sharepp for an input pp steps above a levelaverage output equals the input
Error, rectangulare2‾=p(1−p) Δ2\overline{e^2}=p(1-p)\,\Delta^2zero on a level; Δ2/4\Delta^2/4 halfway; mean Δ2/6\Delta^2/6 over pp
Error, triangulare2‾=Δ2/4\overline{e^2}=\Delta^2/4the same for every pp: no noise modulation
Noise cost, rectangular1/61/12=2\dfrac{1/6}{1/12}=2+3.01+3.01 dB
Noise cost, triangular1/41/12=3\dfrac{1/4}{1/12}=3+4.77+4.77 dB
Subtractive ditheradd vv, round, subtract vverror Δ2/12\Delta^2/12; both ends need the same noise

End of lesson 11.2

Where to go next.

Phasorium
LibraryEvery lesson, in order

Parts

About Phasorium
Look