Take a smooth gradient from dark to light, 256 pixels wide, and store it with only 8 shades. Watch hard bands appear where the shade jumps, and then watch them dissolve into fine grain.
Bands in a smooth sky
A gradient from dark to light, 256 pixels wide.
A smooth sky, from dark to light.
Describe this picture
The picture has no controls and plays once. It shows a smooth gradient, then the same gradient stored with 8 shades, in hard bands, then the 8-shade gradient with dither, where the bands become fine grain. Under it a strip plots brightness against position. The readout “shades available” first shows “any”, then “8”.
Bands in a smooth sky
Quantization and noise (11.1) ended on a quiet sine. The error of a quiet or simple signal repeats with the signal, so it does not sound like hiss. It sounds like extra harmonics, a distortion that follows the music. This page is about a cure that sounds strange at first: add a little noise before you round.
I started with a picture, because a picture shows the problem fastest. In What a signal is (1.1) you saw that one row of a picture is a signal: brightness at each pixel position. Store a smooth gradient with only 8 shades and the rounding error is large, and it is not random. It is the same small sawtooth each time the brightness crosses a level, so your eye sees it as banding: hard edges where the shade jumps.
In the last step of the picture at the top I add dither: a small random value, spread evenly over plus or minus half a shade, added to each pixel before rounding. The picture still has only 8 shades, but the bands dissolve into fine grain.
Look at the strip of brightness against position under the picture at the top. With dither, the rounded values flicker between neighbouring shades instead of switching once. That flicker is what your eye averages into a smooth gradient. A newspaper photo works the same way: each dot is black or white, and the grey you see is the average over a small area.
Right on average
The picture suggests an idea about averages, and I want you to see it with numbers. Dither adds noise, so you would expect the output to get worse. Before you run the next instrument, decide: I will add noise to a quiet signal before rounding it. Will the result get better or worse?
Here is the setup. The input is a sine 1.4 steps tall, with 32 samples per period. A step is the distance between neighbouring levels, and the levels sit at whole steps: Without dither, the quantizer rounds each sample to the nearest level. Every period gives the same output, a fixed staircase.
With dither, we round instead, where is a fresh random value for each sample, spread evenly between and step. This is rectangular dither. The question is what one run does to a sample, and what many runs do on average.
Right on average
A sine 1.4 steps tall, 32 samples per period, rounded to whole steps.
One run each. Without dither: a fixed staircase. With dither: a noisier staircase.
largest gap from the input (steps)
Describe this picture
Two stacked panels, “without dither” and “with dither”, each drawing the input as a dashed line labelled “input” and the average of the runs as stems. The picture plays by itself, and the readout “runs averaged” climbs from 1 to 100. Under the heading “largest gap from the input (steps)”, the readouts “without dither” and “with dither” give each panel’s largest gap. Two buttons, “Hear without dither” and “Hear with dither”, play sound and do not change the picture.
With one run, the gaps read 0.464 without dither and 0.778 with it, so dither is worse for a single run. Watch the stems as the count rises. At 10 runs the gap with dither has fallen to 0.273, and at 100 runs to 0.102. Without dither the stems never move: every run is the same, so the gap stays at 0.464.
Here is why. Suppose the input sits a share of the way from one level to the next, so it is steps above the lower level, with . The sum spreads over step around the input. The output rounds up when , a stretch of values of length out of a total length 1, so it rounds up in a share of the runs. So the output is the upper level in a share of the runs and the lower level in the rest.
Average those outputs and you get steps above the lower level, which is the input. The average of many runs closes in on the input. That is what “right on average” means.
Press Hear without dither and Hear with dither. Each plays that panel’s output for a single run, for 2 seconds: a 1.5 kHz tone at 48 kHz, with a step of 0.05 of full scale. The second has a steady hiss. I think of it this way: the error without dither repeats every period, so it is made of harmonics (see Signals as sums of sinusoids, 7.1), and the third harmonic of the error is 0.273 of a step. The error with dither never repeats, so it is hiss, which the ear tolerates far better than a tone that follows the music.
Dither has a cost. The rounding error’s mean square is , as in 11.1. The dither adds its own mean square, which for a value spread evenly over step is also . Averaged over the input position, the total is : twice as much, which is dB more noise.
Rectangular or triangular dither
Rectangular dither fixes the average, but one problem remains. How much hiss you get still depends on where the input sits between two levels. Put the input exactly on a level and the dither never pushes it across the halfway point, so the output is exact and there is no error. Put it halfway between two levels and the error is as large as it gets.
Here is the mean square of the error for an input steps above a level. The output is up by steps from the input with probability , and down by steps with probability , so
It is zero at and at , and largest at , where it is .
When the input is a slow fade, the hiss swells and shrinks with it. That is noise modulation, hiss that breathes with the music. A fading piano note is the classic case: the background hiss pulses along with the note.
The fix is triangular dither: the sum of two independent rectangular values. It spreads over step and is most often near 0. For this shape the mean square of the error is for every input. Watch both kinds of dither as a steady input moves from one level to the next.
Rectangular or triangular dither
A steady input moving from one level to the next. Each value is averaged over all dither values.
Input exactly on a level. Rectangular dither leaves no error at all here; triangular dither adds its hiss anyway.
Describe this picture
The mean square of the error, divided by , from 0 to 0.3, against the input level in steps above a level, from 0 to 1; the picture has no controls. Each value is averaged over all dither values, so these are exact expectations, not random draws. A solid curve is labelled “rectangular” and a dashed line “triangular”; a dot on the curve and a diamond on the line follow the current input. The readouts “input level”, “rectangular” and “triangular” start at 0.00, 0.000 and 0.250, both read 0.250 at 0.50, and at 1.00 the rectangular one is back at 0.000 while the triangular one still reads 0.250.
The rectangular curve is an arch from 0 up to 0.25 and back, and the triangular line stays flat at 0.25. A steady hiss does not breathe, and that is the reason triangular dither is the usual choice.
Its cost is against the undithered : three times the noise, or dB. Compare rectangular dither’s 3.01 dB.
There is one way to avoid most of that cost. In subtractive dither (Roberts’ method, in MIT 6.003 lecture 22) you add known noise, round, and subtract the same noise afterwards. The error is then and does not depend on the input. But both ends must share the noise, so most audio uses triangular dither instead. When a 24-bit master is reduced to 16 bits, triangular dither is the usual choice.
Worked example
- No dither. Take an input steps above a level. It always rounds down, so the error is steps every time, and the mean square is in units of .
- Rectangular dither. The output is the upper level in a share of the runs, so the average output is steps above the lower level, the input. The mean square of the error is .
- Triangular dither. The mean square is here, as for every input. The rectangular value moves with the input, and the triangular one does not.
- Averaged over the input. Undithered, . Rectangular, , a factor 2, so dB. Triangular, , a factor 3, so dB.
Where you’ll meet this
Every time audio is reduced to fewer bits, dither is applied first, usually triangular. The same idea prints photographs with a few ink colours, and it keeps a slowly fading screen gradient from showing bands.
Dither makes the error steady, but it does not make it smaller. The next page, Oversampling and noise shaping (11.3), shows how to move that hiss to frequencies where it matters less.
Reference card
| Quantity | Formula | Notes |
|---|---|---|
| Undithered output | error repeats for a quiet or simple signal | |
| Dithered output | is a fresh random value for each sample | |
| Rectangular dither | spread evenly over step | the dither’s own mean square is |
| Triangular dither | sum of two rectangular values; step | the dither’s own mean square is |
| Round-up share | for an input steps above a level | average output equals the input |
| Error, rectangular | zero on a level; halfway; mean over | |
| Error, triangular | the same for every : no noise modulation | |
| Noise cost, rectangular | dB | |
| Noise cost, triangular | dB | |
| Subtractive dither | add , round, subtract | error ; both ends need the same noise |