A loud sound hides quieter sounds close to it in frequency, so a coder need not send them. Watch the quieter tone below slide down beside a loud one: once it falls under the shaded curve, it is masked.
A loud tone hides its neighbours
A 1 kHz masker at 70 dB SPL and a 1200 Hz probe; the model of the text (levels are model levels; your volume sets the real ones).
The probe at 60 dB stands 18.7 dB above the masked threshold at 1200 Hz: audible.
Describe this picture
A 1 kHz masker at 70 dB SPL and a 1200 Hz probe, in the model of the text; levels are model levels, and your volume sets the real ones. One panel, level from −10 to 80 dB SPL against frequency on a logarithmic scale from 100 to 8000 Hz. The threshold in quiet is a dashed curve; the masked threshold is a solid curve with the area under it shaded. The masker is a stem with a square head, and the probe a stem with a round head, filled while the probe is above the masked threshold and open when it is under. The readouts are the probe level and the threshold there, in dB SPL, and the margin in dB, all with one decimal. The 13 s clip opens with the probe at 60 dB SPL, 18.7 dB above the masked threshold at 1200 Hz, 41.3 dB SPL: audible. At 3.5 s it slides down to 45 dB SPL, only 3.7 dB above: still audible, just. At 8 s it slides to 35 dB SPL, 6.3 dB under the threshold: masked, and a coder need not send it. After the clip, the slider “Probe level” sets the probe anywhere from 0 to 70 dB SPL. The “Hear it” button plays the masker and the probe together for 1 s, then the probe alone for 1 s, with their amplitudes in the model’s ratio.
A loud tone hides its neighbours
Sixteen-bit sound at 16 kHz takes 256 kbit/s. The coder on this page sends the same sound in about 38 kbit/s, by using one fact about hearing.
A loud sound hides quieter sounds close to it in frequency. A perceptual coder uses this. It spends bits only on what a listener can hear, and none on what is hidden. MP3 and AAC work this way.
Every signal on this page is synthetic: tones and noise levels set by formula, with no recordings. The levels are model levels. When you press a button to hear a tone, your own volume sets how loud it really is.
Levels in dB SPL
A sound is a small, fast change in air pressure. Its level is given in dB SPL, decibels of sound pressure level, against a reference pressure of 20 µPa:
Pressure is an amplitude, so the factor is 20, as in “Signal and noise, in decibels” of How big is a signal (1.3). A tone’s level in dB SPL is written .
Critical bands and the Bark scale
The inner ear sorts sound by frequency, a little like a bank of band-pass filters. Each of these filters is called a critical band. Their widths are not equal in hertz: about 100 Hz wide at low frequencies, and wider and wider above about 500 Hz.
So hearing needs its own frequency scale, the Bark scale, on which every critical band is 1 Bark wide. I use Zwicker’s formula for the critical-band number:
with in hertz. I write with its subscript, because a bare is the z-transform variable. From 0 to 20 kHz the scale runs from 0 to 24.6 Bark, so hearing has about 25 critical bands.
At 1 kHz:
The threshold in quiet
The quietest tone a listener can hear depends on its frequency. That level is the threshold in quiet, . I use Terhardt’s fit, in dB SPL:
It is 23.0 dB SPL at 100 Hz and 3.4 dB SPL at 1 kHz. Its lowest point is −5.0 dB SPL near 3.3 kHz, where hearing is most sensitive. Above that it climbs again.
Masking
Now play a loud tone, the masker. It raises the threshold of hearing around its own frequency. A quieter tone nearby that falls under the raised threshold is not heard. This is masking.
How far the threshold rises follows a spreading function, measured in Bark from the masker. Let be the distance in Bark from the masker’s frequency . To keep the formula short, I write . Then Schroeder’s spreading function is
in dB. here has no subscript; it is not a power spectral density. At the masker, dB.
A masker of level at then sets a threshold at frequency of
The last term is Johnston’s offset for a tone masker. It puts the threshold’s peak under the masker’s level by 14.5 dB plus the masker’s Bark number. The masked threshold combines this with the threshold in quiet. I turn each of the two levels into a power, add the powers, and go back to dB. Where one is far above the other, the sum is that one to a fraction of a dB.
For a 1 kHz masker at 70 dB SPL, the offset is dB. So the threshold peaks at dB SPL.
The curve is lopsided. Between 2 and 3 Bark from the masker it falls 23.1 dB per Bark going down in frequency, but only 9.1 dB per Bark going up. So, in this model, a masker hides more above its frequency than below it.
For the 1 kHz masker, the masked threshold is 43.6 dB SPL at 900 Hz and 45.1 dB SPL at 1100 Hz. Further up it is 41.3 at 1200 Hz, 28.5 at 1500 Hz and 10.8 at 2000 Hz. The masker’s threshold and the threshold in quiet are equal at 637 Hz and at 2467 Hz. Outside those two frequencies the threshold in quiet is the larger, and the masked threshold soon follows it.
This is one model among several. Real thresholds vary from listener to listener, and real coders use more careful models. But every number here can be recomputed from the four formulas, which is why I use it.
The masker and a probe
The picture at the top of the page puts a 1 kHz masker at 70 dB SPL beside a quieter tone at 1200 Hz, the probe. Watch the probe’s margin, its level minus the masked threshold at 1200 Hz, 41.3 dB SPL. At 60 dB SPL the probe stands 18.7 dB above it and at 45 dB only 3.7 dB; at 35 dB it is 6.3 dB under, and masked.
Press “Hear it” to listen for yourself. Whether you hear the probe vanish depends on your ears, your headphones and your volume. The page’s verdict, audible or masked, is the model’s.
What a coder takes from this
Below 41.3 dB SPL, by this model, the probe could be left out of the sound altogether. A coder can do better than drop whole tones, though. Any error it makes, such as rounding noise, is also hidden as long as it stays under the masked threshold.
This is the hearing version of “A weak tone beside a strong one” in Windowing & spectral leakage (15.1). There, a strong tone’s leakage could bury a weak tone in the spectrum. Here, the ear does the burying.
Bits only where the sound shows
A coder works on short frames of sound. For each frame it needs the levels in each band, a masked threshold for each band, and a rule that turns the two into bits.
One frame, 21 bands
I take one frame of 32 ms at a sampling rate of 16 kHz: 512 samples. A transform called the MDCT, which I come back to after the instrument, turns it into 256 coefficients. Coefficient stands for the frequency Hz.
Frames overlap by half, so each frame brings 256 new samples, 16 ms of sound. Sent as plain 16-bit samples, those would take bits.
I group the coefficients into one-Bark bands. Band 0 runs from 0 to 1 Bark, band 1 from 1 to 2 Bark, and so on up to band 20, which ends at 21 Bark, 7617 Hz. A coefficient belongs to the band its frequency falls in.
The low bands are narrow in hertz, so they hold few coefficients: 3 in band 0, 5 in band 8. The high ones are wide: 39 in band 20. The 12 coefficients above 7617 Hz fall in no band, and get no bits here.
The sound in the frame
The frame holds two tones: 1 kHz at 70 dB SPL and 3.2 kHz at 52 dB SPL. Under them is a quiet noise floor, 20.5 dB SPL in every band.
The 1 kHz tone lies in band 8, from 922 to 1079 Hz. The 3.2 kHz tone lies in band 15, from 2711 to 3212 Hz, near its top edge. To add a tone to the noise in its band I add their powers, then go back to dB. Both tone bands come out at the tone’s level to one decimal, 70.0 and 52.0 dB SPL.
A mask for each band
Each band’s mask is computed at one frequency, its centre in hertz: halfway between its two edges. Three thresholds count there: the threshold in quiet, the 1 kHz tone’s spread threshold, and the 3.2 kHz tone’s.
I combine the three as in the first instrument: add their powers, then go back to dB.
The masks run from 2.8 dB SPL, in band 19, to 47.0 dB SPL, in band 8 around the loud tone.
Noise under the mask
How far a band’s level stands above its mask is its signal-to-mask ratio, in dB, now with both levels the band’s.
“Each bit is worth 6 dB” in Quantization & noise (11.1) showed that each bit pushes rounding noise down by 6.02 dB. If I round every coefficient in the band with bits, its noise sits about dB under the band’s level.
The noise is hidden when is at least the signal-to-mask ratio. So the fewest bits that do the job are
where rounds up to a whole number. A band under its mask gets : nothing is sent for it. Choosing band by band like this is bit allocation.
It is the move of “Smooth blocks need few numbers” in Image compression (28.4): round transform coefficients as coarsely as the eye or ear allows.
The bit chart
Bits only where the sound shows
One 32 ms frame: tones at 1 kHz (70 dB SPL) and 3.2 kHz (52 dB SPL) over a 20.5 dB noise floor in every band; 21 one-Bark bands.
Band levels against the masked threshold: both tones stand far above it, 14 noise bands by less, and 5 bands sit under it.
Describe this picture
One 32 ms frame: tones at 1 kHz (70 dB SPL) and 3.2 kHz (52 dB SPL) over a 20.5 dB noise floor in every band, in 21 one-Bark bands. Two stacked panels share the band axis, 0 to 21 Bark. The first, the levels, runs from −10 to 80 dB SPL: each band’s level is an outlined bar and its mask a filled bar. The second shows the bits per coefficient of each band as bars, from 0 to 8. The readouts are the bands with bits, the frame’s bits and the bits of a 16-bit frame. The 12 s clip opens on the levels and the masks, with no bits yet: both tones stand far above the masked threshold, 14 noise bands by less, and 5 bands sit under it. From 3 s the bits grow band by band. At 8 s it holds on the 1 kHz band, 23.0 dB above its mask, so 3.82 rounded up: 4 bits per coefficient. At the end 16 of 21 bands need bits, and the frame needs 605 bits against 4096 at 16 bits each, 6.8 times fewer.
Watch the bits grow band by band. Both tones get bits, and so do most bands of the quiet noise; the five bands under their masks get none.
Reading the chart
The 3.2 kHz band gets the most bits. Its level, 52.0 dB SPL, stands 32.0 dB over its mask of 20.0 dB SPL. That needs 6 bits per coefficient, since 5 bits cover only 30.1 dB.
The loud 1 kHz tone gets only 4. It is louder, but it also raises its own mask more, to 47.0 dB SPL.
From band 0 to band 20, the bits per coefficient are 0, 1, 2, 3, 3, 3, 1, 0, 4, 0, 0, 0, 1, 3, 3, 6, 1, 2, 3, 3, 3. Bands 0, 7, 9, 10 and 11 are under their masks and cost nothing. Bands 7, 9, 10 and 11 sit near the loud tone; band 0 is under the high threshold in quiet at low frequencies.
The total is the sum over the bands of coefficients times bits: 605 bits, against 4096 for plain 16-bit samples, 6.8 times fewer. Most of it does not go to the tones. The two tone bands take 116 bits; the four highest bands of noise floor take 342, because they hold the most coefficients.
The 605 bits count only the coefficients. A real coder also sends, for every band, its bits per coefficient and a scale, and then shrinks everything with Huffman codes. So this is the allocation step, not a full coder.
The maths behind it · reverse water-filling
Bit allocation is reverse water-filling from rate–distortion theory. Spend bits on the components whose variance stands above the allowed distortion, and none on the rest. Here the allowed distortion in each band is its mask.
The MDCT behind the bands
The MDCT, the modified discrete cosine transform, appeared in “One prototype, eight channels, a perfect sum” of Filter banks (23.1). Frames of samples hop by , and each frame gives coefficients. In 23.1 that count was ; here it is , the channel count of a critically sampled bank. On this page .
For one frame , ,
for , with the sine window
The halves in and put the cosines’ mirror points between samples, as in “What the mirror does to the DFT” of The discrete cosine transform (13.6).
To go back, each frame computes
for , and the frames are overlap-added, samples apart.
numbers cannot hold samples. So one frame’s is its windowed samples plus a mirrored copy of them, an alias in time. In the first half of the frame the copy is subtracted; in the second half it is added.
Where two frames overlap, the earlier frame’s second half adds its copy, and the later frame’s first half subtracts the same copy. They cancel: time-domain alias cancellation. The sine window also passes the samples through whole, since .
So overlap-add gives the signal back exactly. On 192 ms of the page’s two tones, with , the largest error is below (rounding).
The maths behind it · orthogonal transforms
The MDCT is not square: samples in, numbers out. Yet consecutive frames together form an orthogonal transform of the whole signal. Each frame’s aliasing lies in a subspace that the next frame’s cancels.
MP3 puts an MDCT after a 32-band filter bank. AAC uses the MDCT alone, with 1024 coefficients per frame, or 128 for short frames.
I left out a good deal. Masking also works in time, just before and after a loud sound. A sudden attack can smear quantisation noise in front of it, called pre-echo, which is why coders switch to short frames there. I also left out stereo coding and the Huffman tables.
Worked example
1. Hertz to Bark. At 1000 Hz, and . Then Bark.
2. The threshold at the probe. At 1200 Hz, , so . The spreading function gives dB, with . The masker’s threshold is dB SPL. The threshold in quiet there is 2.7 dB SPL, 38.6 dB lower, so the power sum is still 41.3 dB SPL.
3. Bits for a band. The 3.2 kHz band stands 32.0 dB over its mask. Then , which rounds up to 6 bits per coefficient. Its 16 coefficients take bits.
4. The saving. The frame needs 605 bits for 256 new samples. Plain 16-bit samples need , and , so 6.8 times fewer.
Where you’ll meet this
MP3 and AAC, in music files, streaming and broadcast, both shape their noise with a masking model and an MDCT, as on this page. Dolby Digital, in film and television sound, uses an MDCT too. So does Opus, the codec of many voice and video calls, in its music mode.
Bluetooth headphones code sound too. Their basic codec, SBC, splits it into 4 or 8 bands with a filter bank and allots bits band by band. Many headphones can also receive AAC.
Speech codecs work differently: they model the voice that made the sound. The source–filter model of speech (30.1) takes that up.
For more, see Zwicker and Fastl, Psychoacoustics (3rd ed., 2007), chapters 4 and 6, J. D. Johnston, “Transform coding of audio signals using perceptual noise criteria” (IEEE J. Sel. Areas Commun., 1988), and Bosi and Goldberg, Introduction to Digital Audio Coding and Standards (2003), chapters 5 to 7.
Reference card
| Quantity | Formula | Notes |
|---|---|---|
| Bark | Zwicker; 1 kHz is 8.51 | |
| Threshold in quiet | , Terhardt’s fit | lowest −5.0 dB SPL near 3.3 kHz |
| Spreading | , Schroeder’s | falls faster below the masker |
| Tone’s threshold | one tonal masker, Johnston | |
| Masked threshold | : power sum of and each masker’s threshold | add powers, back to dB |
| Bits | per coefficient | 0 under the mask |
| MDCT | samples in, out, sine window | overlap-add is exact |