Skip to content

Perceptual audio coding

A loud tone hides quieter ones nearby; MP3 and AAC give each MDCT band just enough bits to keep its noise hidden.

Before this23.1 · 28.4 · 4 more
Chapter 29 · Lesson 3 of 3

First, the picture

A loud sound hides quieter sounds close to it in frequency, so a coder need not send them. Watch the quieter tone below slide down beside a loud one: once it falls under the shaded curve, it is masked.

A loud tone hides its neighbours

A 1 kHz masker at 70 dB SPL and a 1200 Hz probe; the model of the text (levels are model levels; your volume sets the real ones).

The probe at 60 dB stands 18.7 dB above the masked threshold at 1200 Hz: audible.

probe level
60.0 dB SPL
threshold there
41.3 dB SPL
margin
18.7 dB
0.00 / 13.00 s
Describe this picture

A 1 kHz masker at 70 dB SPL and a 1200 Hz probe, in the model of the text; levels are model levels, and your volume sets the real ones. One panel, level from −10 to 80 dB SPL against frequency on a logarithmic scale from 100 to 8000 Hz. The threshold in quiet is a dashed curve; the masked threshold is a solid curve with the area under it shaded. The masker is a stem with a square head, and the probe a stem with a round head, filled while the probe is above the masked threshold and open when it is under. The readouts are the probe level and the threshold there, in dB SPL, and the margin in dB, all with one decimal. The 13 s clip opens with the probe at 60 dB SPL, 18.7 dB above the masked threshold at 1200 Hz, 41.3 dB SPL: audible. At 3.5 s it slides down to 45 dB SPL, only 3.7 dB above: still audible, just. At 8 s it slides to 35 dB SPL, 6.3 dB under the threshold: masked, and a coder need not send it. After the clip, the slider “Probe level” sets the probe anywhere from 0 to 70 dB SPL. The “Hear it” button plays the masker and the probe together for 1 s, then the probe alone for 1 s, with their amplitudes in the model’s ratio.

A loud tone hides its neighbours

Sixteen-bit sound at 16 kHz takes 256 kbit/s. The coder on this page sends the same sound in about 38 kbit/s, by using one fact about hearing.

A loud sound hides quieter sounds close to it in frequency. A perceptual coder uses this. It spends bits only on what a listener can hear, and none on what is hidden. MP3 and AAC work this way.

Every signal on this page is synthetic: tones and noise levels set by formula, with no recordings. The levels are model levels. When you press a button to hear a tone, your own volume sets how loud it really is.

Levels in dB SPL

A sound is a small, fast change in air pressure. Its level is given in dB SPL, decibels of sound pressure level, against a reference pressure of 20 µPa:

L=20log⁡10prms20 μPa dB SPL.L=20\log_{10}\frac{p_\text{rms}}{20\,\mu\text{Pa}}\ \text{dB SPL}.

Pressure is an amplitude, so the factor is 20, as in “Signal and noise, in decibels” of How big is a signal (1.3). A tone’s level in dB SPL is written LSPLL_\text{SPL}.

Critical bands and the Bark scale

The inner ear sorts sound by frequency, a little like a bank of band-pass filters. Each of these filters is called a critical band. Their widths are not equal in hertz: about 100 Hz wide at low frequencies, and wider and wider above about 500 Hz.

So hearing needs its own frequency scale, the Bark scale, on which every critical band is 1 Bark wide. I use Zwicker’s formula for the critical-band number:

zBark(f)=13arctan⁡(0.00076 f)+3.5arctan⁡ ⁣((f/7500)2),\begin{aligned} z_\text{Bark}(f)&=13\arctan(0.00076\,f)\\ &\quad+3.5\arctan\!\big((f/7500)^2\big), \end{aligned}

with ff in hertz. I write zBarkz_\text{Bark} with its subscript, because a bare zz is the z-transform variable. From 0 to 20 kHz the scale runs from 0 to 24.6 Bark, so hearing has about 25 critical bands.

At 1 kHz:

zBark(1000)=13arctan⁡(0.76)+3.5arctan⁡(0.0178)=8.448+0.062=8.51.\begin{aligned} z_\text{Bark}(1000)&=13\arctan(0.76)\\ &\quad+3.5\arctan(0.0178)\\ &=8.448+0.062=8.51. \end{aligned}

The threshold in quiet

The quietest tone a listener can hear depends on its frequency. That level is the threshold in quiet, LATH(f)L_\text{ATH}(f). I use Terhardt’s fit, in dB SPL:

LATH(f)=3.64 (f/kHz)−0.8−6.5 e−0.6 (f/kHz−3.3)2+10−3 (f/kHz)4.\begin{aligned} L_\text{ATH}(f)&=3.64\,(f/\text{kHz})^{-0.8}\\ &\quad-6.5\,e^{-0.6\,(f/\text{kHz}-3.3)^2}\\ &\quad+10^{-3}\,(f/\text{kHz})^4. \end{aligned}

It is 23.0 dB SPL at 100 Hz and 3.4 dB SPL at 1 kHz. Its lowest point is −5.0 dB SPL near 3.3 kHz, where hearing is most sensitive. Above that it climbs again.

Masking

Now play a loud tone, the masker. It raises the threshold of hearing around its own frequency. A quieter tone nearby that falls under the raised threshold is not heard. This is masking.

How far the threshold rises follows a spreading function, measured in Bark from the masker. Let ΔzBark=zBark(f)−zBark(fm)\Delta z_\text{Bark}=z_\text{Bark}(f)-z_\text{Bark}(f_\text{m}) be the distance in Bark from the masker’s frequency fmf_\text{m}. To keep the formula short, I write a=ΔzBark+0.474a=\Delta z_\text{Bark}+0.474. Then Schroeder’s spreading function is

S(ΔzBark)=15.81+7.5 a−17.51+a2\begin{aligned} S(\Delta z_\text{Bark})&=15.81+7.5\,a\\ &\quad-17.5\sqrt{1+a^2} \end{aligned}

in dB. SS here has no subscript; it is not a power spectral density. At the masker, S(0)=0.00S(0)=0.00 dB.

A masker of level LSPLL_\text{SPL} at fmf_\text{m} then sets a threshold at frequency ff of

LSPL+S(ΔzBark)−(14.5+zBark(fm)) dB SPL.\begin{aligned} &L_\text{SPL}+S(\Delta z_\text{Bark})\\ &\quad-\big(14.5+z_\text{Bark}(f_\text{m})\big)\ \text{dB SPL}. \end{aligned}

The last term is Johnston’s offset for a tone masker. It puts the threshold’s peak under the masker’s level by 14.5 dB plus the masker’s Bark number. The masked threshold Lmask(f)L_\text{mask}(f) combines this with the threshold in quiet. I turn each of the two levels into a power, add the powers, and go back to dB. Where one is far above the other, the sum is that one to a fraction of a dB.

For a 1 kHz masker at 70 dB SPL, the offset is 14.5+8.51=23.014.5+8.51=23.0 dB. So the threshold peaks at 70−23.0=47.070-23.0=47.0 dB SPL.

The curve is lopsided. Between 2 and 3 Bark from the masker it falls 23.1 dB per Bark going down in frequency, but only 9.1 dB per Bark going up. So, in this model, a masker hides more above its frequency than below it.

For the 1 kHz masker, the masked threshold is 43.6 dB SPL at 900 Hz and 45.1 dB SPL at 1100 Hz. Further up it is 41.3 at 1200 Hz, 28.5 at 1500 Hz and 10.8 at 2000 Hz. The masker’s threshold and the threshold in quiet are equal at 637 Hz and at 2467 Hz. Outside those two frequencies the threshold in quiet is the larger, and the masked threshold soon follows it.

This is one model among several. Real thresholds vary from listener to listener, and real coders use more careful models. But every number here can be recomputed from the four formulas, which is why I use it.

The masker and a probe

The picture at the top of the page puts a 1 kHz masker at 70 dB SPL beside a quieter tone at 1200 Hz, the probe. Watch the probe’s margin, its level minus the masked threshold at 1200 Hz, 41.3 dB SPL. At 60 dB SPL the probe stands 18.7 dB above it and at 45 dB only 3.7 dB; at 35 dB it is 6.3 dB under, and masked.

Press “Hear it” to listen for yourself. Whether you hear the probe vanish depends on your ears, your headphones and your volume. The page’s verdict, audible or masked, is the model’s.

What a coder takes from this

Below 41.3 dB SPL, by this model, the probe could be left out of the sound altogether. A coder can do better than drop whole tones, though. Any error it makes, such as rounding noise, is also hidden as long as it stays under the masked threshold.

This is the hearing version of “A weak tone beside a strong one” in Windowing & spectral leakage (15.1). There, a strong tone’s leakage could bury a weak tone in the spectrum. Here, the ear does the burying.

Bits only where the sound shows

A coder works on short frames of sound. For each frame it needs the levels in each band, a masked threshold for each band, and a rule that turns the two into bits.

One frame, 21 bands

I take one frame of 32 ms at a sampling rate of 16 kHz: 512 samples. A transform called the MDCT, which I come back to after the instrument, turns it into 256 coefficients. Coefficient kk stands for the frequency (k+12)×31.25(k+\tfrac12)\times31.25 Hz.

Frames overlap by half, so each frame brings 256 new samples, 16 ms of sound. Sent as plain 16-bit samples, those would take 16×256=409616\times256=4096 bits.

I group the coefficients into one-Bark bands. Band 0 runs from 0 to 1 Bark, band 1 from 1 to 2 Bark, and so on up to band 20, which ends at 21 Bark, 7617 Hz. A coefficient belongs to the band its frequency falls in.

The low bands are narrow in hertz, so they hold few coefficients: 3 in band 0, 5 in band 8. The high ones are wide: 39 in band 20. The 12 coefficients above 7617 Hz fall in no band, and get no bits here.

The sound in the frame

The frame holds two tones: 1 kHz at 70 dB SPL and 3.2 kHz at 52 dB SPL. Under them is a quiet noise floor, 20.5 dB SPL in every band.

The 1 kHz tone lies in band 8, from 922 to 1079 Hz. The 3.2 kHz tone lies in band 15, from 2711 to 3212 Hz, near its top edge. To add a tone to the noise in its band I add their powers, then go back to dB. Both tone bands come out at the tone’s level to one decimal, 70.0 and 52.0 dB SPL.

A mask for each band

Each band’s mask is computed at one frequency, its centre in hertz: halfway between its two edges. Three thresholds count there: the threshold in quiet, the 1 kHz tone’s spread threshold, and the 3.2 kHz tone’s.

I combine the three as in the first instrument: add their powers, then go back to dB.

The masks run from 2.8 dB SPL, in band 19, to 47.0 dB SPL, in band 8 around the loud tone.

Noise under the mask

How far a band’s level stands above its mask is its signal-to-mask ratio, LSPL−LmaskL_\text{SPL}-L_\text{mask} in dB, now with both levels the band’s.

“Each bit is worth 6 dB” in Quantization & noise (11.1) showed that each bit pushes rounding noise down by 6.02 dB. If I round every coefficient in the band with BB bits, its noise sits about 6.02 B6.02\,B dB under the band’s level.

The noise is hidden when 6.02 B6.02\,B is at least the signal-to-mask ratio. So the fewest bits that do the job are

B=⌈LSPL−Lmask6.02⌉,B=\left\lceil\frac{L_\text{SPL}-L_\text{mask}}{6.02}\right\rceil,

where ⌈⋅⌉\lceil\cdot\rceil rounds up to a whole number. A band under its mask gets B=0B=0: nothing is sent for it. Choosing BB band by band like this is bit allocation.

It is the move of “Smooth blocks need few numbers” in Image compression (28.4): round transform coefficients as coarsely as the eye or ear allows.

The bit chart

Bits only where the sound shows

One 32 ms frame: tones at 1 kHz (70 dB SPL) and 3.2 kHz (52 dB SPL) over a 20.5 dB noise floor in every band; 21 one-Bark bands.

Band levels against the masked threshold: both tones stand far above it, 14 noise bands by less, and 5 bands sit under it.

bands with bits
—
frame bits
—
16-bit frame
—
0.00 / 12.00 s
Describe this picture

One 32 ms frame: tones at 1 kHz (70 dB SPL) and 3.2 kHz (52 dB SPL) over a 20.5 dB noise floor in every band, in 21 one-Bark bands. Two stacked panels share the band axis, 0 to 21 Bark. The first, the levels, runs from −10 to 80 dB SPL: each band’s level is an outlined bar and its mask a filled bar. The second shows the bits per coefficient of each band as bars, from 0 to 8. The readouts are the bands with bits, the frame’s bits and the bits of a 16-bit frame. The 12 s clip opens on the levels and the masks, with no bits yet: both tones stand far above the masked threshold, 14 noise bands by less, and 5 bands sit under it. From 3 s the bits grow band by band. At 8 s it holds on the 1 kHz band, 23.0 dB above its mask, so 3.82 rounded up: 4 bits per coefficient. At the end 16 of 21 bands need bits, and the frame needs 605 bits against 4096 at 16 bits each, 6.8 times fewer.

Watch the bits grow band by band. Both tones get bits, and so do most bands of the quiet noise; the five bands under their masks get none.

Reading the chart

The 3.2 kHz band gets the most bits. Its level, 52.0 dB SPL, stands 32.0 dB over its mask of 20.0 dB SPL. That needs 6 bits per coefficient, since 5 bits cover only 30.1 dB.

The loud 1 kHz tone gets only 4. It is louder, but it also raises its own mask more, to 47.0 dB SPL.

From band 0 to band 20, the bits per coefficient are 0, 1, 2, 3, 3, 3, 1, 0, 4, 0, 0, 0, 1, 3, 3, 6, 1, 2, 3, 3, 3. Bands 0, 7, 9, 10 and 11 are under their masks and cost nothing. Bands 7, 9, 10 and 11 sit near the loud tone; band 0 is under the high threshold in quiet at low frequencies.

The total is the sum over the bands of coefficients times bits: 605 bits, against 4096 for plain 16-bit samples, 6.8 times fewer. Most of it does not go to the tones. The two tone bands take 116 bits; the four highest bands of noise floor take 342, because they hold the most coefficients.

The 605 bits count only the coefficients. A real coder also sends, for every band, its bits per coefficient and a scale, and then shrinks everything with Huffman codes. So this is the allocation step, not a full coder.

The maths behind it · reverse water-filling

Bit allocation is reverse water-filling from rate–distortion theory. Spend bits on the components whose variance stands above the allowed distortion, and none on the rest. Here the allowed distortion in each band is its mask.

The MDCT behind the bands

The MDCT, the modified discrete cosine transform, appeared in “One prototype, eight channels, a perfect sum” of Filter banks (23.1). Frames of 2M2M samples hop by MM, and each frame gives MM coefficients. In 23.1 that count was NN; here it is MM, the channel count of a critically sampled bank. On this page M=256M=256.

For one frame x[n]x[n], n=0,…,2M−1n=0,\dots,2M-1,

XMDCT[k]=∑n=02M−1w[n] x[n]×cos⁡ ⁣(πM(n+12+M2)×(k+12)),\begin{aligned} X_\text{MDCT}[k]&=\sum_{n=0}^{2M-1}w[n]\,x[n]\\ &\quad\times\cos\!\Big(\frac{\pi}{M}\big(n+\tfrac12+\tfrac M2\big)\\ &\qquad\quad\times\big(k+\tfrac12\big)\Big), \end{aligned}

for k=0,…,M−1k=0,\dots,M-1, with the sine window

w[n]=sin⁡ ⁣(π (n+12)2M).w[n]=\sin\!\Big(\frac{\pi\,(n+\frac12)}{2M}\Big).

The halves in n+12n+\tfrac12 and k+12k+\tfrac12 put the cosines’ mirror points between samples, as in “What the mirror does to the DFT” of The discrete cosine transform (13.6).

To go back, each frame computes

y[n]=2M w[n]∑k=0M−1XMDCT[k]×cos⁡ ⁣(πM(n+12+M2)×(k+12)),\begin{aligned} y[n]&=\frac{2}{M}\,w[n]\sum_{k=0}^{M-1}X_\text{MDCT}[k]\\ &\quad\times\cos\!\Big(\frac{\pi}{M}\big(n+\tfrac12+\tfrac M2\big)\\ &\qquad\quad\times\big(k+\tfrac12\big)\Big), \end{aligned}

for n=0,…,2M−1n=0,\dots,2M-1, and the frames are overlap-added, MM samples apart.

MM numbers cannot hold 2M2M samples. So one frame’s y[n]y[n] is its windowed samples plus a mirrored copy of them, an alias in time. In the first half of the frame the copy is subtracted; in the second half it is added.

Where two frames overlap, the earlier frame’s second half adds its copy, and the later frame’s first half subtracts the same copy. They cancel: time-domain alias cancellation. The sine window also passes the samples through whole, since w[n]2+w[n+M]2=sin⁡2+cos⁡2=1w[n]^2+w[n+M]^2=\sin^2+\cos^2=1.

So overlap-add gives the signal back exactly. On 192 ms of the page’s two tones, with M=256M=256, the largest error is below 10−1210^{-12} (rounding).

The maths behind it · orthogonal transforms

The MDCT is not square: 2M2M samples in, MM numbers out. Yet consecutive frames together form an orthogonal transform of the whole signal. Each frame’s aliasing lies in a subspace that the next frame’s cancels.

MP3 puts an MDCT after a 32-band filter bank. AAC uses the MDCT alone, with 1024 coefficients per frame, or 128 for short frames.

I left out a good deal. Masking also works in time, just before and after a loud sound. A sudden attack can smear quantisation noise in front of it, called pre-echo, which is why coders switch to short frames there. I also left out stereo coding and the Huffman tables.

Worked example

1. Hertz to Bark. At 1000 Hz, 0.00076×1000=0.760.00076\times1000=0.76 and (1000/7500)2=0.0178(1000/7500)^2=0.0178. Then 13arctan⁡(0.76)+3.5arctan⁡(0.0178)=8.448+0.062=8.5113\arctan(0.76)+3.5\arctan(0.0178)=8.448+0.062=8.51 Bark.

2. The threshold at the probe. At 1200 Hz, zBark=9.70z_\text{Bark}=9.70, so ΔzBark=9.70−8.51=1.19\Delta z_\text{Bark}=9.70-8.51=1.19. The spreading function gives S(1.19)=−5.69S(1.19)=-5.69 dB, with a=1.191+0.474=1.665a=1.191+0.474=1.665. The masker’s threshold is 70−5.69−23.01=41.3070-5.69-23.01=41.30 dB SPL. The threshold in quiet there is 2.7 dB SPL, 38.6 dB lower, so the power sum is still 41.3 dB SPL.

3. Bits for a band. The 3.2 kHz band stands 32.0 dB over its mask. Then 32.0/6.02=5.3232.0/6.02=5.32, which rounds up to 6 bits per coefficient. Its 16 coefficients take 16×6=9616\times6=96 bits.

4. The saving. The frame needs 605 bits for 256 new samples. Plain 16-bit samples need 16×256=409616\times256=4096, and 4096/605=6.774096/605=6.77, so 6.8 times fewer.

Where you’ll meet this

MP3 and AAC, in music files, streaming and broadcast, both shape their noise with a masking model and an MDCT, as on this page. Dolby Digital, in film and television sound, uses an MDCT too. So does Opus, the codec of many voice and video calls, in its music mode.

Bluetooth headphones code sound too. Their basic codec, SBC, splits it into 4 or 8 bands with a filter bank and allots bits band by band. Many headphones can also receive AAC.

Speech codecs work differently: they model the voice that made the sound. The source–filter model of speech (30.1) takes that up.

For more, see Zwicker and Fastl, Psychoacoustics (3rd ed., 2007), chapters 4 and 6, J. D. Johnston, “Transform coding of audio signals using perceptual noise criteria” (IEEE J. Sel. Areas Commun., 1988), and Bosi and Goldberg, Introduction to Digital Audio Coding and Standards (2003), chapters 5 to 7.

Reference card

QuantityFormulaNotes
BarkzBark(f)=13arctan⁡(0.00076f)+3.5arctan⁡((f/7500)2)z_\text{Bark}(f)=13\arctan(0.00076f)+3.5\arctan\big((f/7500)^2\big)Zwicker; 1 kHz is 8.51
Threshold in quietLATH(f)L_\text{ATH}(f), Terhardt’s fitlowest −5.0 dB SPL near 3.3 kHz
SpreadingS(ΔzBark)S(\Delta z_\text{Bark}), Schroeder’sfalls faster below the masker
Tone’s thresholdLSPL+S(ΔzBark)−(14.5+zBark(fm))L_\text{SPL}+S(\Delta z_\text{Bark})-\big(14.5+z_\text{Bark}(f_\text{m})\big)one tonal masker, Johnston
Masked thresholdLmaskL_\text{mask}: power sum of LATHL_\text{ATH} and each masker’s thresholdadd powers, back to dB
BitsB=⌈(LSPL−Lmask)/6.02⌉B=\big\lceil(L_\text{SPL}-L_\text{mask})/6.02\big\rceil per coefficient0 under the mask
MDCT2M2M samples in, MM out, sine windowoverlap-add is exact

End of lesson 29.3

Where to go next.

Phasorium
LibraryEvery lesson, in order

Parts

About Phasorium
Look