We’ll begin with acoustics and psychoacoustics because binaural audio, ambisonics, and Dolby Atmos all depend on how sound behaves and how we perceive it.
Over the semester, we’ll work with stereo, binaural audio, ambisonics, 5.1, and Dolby Atmos.
Today’s class ends with a listening session that compares these formats.
Fundamentals of sound
We can study sound in three ways: as physical vibrations and waves, as activity in the ear, and as a perceptual experience.
Sound, noise, and music are contextual labels. The same signal may fit any of them, depending on its purpose and setting.
How you locate a sound
According to Rayleigh’s duplex theory from 1907, interaural time differences help us locate low frequencies, while interaural level differences help us locate high frequencies.
The maximum time difference across a human head is about 650 microseconds.
The head shadow depends on frequency. For a sound 90 degrees to one side, the shadow measures about 10 dB at 3 kHz, 20 dB at 6 kHz, and 35 dB at 10 kHz. Below roughly 2 kHz, sound waves bend around the head and weaken this cue. That is why hearing relies on both time and level differences.
The cone of confusion
Sounds arriving from several directions can produce identical time and level differences. Those two cues alone cannot distinguish front from back or up from down.
The pinna resolves elevation with a spectral notch that slides from about 5 kHz for sounds straight ahead to about 10 kHz for sounds overhead.
Head movement provides more information. When you turn your head, both cues change and help reveal the sound’s direction.
How you judge distance
In a free field, the sound level falls by 6 dB each time the distance doubles. To make a sound seem half as far away, however, its level must increase by about 10 dB.
Indoors, the strongest cue is the ratio of reverberant to direct sound.
Familiar sounds can override these cues. A shout tends to sound far away, while a whisper tends to sound close, regardless of playback level.
Immersive audio technologies
Dolby Atmos
Atmos is an object-based format. Sounds carry position metadata, and a renderer maps them to the available speakers.
Ambisonics
Ambisonics is a scene-based format. It uses spherical harmonics to encode a full sphere of sound, which can then be decoded for different speaker layouts.
Listening session
We’ll listen to some 5.1 mixes with the time we have left.
Source
The localization and distance material comes from Elizabeth Wenzel, Durand Begault, and Martine Godfroy-Cooper, “Perception of Spatial Sound,” ch. 1 of Agnieszka Roginska and Paul Geluso, eds., Immersive Sound (Routledge, 2017).