Spatial Audio Design: How Do You Build the Sphere?
Confining sound not to speakers, but to a three-dimensional sphere: Ambisonics, HRTF, and the math behind how our ears locate a sound source.
Traditional stereo and 5.1 mixes tie sound to fixed speaker positions. Ambisonics takes the opposite approach: it encodes sound not to specific speakers, but to an abstract sphere surrounding the listener. That sphere can then be 'decoded' to whatever speaker layout or headphones are being used — which is exactly what makes Ambisonics ideal for VR and 360° video.
B-Format: W, X, Y, Z
First-order Ambisonics' core recording format consists of four channels known as B-Format: W (the omnidirectional pressure component — overall level), X (front-back), Y (left-right) and Z (up-down). Together, these four channels mathematically encode the sound scene's three-dimensional direction and depth — this is the foundation that keeps a sound source correctly positioned even as the listener turns their head.
ITD and ILD: How Our Ears Locate Direction
- ―ITD (Interaural Time Difference): a low-frequency sound reaches the two ears at slightly different moments; the brain converts that microsecond delay into directional information.
- ―ILD (Interaural Level Difference): a high-frequency sound arrives weaker at the far ear because the head casts an acoustic 'shadow'; the brain uses that amplitude difference as another directional cue too.
HRTF: A Personal Signature
The Head-Related Transfer Function (HRTF) is an algorithm that mathematically models how sound changes as it passes over our shoulders, the shape of our head, and every fold of the outer ear. HRTF is the key to making a sound played over headphones feel like it's positioned outside the head at a specific point, rather than 'inside' it — and because everyone's ear shape is slightly different, 'personalised HRTF' remains an active area of research in immersive audio technology.