Cem Tuncer logoCem TuncerSound & Music
← Back to Guide
Sound Design4 August 2026 · 8 min read

Spatial Audio Design: How Do You Build the Sphere?

Confining sound not to speakers, but to a three-dimensional sphere: Ambisonics, HRTF, and the math behind how our ears locate a sound source.

Traditional stereo and 5.1 mixes tie sound to fixed speaker positions. Ambisonics takes the opposite approach: it encodes sound not to specific speakers, but to an abstract sphere surrounding the listener. That sphere can then be 'decoded' to whatever speaker layout or headphones are being used — which is exactly what makes Ambisonics ideal for VR and 360° video.

B-Format: W, X, Y, Z

First-order Ambisonics' core recording format consists of four channels known as B-Format: W (the omnidirectional pressure component — overall level), X (front-back), Y (left-right) and Z (up-down). Together, these four channels mathematically encode the sound scene's three-dimensional direction and depth — this is the foundation that keeps a sound source correctly positioned even as the listener turns their head.

First-order Ambisonics has limited spatial resolution — four channels produce a relatively coarse map of the sphere. Second- and third-order Ambisonics (HOA — Higher Order Ambisonics) add extra channels on top of W/X/Y/Z, slicing the sphere more finely for sharper spatial separation, at the cost of file size and processing load multiplying too. The fact that VR productions often push to third or fourth order is a good indicator of how that trade-off gets managed in practice.

ITD and ILD: How Our Ears Locate Direction

  • ITD (Interaural Time Difference): a low-frequency sound reaches the two ears at slightly different moments; the brain converts that microsecond delay into directional information.
  • ILD (Interaural Level Difference): a high-frequency sound arrives weaker at the far ear because the head casts an acoustic 'shadow'; the brain uses that amplitude difference as another directional cue too.
Δt delayInteraural Time Difference< 1500 Hz — low frequencyhead shadowInteraural Level Difference> 1500 Hz — high frequency
Duplex Theory: below 1500 Hz the brain relies on arrival-time difference (ITD); above it, on the level difference created by the head's acoustic shadow (ILD) — exactly the dual mechanism Ambisonics decoders try to imitate.

HRTF: A Personal Signature

The Head-Related Transfer Function (HRTF) is an algorithm that mathematically models how sound changes as it passes over our shoulders, the shape of our head, and every fold of the outer ear. HRTF is the key to making a sound played over headphones feel like it's positioned outside the head at a specific point, rather than 'inside' it — and because everyone's ear shape is slightly different, 'personalised HRTF' remains an active area of research in immersive audio technology.

The biggest practical obstacle to personalized HRTF research is measurement difficulty: accurately extracting someone's real HRTF means measuring them in a specialized acoustic chamber with test tones from dozens of different angles — not a process that scales for consumer products. That's why companies are now developing machine-learning models that try to generate a generalized HRTF estimate from a photo of someone's ear, or from a handful of simple measurements.

Object-based formats like Dolby Atmos take Ambisonics' spherical logic a step further, positioning each sound as an individual 'object' — the mix is automatically recalculated regardless of what system the listener is using.
Spatial AudioAmbisonicsImmersive Audio
← Back to Guide