Sound doesn't exist in a vacuum.

It lives in rooms. It bounces off walls. It wraps around your head in ways that are entirely, irrevocably yours.

So why do we settle for generic?

Most spatial audio gives you an approximation. An average. A calculated guess at how sound might reach human ears based on statistical models and dummy head measurements.

But you're not average. Your ears aren't average. The way sound reaches your left ear milliseconds before your right, the unique folds and curves that color everything you hear—these are yours alone.

The Question

What if headphones could deliver not a simulation, but a translation?

Your actual room. Your actual speakers. Your actual hearing. Preserved and transported to wherever you happen to be listening.

This was the starting point for 12to2.

Not "how do we make spatial audio?"

But "how do we honor the space that already exists?"

The Capture

It begins with listening.

Not with equipment lists or technical specifications. With the simple act of placing tiny microphones where your ears actually are—deep in the canal, where sound naturally arrives—and recording what happens when each speaker speaks.

Twelve speakers. Seven surround positions plus four overhead channels plus the sub that you feel more than hear.

Each one sends out a logarithmic sweep, 20Hz to 20kHz, a sonic flashlight illuminating the path from speaker to ear. The microphones capture not just the direct sound, but the reflections. The room's character. The way your particular geometry shapes everything.

This is where most systems stop. They use databases. Averages. The CIPIC dataset. The KEMAR dummy head.

But databases don't capture your front-left speaker positioned exactly 30 degrees off-center in your room with your specific absorption characteristics.

They don't capture the slight timing difference between your ears—the interaural time difference that your brain has spent a lifetime learning to interpret as direction and distance.

So we capture it all.

The Discovery

Raw impulse responses are messy.

They contain the truth, but also the noise. The late reflections that smear the image. The artifacts of recording. The inevitable imperfections of any real-world measurement.

The question becomes: what do we keep?

Through experimentation, a window emerges. About 120 milliseconds—enough to preserve the direct sound and early reflections that give a room its character, while discarding the late reverb that clouds localization.

Timing matters. All twelve speakers must speak with one voice, synchronized to the same moment. Peak alignment ensures that when the center channel says "now," it means the same "now" as the left surround.

And critically: the left and right ears must be processed together. Not normalized independently, but as a pair, preserving the level differences that tell your brain whether a sound is above, below, or straight ahead.

The Real-Time Renderer

Now comes the impossible task.

Take twelve channels of immersive audio. Convolve each with a personalized 8192-sample impulse response (left ear and right ear, separately). Sum to stereo. Apply headphone compensation. Apply personal EQ.

Do it in real-time. With no perceptible delay.

The naive approach—processing the full impulse response in one go—adds 170 milliseconds of latency. Too slow. The brain notices.

The solution comes in partitions.

Break the 8192 samples into 256-sample chunks. Process each chunk in the frequency domain using FFT convolution. Sum the results with overlap-save. The latency drops to 18 milliseconds—below the threshold of perception.

12ch Atmos → Per-Speaker HRTF → Binaural Mix → Headphone FIR → David G EQ → Stereo

It's not magic. It's just listening carefully to what the math reveals.

The Personal Touch

Headphones have their own voice.

Every model colors sound differently. Some boost the bass. Some emphasize the treble. Some are surprisingly flat. To translate accurately, we must first understand the translator.

So we capture the headphone's own frequency response. Measure what it does to a known signal. Create an inverse filter that restores neutrality.

But neutrality isn't the end goal.

David Griesinger—researcher, psychoacoustician, explorer of human hearing—discovered something crucial about how we perceive loudness. Our ears adapt. Given a constant tone, they adjust, normalizing, making everything seem equally loud over time.

The solution? Rapid comparison. Don't give the ear time to adapt. Alternate between a reference tone and a test tone so quickly that the brain must judge them in the moment, before auto-gain control kicks in.

The result is a personalized EQ curve. Twelve frequency bands, from 100Hz to 10kHz, tuned to your specific hearing. Not a generic "smiley face" EQ. Your curve. Your compensation.

What It Sounds Like

Close your eyes.

The sound isn't "in your head." It's out there. Front left. Rear right. Directly overhead—a position that standard binaural processing struggles to render convincingly.

You can point to where the sound is coming from. Not approximate. Precise.

This is the difference between remembering a place and being there. Between seeing a photograph and looking out a window.

Generic HRTFs give you spatial audio. Personalized HRTFs give you your room, your speakers, your hearing.

Why This Matters

As a mixing engineer, I need to trust what I hear.

Headphones are essential for detail work—the subtle panning, the micro-dynamics, the elements that get lost in room reflections. But standard headphone monitoring doesn't translate to speaker systems. What sounds spacious on headphones often sounds cluttered on speakers. What sounds intimate often sounds distant.

The personalized approach bridges this gap.

What I hear on headphones is what listeners will hear on their systems. Not because I've learned to compensate for the difference, but because there is no difference. It's the same acoustic information, just routed through different transducers.

But there's something deeper here.

Spatial audio is personal. Our ears are as unique as our fingerprints. Tools that acknowledge this individuality—rather than forcing everyone into an average—produce more accurate results.

More importantly, they produce more satisfying results. Because satisfaction comes from truth. From alignment between what we hear and what actually is.

What's Next

The system works.

It captures. It processes. It renders. It sounds like being in the room.

But there's always further to go. A Tauri GUI to make device selection intuitive. A full Rust migration for consistency. Spectral denoising to extract cleaner impulse responses from imperfect captures.

The work continues.

Not because the current version is lacking, but because exploration is its own reward. Because each refinement reveals something new about how we hear, how we perceive, how we experience sound.


C
Dolby Atmos Mixing Engineer
Exploring the boundaries of spatial audio and personalized listening experiences from New Zealand.