Written & Presented by Associate Professor Lorne Bregitzer
Over the past couple of years, there has been an increased availability of 3D or Spatial audio. You can hear this 3D audio on headphones and earbuds from Apple, Sony, VR headsets, and many others. How does this 3D sound work, and how do they allow us to hear audio all around us just through two speakers on our ears? I’m Lorne Bregitzer, the Audio Professor, and I’m here to explain it to you.
Most of us can recognize the location of audio in our day-to-day lives primarily through our two ears. We can hear which direction a car is coming from, somebody calling our name from behind us, or an airplane flying overhead. With two ears, we can localize the direction of the sound. Not just in 360 degrees around us, but also in sound above us and below.
Normally when we listen to music or movies over headphones, the audio is pretty much presented in a 180 degree spread in front of us. The position of the audio we’re hearing is basically determined by the difference in level between our left and right ears.
3D or spatial audio incorporates the ability to localize audio in 360 degrees around us, as well as the elevation.
So how does sound get translated into 3D audio when we listen to it on headphones?
In order to create Spatial Audio, we need to understand how we perceive the localization of sound with our ears in the real world.
Well there are three main factors on how we perceive the position or localization of sound. They consist of the Inter-Aural Time Difference, Inter-Aural Level Difference, and the timbre change or frequency difference of a sound in between our left and right ears.
Inter-Aural Level Difference, or ILD, is the difference in level between our two ears, which is decoded by our brain to tell us where the audio is localized. This is how we regularly hear audio positioned in stereo sound over headphones.
If an audio source is on our left, it’ll be slightly louder in the left ear than the
Audio Panning Clip
Delayed audio clip right ear. Then our brain tells us the sound is coming from the left.
Inter-Aural Time Difference, is the slight difference in time when audio hits our ears. Since our ears are approximately 20cm apart, when there’s a sound on our right, it’ll hit the left ear at a slightly later time than the right ear. The time difference
Head Shadow clip is approximately 6 milliseconds or 6/thousandths of a second.
The third factor is the frequency or timbre difference between our two ears. If there’s a sound coming from the left side, the frequency content of of that sound is different in the right ear than it is in the left. This is referred to as the “head shadow” effect. The head will block higher frequencies while lower frequencies
Shadow Clip will wrap around our head.
In the same way our head will block light from one side, it does something similar to the audio. However, it doesn’t completely block the sound, it changes the frequency content of that sound in each ear.
With two different frequency responses in different ears, our brain will localize
Audio elevation clip the sound accordingly.
The height, or elevation of the sound is decoded by our brains based upon slight variations of the head shadow effect. This changes based upon the height of the sound source.
HRTF Frequency Graphic
In this graphic, you can see the difference in frequencies of a single sound source heard in both the left and right ears coming from the left side..
It’s the combination of these three of these localization techniques that is used to reproduce the sound to create 3D spatial audio, binaurally.
Binaural Microphones, historical photos
I used that term Binaural. 3D audio is the same thing as listening to binaural audio. Binaural audio has been around since the late 1800s. It involves placing microphones in such a way to mimic the position of the human ears. When played back in headphones, the Time, Level, and Frequency differences are captured with the audio.
In order to artificially recreate this binaural audio, a Head-Related Transfer Function or HRTF is created which captures all of these localization parameters across all heights and angles.
The audio is then processed with these HRTFs to transform them from various mono sound sources to somewhere in a 3 dimensional space in your headphones.
This is an example of just one of the types of software that can be used to artificially create a 3D binaural sound from a single sound source.
Regular headphones can playback 3D encoded audio. However, to have that encoded in real-time requires additional processing. This can be done in the playback device, like a video game console, such as when I listen to 3D audio through Sony’s Pulse 3D headphones while playing my PS5. Or it can be done in the headphones themselves, such as Apple’s AirPods Pro, when I’m watching a movie.
#spatialaudio #3Daudio