# Sound localization **Sound localization** is a listener's ability to identify the location of a [[Sound|sound]] source in direction and distance from the pattern it produces at the two ears. A single ear on its own gives almost no directional information; two ears, spaced roughly 18 cm (7.1 in) apart on either side of the head, give the brain two slightly different signals to compare, and it is the difference between them — not either signal alone — that carries direction. The microsim *Sound localization* draws a head from above with wavefronts sweeping across it: moving the source's azimuth and the tone's frequency changes which of the two cues, an arrival-time difference or a level difference, is doing the work at a given moment. The two cues have distinct physics. At low frequencies the wavelength is long compared with the head, so [[Diffraction|diffraction]] bends the sound around it largely undiminished, and the ears mainly disagree about *when* the wave arrives — the interaural time difference, or ITD. At high frequencies the head is large compared with the wavelength and casts a real [[Acoustic_shadow|acoustic shadow]], so the ears mainly disagree about *how loud* the wave is — the interaural level difference, or ILD. This split, long known as the duplex theory of sound localization, was proposed by Lord Rayleigh in 1907 from experiments with tuning forks of different pitch.[^rayleigh1907] Beyond the two ears, the folds of the [[Auditory_system|outer ear]] add direction-dependent notches to the spectrum that help resolve front from back and above from below, and animals with differently placed or independently movable ears solve the same geometry in their own ways. ## How sound reaches the brain Both ears receive the same waveform delayed and attenuated by different amounts depending on where the source sits. The [[Auditory_system|auditory system]] extracts the interaural time difference and the interaural level difference from this pair of signals and combines them, with the two ears acting jointly as the primary organ of directional hearing; the outer ear, head and torso reshape the sound on its way in and add cues neither ear alone could supply. ## Neural interactions The signals from the two ears converge early, in the superior olivary complex of the brainstem, where neurons compare arrival time and level across the two channels rather than analyzing each ear's signal separately first — an early stage of [[Neural_encoding_of_sound|neural encoding of sound]] specific to binaural direction. The medial superior olive is the classic site proposed for timing comparisons and the lateral superior olive for level comparisons, mapping the raw ITD and ILD onto a neural representation of azimuth before the signal ever reaches the cortex. ## Human auditory system Human sound localization works over the full range of [[Hearing_range|audible frequencies]], roughly 20 Hz to 20 kHz, but the two duplex cues below are most reliable well within that range, where the head's dimensions are neither too small nor too large relative to the wavelength involved. ### Duplex theory The interaural time difference for a source directly to one side and a spherical head of radius *a* is given by a formula credited to Woodworth (1938): ITD = (*a*/*c*)(θ + sin θ), where θ is the azimuth in radians measured from straight ahead and *c* is the [[Speed_of_sound|speed of sound]] in air, about 343 m/s (1,130 ft/s) at 20 °C.[^ups17][^woodworth1938] For a head radius of 8.75 cm and a source at 60°, this gives an ITD of about 0.49 ms and a path difference around the head of 16.7 cm — small numbers, but well inside the roughly 20 μs discrimination the ear achieves at its best.[^woodworth1938] The interaural level difference behaves oppositely: it is near zero at low frequencies, where diffraction lets sound pass the head almost unimpeded, and grows with frequency as the head becomes acoustically large — at 1 kHz and 60° azimuth the shadow is still weak, on the order of 2 dB, while at several kilohertz it can reach 20 dB or more.[^rayleigh1907] Rayleigh's duplex theory holds that ITD dominates below roughly 1.5 kHz, where the wavelength (23 cm at 1.5 kHz) is long enough for the auditory system to track phase unambiguously, and ILD dominates above it, where phase becomes ambiguous but the shadow is strong. *Try: move the azimuth slider around the head and slide the frequency control past 1.5 kHz; the ITD and ILD readouts trade off exactly as the duplex theory predicts, and the wavefront field shows the shadow forming behind the head.* ### Pinna filtering effect The pinna, the visible folds of the outer ear, reflects sound within its own small cavities with delays of a fraction of a millisecond, adding peaks and notches to the spectrum that depend on the sound's elevation and front-back position rather than its left-right azimuth. Because these spectral cues do not depend on having two ears, they are the principal way a listener resolves elevation and distinguishes a source in front from its mirror image behind — a distinction the ITD and ILD alone cannot make. ### Other cues Head movement converts a static, ambiguous pair of ITD and ILD readings into a changing sequence: turning the head toward a sound changes its azimuth relative to the head in a predictable way, and the auditory system uses that change to break ties that a single snapshot cannot resolve. Visual and other sensory cues can also bias or override acoustic localization, a phenomenon exploited deliberately in ventriloquism and in film sound, where a voice is heard as coming from an actor's moving lips rather than from the loudspeaker actually producing it. ### Distance of the sound source Interaural cues carry direction well but distance poorly, because ITD and ILD depend on azimuth, not range. Distance is instead judged from the loss of high [[Frequency|frequencies]] with distance from atmospheric absorption, the changing ratio of direct to reverberant sound as a listener moves through a room, and, for very close sources within about a meter, an ILD that grows unusually large because the head then intercepts a real fraction of the wavefront's curvature rather than acting on a nearly plane wave. ### Signal processing Because the ITD and ILD depend continuously on frequency and azimuth, artificial sound reproduction can recreate them from a recorded or synthesized signal by filtering it appropriately for a chosen direction — the basis of binaural recording and of the [[Head-related_transfer_function|head-related transfer function]], which captures a listener's or a dummy head's own filtering of every direction as a pair of impulse responses. ## Specific techniques with applications ### Auditory transmission stereo system Two-channel stereo reproduces a left-right ILD (and, in some formats, a small ITD) between two loudspeakers or headphone channels, giving a listener a left-to-right image of a recorded sound stage without attempting to reproduce elevation or front-back cues. ### 3D para-virtualization stereo system Binaural synthesis over headphones applies a measured or modeled [[Head-related_transfer_function|head-related transfer function]] to a signal so that, even through two channels, a listener perceives a source at an arbitrary azimuth and elevation, including behind or above — cues ordinary stereo cannot deliver because it reproduces only azimuth. *Try: the sim's shadow field shows why azimuth alone is not enough — see the [[3D_sound_localization|cone of confusion]] variant, where a source behind the head and its front mirror image share the same ITD and ILD.* ### Multichannel stereo virtual reproduction Surround-sound formats add loudspeakers behind and to the side of the listener so that ITD and ILD cues can be generated directly by real sources at the intended angles rather than synthesized through headphone filtering, extending localizable direction beyond what two front channels can provide. ## Animals Animals face the same physics with ears set at different spacings, on different axes, or with independent mobility, so the balance between ITD and ILD, and the tricks used to resolve front-back ambiguity, differ by species and habitat. ### Lateral information (left, ahead, right) For most terrestrial animals, as for humans, the basic left-right cue is the pair of ITD and ILD generated by two ears spaced across the head, with the crossover frequency between the two cues set by the animal's own head size — a smaller head shifts the crossover higher because it takes a shorter wavelength before the head becomes acoustically large. ### In the median plane (front, above, back, below) Sources on the vertical midline plane produce equal ITD and ILD at both ears, so animals resolve front, back, above and below chiefly through pinna filtering and head movement, the same strategy used for elevation in humans, though the shape and mobility of the outer ear varies widely between species. ### Head tilting Many birds and small mammals tilt or bob the head when localizing a sound, converting the ambiguous cues from a single, still position into a changing sequence across a known head movement, functionally similar to the way a human listener turns toward an uncertain sound. ### Localization with coupled ears (flies) The parasitoid fly *Ormia ochracea* has ears spaced only about 0.5 mm apart — far too close to generate a usable ITD by direct propagation delay alone — yet it localizes its cricket host's calling song with sub-degree precision because its two eardrums are mechanically coupled through a connecting structure that amplifies the tiny time difference into a much larger one, a solution that has inspired miniature directional microphones. ### Bi-coordinate sound localization (owls) Barn owls localize prey in complete darkness using both ITD and ILD simultaneously as independent cues to two different coordinates: because the owl's right ear opening is asymmetrically higher than the left, ITD chiefly signals azimuth while ILD chiefly signals elevation, letting the two ears together fix both coordinates of a source rather than only one. ### Dolphins Dolphins and other odontocetes localize primarily through active [[Animal_echolocation|echolocation]] rather than passive listening to ambient sound, sending out clicks and judging the direction and range of an object from the returning echo, much as [[Sonar|sonar]] does artificially; underwater sound travels at roughly 1,500 m/s, more than four times its speed in air, which changes the timing of every cue relative to the terrestrial case. ## History The duplex theory's clearest early statement is Lord Rayleigh's 1907 paper "On our perception of sound direction," which reported that tuning forks of different pitch were localized with different reliability and traced the difference to the relative importance of phase (time) and intensity cues at different frequencies — the foundation on which the modern ITD/ILD account still rests.[^rayleigh1907] A quantitative geometric model of the time difference itself, treating the head as a rigid sphere, was worked out by Robert Woodworth and published in his 1938 textbook *Experimental Psychology*; the formula in the microsim above is Woodworth's.[^woodworth1938] ## Minnesota *This section is specific to Wikitube.* No sourced Minnesota-specific case for sound localization research has been verified for this article; none is asserted here. *Citation needed* — a Minnesota audiology, hearing-aid, or bioacoustics laboratory working on directional hearing would settle this. ## See also - [[Head-related_transfer_function]] - [[3D_sound_localization]] - [[Hearing]] - [[Underwater_acoustics]] - [[Bioacoustics]] - [[Doppler_effect]] ## References [^ups17]: OpenStax, *University Physics Volume 1* (2016), §17.2 "Speed of Sound," speed of sound at sea level 343 m/s; §17.3 "Sound Intensity." [^rayleigh1907]: Strutt, John William (Lord Rayleigh) (1907). "On our perception of sound direction." *Philosophical Magazine* 13 (74): 214–232. https://doi.org/10.1080/14786440709463595 [^woodworth1938]: Woodworth, Robert S. (1938). *Experimental Psychology*. Henry Holt and Company. The interaural-time-difference formula for a rigid spherical head, ITD = (a/c)(θ + sin θ), is standardly credited to this source; Wikitube has not independently confirmed the page number and marks the attribution *citation needed* for the exact passage. <!-- ACOUSIM:BEGIN g22 — Acoustics portal microsim (framework build, specs/acoustics/sims/Sound_localization.json); do not hand-edit inside --> **Microsim — three.js (Wikitube framework):** *Sound localization* <div class="wt-sim" data-src="https://wikitube-3d-microsims.netlify.app/acoustics/Sound_localization.html" data-title="Sound localization"></div> *Built from `MICROSIM_GUIDE/specs/acoustics/sims/Sound_localization.json`; part of the [[PORTAL_Acoustics|Acoustics portal]] spine (section sims and See-also variants).* <!-- ACOUSIM:END --> ## Wikipedia : Wikitube **Strict pair:** [Wikipedia](https://en.wikipedia.org/wiki/Sound_localization) : [Wikitube](https://en.wikitube.io/wiki/Sound_localization) - skeleton pinned to revision 1372001115 (2026-09-11). <!-- hub tags: GENERATIVE; Centers_of_Excellence; PORTAL_Acoustics section 21 -->