# Formant > *For the singer's high-frequency resonance cluster, see the section [Singer's formant](#singers-formant) below.* A **formant** is a [[Resonance|resonance]] frequency of the vocal tract that reinforces certain harmonics of a voiced sound while leaving others comparatively weak, shaping the spectral envelope that a listener hears as a particular vowel or vowel-like quality. The vocal folds buzz out a source signal rich in harmonics of the [[Pitch_(music)|pitch]] frequency, f0, and the tract above the larynx — throat, mouth and sometimes nose — acts as a filter whose resonances lift the harmonics nearest them and let the rest fall away; changing the shape of that tract, by moving the tongue and lips, moves the formants and changes the vowel while the pitch stays whatever the speaker chooses. The microsim *Formant* builds a vowel this way in two stages, a bar chart of the glottal source's harmonics and a filter curve of the tract's resonances, and multiplies them to make the vowel's own spectrum. Formants are conventionally numbered from the lowest, F1, upward; for an adult male voice the first two or three formants of a given vowel typically fall in a range from a few hundred hertz to a few thousand, and it is chiefly the position of F1 and F2 relative to each other that a listener uses to identify which vowel is being spoken.[^peterson1952] Because the tract's resonances depend only on its shape and length, not on how fast the vocal folds vibrate, a soprano and a bass singing the "same" vowel at very different pitches keep formants in roughly the same place while their harmonic spacing is entirely different — the [[Source–filter_model|source-filter]] separation that gives the model its name.[^fant1960] ## History The source-filter account of vowel production was formalized by the Swedish phonetician Gunnar Fant in his 1960 monograph *Acoustic Theory of Speech Production*, which modeled the vocal tract as an acoustic tube and derived its resonances from the tube's changing cross-section, putting on a rigorous footing an idea already implicit in nineteenth-century acoustics — that the throat and mouth act as a [[Acoustic_resonance|resonant]] cavity shaping the voice's raw buzz.[^fant1960] An early physical description along these lines appears in standard acoustics teaching, which treats the throat and mouth as an air column closed at one end that resonates in response to vibration at the vocal folds, with the growth of the larynx at puberty explaining why the same vowel sits at different frequencies in adult men's and women's voices.[^ups175] A large, still-cited body of measured vowel formants for American English was published by Gordon Peterson and Harold Barney in 1952, based on recordings from 76 speakers, and their tabulated values are the reference numbers many later textbooks and speech systems still quote.[^peterson1952] ## Phonetics In phonetics the formants, especially F1 and F2, are the primary acoustic correlates of vowel quality: F1 correlates inversely with vowel height (a low, open vowel like the one in "hod" has a comparatively high F1, while a high, close vowel like the one in "heed" has a low F1), and F2 correlates with how far forward or back the tongue is (a front vowel raises F2, a back vowel lowers it).[^ups175] For the vowel in "heed," Peterson and Barney's adult male averages put F1 near 270 Hz and F2 near 2,290 Hz — a low F1 with the highest F2 in the vowel set — while for the vowel in "hod," F1 rises to roughly 730 Hz and F2 falls to roughly 1,090 Hz, the opposite corner.[^peterson1952] A neutral, unstressed vowel such as schwa corresponds to a vocal tract with no narrowing anywhere along its length, modeled as a uniform tube closed at the glottis and open at the lips whose resonances fall at the odd quarter-wave frequencies, *c*/4*L*, 3*c*/4*L*, 5*c*/4*L*, and so on; for a tract length of 17.5 cm and a warm, moist speed of sound of about 350 m/s this gives formants near 500 Hz, 1,500 Hz and 2,500 Hz.[^ups175] *Try: set the vowel control to "schwa (tube)" and watch the filter curve sit exactly on those odd multiples; then switch through the other vowel choices and watch F1 and F2 move to the measured Peterson–Barney positions while the harmonic bars underneath stay fixed by f0 alone.* The tract's resonances scale with its physical length: because the quarter-wave frequencies of a closed tube go as *c*/4*L*, shortening the tract raises every formant in proportion. A tract shortened to 14 cm from the reference 17.5 cm used above raises the schwa's resonances from 500, 1,500 and 2,500 Hz to about 625, 1,875 and 3,125 Hz — a simplified, uniform-scaling account of why children's voices and, on average, women's voices carry higher formants than adult men's, since a shorter tract is the dominant anatomical difference the model captures, even though real tracts also differ in shape, not only in length.[^ups175] Formants are not infinitely sharp: each behaves as a damped resonance with its own [[Q_factor|bandwidth]], commonly modeled at roughly 60 Hz for F1, 90 Hz for F2 and 120 Hz for F3, widening as frequency increases and setting how broad a peak looks in a spectrum rather than an idealized spike at one exact frequency. ## Formant estimation Because a formant is a peak in the spectral envelope rather than a single [[Harmonic|harmonic]], estimating it from a recorded voice means separating that smooth envelope from the fine, harmonic-by-harmonic detail riding on top of it — the same source-filter split the model is named for. Linear predictive coding is the most common practical method: it fits an autoregressive filter to short windows of the recorded waveform, and the resonant peaks of the fitted filter are taken as estimates of the tract's formants, a technique used throughout [[Speech_coding|speech coding]] and speech recognition. Cepstral smoothing of the spectrum is a related, older technique that reaches a similar separation by working in a transformed frequency domain rather than by direct autoregressive fitting. Both methods must contend with the same difficulty the microsim makes visible: at a high [[Pitch_(music)|pitch]] the harmonics sit hundreds of hertz apart, so the envelope is sampled only sparsely by the available partials, and an estimator — human or algorithmic — has less evidence to fix a formant's exact position; this is part of why formant estimation is markedly harder for high-pitched voices, such as sopranos singing above the range where their own formants were originally measured, than for typical speech. ## Formant plots Plotting the first two formants of a set of vowels against each other, conventionally with F1 on a downward-increasing vertical axis and F2 on a reversed horizontal axis, produces a vowel space whose overall shape resembles an inverted, skewed triangle or quadrilateral matching the physiological range of tongue positions from high-front to low-back; this F1–F2 plot is the standard way phoneticians summarize a speaker's or a language's vowel inventory in [[Psychoacoustics|psychoacoustics]] and phonetics alike, and Peterson and Barney's original 1952 measurements are still one of the datasets most often plotted this way.[^peterson1952] ## Singer's formant Trained singers, particularly in Western classical opera and choral technique studied in [[Musical_acoustics|musical acoustics]], can produce an additional strong resonance cluster around 2.8–3.5 kHz, called the singer's formant, that does not correspond to any single one of the usual vowel formants but arises when the third, fourth and fifth formants are brought close together by a lowered larynx and a widened space just above it (the laryngeal ventricle and epilaryngeal tube).[^fant1960] Because the singer's formant sits in a [[Hearing_range|frequency band]] where the ear is highly sensitive and where an orchestra's own energy is comparatively weak, it lets an unamplified operatic voice project over a full orchestra without simply singing louder across the whole spectrum — an efficient, narrow-band boost rather than a general increase in level. ## Minnesota *This section is specific to Wikitube.* No sourced Minnesota-specific case for formant or vocal-tract acoustics research has been verified for this article; none is asserted here. *Citation needed* — a Minnesota speech-and-hearing science program or opera training program with a documented acoustic study would settle this. ## See also - [[Source–filter_model]] - [[Ohm's_acoustic_law]] - [[Musical_acoustics]] - [[Acoustic_resonance]] - [[Hearing]] - [[Bioacoustics]] ## References [^ups175]: OpenStax, *University Physics Volume 1* (2016), §17.4 "Normal Modes of a Standing Sound Wave" (tube closed at one end: odd-harmonic resonances, f_n = n·v/4L) and §17.5 "Sources of Musical Sound," p. 836, Figure 17.26: the throat and mouth modeled as an air column closed at one end, and the effect of the larynx's growth at puberty on speech frequencies. [^fant1960]: Fant, Gunnar (1960). *Acoustic Theory of Speech Production*. Mouton & Co., The Hague. The source-filter model and the physiological account of the singer's formant (laryngeal ventricle and epilaryngeal tube) both trace to this and to later work summarizing it; Wikitube has not verified the singer's-formant mechanism against Fant's original text page by page and marks the specific physiological description *citation needed* pending that check. [^peterson1952]: Peterson, Gordon E.; Barney, Harold L. (1952). "Control methods used in a study of the vowels." *Journal of the Acoustical Society of America* 24 (2): 175–184. https://doi.org/10.1121/1.1906875 <!-- ACOUSIM:BEGIN g22 — Acoustics portal microsim (framework build, specs/acoustics/sims/Formant.json); do not hand-edit inside --> **Microsim — three.js (Wikitube framework):** *Formant* <div class="wt-sim" data-src="https://wikitube-3d-microsims.netlify.app/acoustics/Formant.html" data-title="Formant"></div> *Built from `MICROSIM_GUIDE/specs/acoustics/sims/Formant.json`; part of the [[PORTAL_Acoustics|Acoustics portal]] spine (section sims and See-also variants).* <!-- ACOUSIM:END --> ## Wikipedia : Wikitube **Strict pair:** [Wikipedia](https://en.wikipedia.org/wiki/Formant) : [Wikitube](https://en.wikitube.io/wiki/Formant) - skeleton pinned to revision 1340709526 (2026-09-11). <!-- hub tags: GENERATIVE; Centers_of_Excellence; PORTAL_Acoustics section 23 -->