Formants and resonance
Formants are the acoustic resonances of your vocal tract. Because the vocal tract acts as a variable-geometry resonator sitting on top of the vibrating vocal folds, each spoken sound has a characteristic pattern of peaks in its spectrum — F1, F2, F3 (and sometimes F4) — that encodes both the phonetic identity of the sound and the size and shape of the tract that produced it. Vocal Analyzer reports formants both as whole-recording medians and, when enough vowels are detected, as per-vowel aggregates.
What each formant carries
- F1 varies mainly with vowel height — how open your jaw is. Low F1 → high vowels (like /i/ in "see"); high F1 → low vowels (like /ɑ/ in "father").
- F2 varies mainly with vowel frontness — how far forward your tongue is. High F2 → front vowels (/i/); low F2 → back vowels (/u/).
- F3 varies with lip rounding and with vocal-tract shape details; it tends to be more stable across vowels than F1 or F2.
- F4 (when present) largely reflects overall tract length.
Across speakers, all formants tend to shift upward as apparent vocal- tract length shortens. This is why the whole-recording formant average of a vowel-varied recording is a reasonable size proxy even though any single vowel's formants encode phonetic information first.
Whole-recording vs per-vowel
Vocal Analyzer's formant section has two aggregation modes:
- Whole recording — the median F1/F2/F3 across every voiced frame, regardless of vowel. Fast, stable, and the historical default. This is the benchmark-continuity baseline. Changing the formant mode mid-session changes the aggregation, not your voice.
- Per vowel — when vowel timing is available, per-class F1/F2 aggregates with per-class stability scores. The current timing source can be heuristic vowel detection or optional aligned vowel intervals. The vowel classes are loosely inspired by the classic Hillenbrand et al. (1995) vowel corpus partitioning.
Both modes can render on the dashboard at once. See Goal profiles for why the whole-recording mode is the continuity anchor.
Formant dispersion
Formant dispersion is the average spacing between successive formants. It serves as a whole-recording resonance proxy that is largely vowel-independent: lower dispersion corresponds to a longer apparent vocal tract, and higher dispersion corresponds to a shorter one. This is the second spectrum bar on the dashboard.
The bands in use (unit: Hz):
| Band | Low | High |
|---|---|---|
| Masculine | 990 | 1100 |
| Androgynous | 1100 | 1175 |
| Feminine | 1175 | 1300 |
Caveat: Formant-dispersion bands are continuity-oriented heuristics, not clinical cutoffs. They track apparent vocal-tract size and are useful for trend tracking, but have weaker direct perceptual validation than speaking F0. Treat the bar as a reference overlay, not a verdict.
The boundaries above center on the overlap region identified in the 2026 scientific audit, with the feminine onset at 1175 Hz aligning with the audit's stated female central tendency of 1120–1210 Hz (Pisanski & Rendall 2011, Reby & McComb 2003).
Vocal tract length
The VTL (cm) tile derives an estimated vocal-tract length from formant dispersion via Reby & McComb's (2003) regression method. Shown as a concrete, physically-interpretable version of the dispersion proxy; it moves monotonically with dispersion, so the two tiles will never disagree — the VTL tile exists for when centimetres are more intuitive than hertz.
See also Fitch (1997) for the acoustic theory linking vocal-tract length to formant spacing.
F1 / F2 / F3 medians
The Detailed Measurements card lists the raw whole-recording medians in Hz. These are the values the dispersion calculation runs on. F4 is reported when the extractor could find a stable fourth formant.
If you want to see which vowels contributed which values, the per- vowel aggregates and the Vowel Space (F1 × F2) chart tell that story.
Per-vowel coverage
When enough distinct vowel classes are detected, a Per-vowel coverage tile reports:
- The number of detected vowel segments.
- The number of distinct vowel classes represented.
- A mean stability score across segments (higher = more stable tracking).
Low class count in a long recording (for example, three classes across 30 s of speech) triggers a "narrow coverage" flag: the per- vowel aggregates exist but are drawn from a small sample of the English vowel space, so treat them as a trend cue rather than a precise characterization.
If aligned vowel intervals are available, they provide the timing windows for the per-vowel path. They do not change the whole-recording formant baseline, and they do not make the per-vowel result a clinical or perceptual claim.
Formant ceiling
The Formant ceiling row in Detailed Measurements is the upper frequency limit used for tracking. Vocal Analyzer adapts this to the speaker when possible (high ceiling for higher-F0 voices, lower for lower-F0 voices). A formant ceiling tuned to the wrong range causes the tracker to merge or split adjacent formants; the adaptive choice reduces that failure mode.
Algorithm
Under the hood formants are extracted with Praat's Burg LPC analysis (via parselmouth). The analyzer:
- Estimates F0 first (see Pitch).
- Picks a formant ceiling from the F0-derived estimate.
- Runs Burg LPC on voiced frames.
- Aggregates the per-frame tracks to medians for whole-recording mode and to per-class aggregates for per-vowel mode.
Caveats and measurement warnings
- Burg instability at high F0. Above roughly 200 Hz, LPC-based formant estimates become less reliable because the harmonics are more widely spaced and the LPC order does not align as cleanly with vocal-tract resonances. The dashboard emits a warning when it detects this condition.
- Source-filter coupling. When F0 is within ~80% of F1, the source (voice folds) and the filter (vocal tract) couple acoustically and both F0 and formant estimates become biased. This case triggers its own warning.
- Formant ceiling selection. If the adaptive ceiling is turned off via the sidebar, the default ceiling may under- or over-shoot the speaker's actual formant range; leave it on unless you have a specific reason to override.
- Sparse harmonics coupling. Very high-F0 voices with few harmonics under the formant ceiling cause formant tracking to collapse onto individual harmonics. This overlaps with the Burg- instability and high-F0-suppression warnings.
- Sample rate below 16 kHz. The formant pipeline is skipped entirely below 16 kHz because the 8 kHz Nyquist ceiling puts F3 / F4 into the noise floor. The dashboard will show unavailable formant tiles and a clear warning in that case.
References
- Pisanski, K., & Rendall, D. (2011). The prioritization of voice fundamental frequency or formants in listeners' assessments of speaker size, masculinity, and attractiveness. The Journal of the Acoustical Society of America 129(4), 2201–2212.
- Reby, D., & McComb, K. (2003). Anatomical constraints generate honesty: acoustic cues to age and weight in the roars of red deer stags. Animal Behaviour 65(3), 519–530.
- Fitch, W. T. (1997). Vocal tract length and formant frequency dispersion correlate with body size in rhesus macaques. The Journal of the Acoustical Society of America 102(2), 1213–1222.
- Hillenbrand, J., Getty, L. A., Clark, M. J., & Wheeler, K. (1995). Acoustic characteristics of American English vowels. The Journal of the Acoustical Society of America 97(5), 3099–3111.
- Gelfer, M. P., & Bennett, Q. E. (2013). Speaking fundamental frequency and vowel formant frequencies: effects on perception of gender. Journal of Voice 27(5), 556–566.