Pitch (F0)
Fundamental frequency — the rate at which your vocal folds vibrate, measured in Hertz — is the strongest single acoustic cue to perceived gender in voice (Leung et al. 2018 report it accounts for ~42% of listener-judgement variance in their English-corpus study). Vocal Analyzer treats pitch as the primary cue with the strongest direct perceptual validation, and it is the only metric whose reference bands are drawn from converging calibrations across several studies (Gelfer & Bennett 2013; Davies et al. 2015; Holmberg et al. 2010).
Reference bands
The spectrum bar at the top of the dashboard overlays the following bands onto your median F0 (in Hz):
| Band | Low | High |
|---|---|---|
| Masculine | 90 | 155 |
| Androgynous | 155 | 165 |
| Feminine | 165 | 255 |
Caveat: Bands are reference overlays, not clinical cutoffs. They describe published central tendencies for English adult speech, not validated predictors of individual listener judgement.
Reported values
Vocal Analyzer reports several pitch statistics, each serving a different question:
Median F0
The central value of voiced F0 across the whole recording. This is what the spectrum bar plots. Less sensitive to occasional creak or falsetto frames than the mean, which is why it — not the mean — is the headline number.
F0 distribution
The 5th percentile, median, and 95th percentile of voiced F0, plus the interquartile range (IQR = p75 − p25). Together these summarise how tightly F0 clusters around the median. A narrow IQR with close p5/p95 values means a fairly monotone speaker; a wide IQR with distant percentiles means a very expressive speaker.
Pitch expressiveness
Standard deviation of F0 in semitones. Vocal Analyzer reports this in semitones rather than Hz because a semitone represents the same perceptual step whether you are speaking at 100 Hz or 250 Hz (Pépiot 2014). Typical conversational English speech sits around 2–4 ST; sustained vowels are usually under 1 ST; highly animated speech can exceed 5–6 ST.
The editable expressiveness overlay defaults to 3–6 ST. This is a deliberately broad continuity target for tracking practice over time, not a gendered or clinical reference range: language, task, affect, and speaking style can all move pitch variation substantially. All reference profiles use the same default, and Settings can reshape it for an individual goal.
Pitch range
Lowest and highest detected voiced F0 in the recording. Reported as
raw min–max in Hz. Sensitive to occasional creak frames at the low
end and to falsetto or noise artifacts at the high end — the F0
distribution tile (above) is a more stable view.
Pitch dynamism quotient (PDQ)
Standard deviation divided by the mean. A scale-invariant measure of variability (Leyns et al. 2024). Reported as a descriptive cue without reference bands.
Percentage voiced
Fraction of the recording that contained voiced speech. Long pauses or unvoiced consonants lower this number. Very low values (under ~40%) usually mean the recording has long silences or a lot of unvoiced sound — the median F0 is still valid if it crosses the minimum- duration threshold for a stable estimate.
Algorithm
Vocal Analyzer uses a two-pass, adaptive F0 extraction based on De Looze & Hirst (2008):
- The first pass uses a wide default range (the settings you see in the sidebar — pitch floor / ceiling). The median of that first pass is used to derive speaker-adaptive bounds.
- The second pass runs with those tightened bounds. Adaptive bounds give stabler F0 tracks and better reject creak and falsetto outliers than a fixed wide range would.
The effective_floor_hz and effective_ceiling_hz fields in the raw
result record the bounds actually used for the second pass; these are
surfaced in the Detailed Measurements card on the dashboard.
Under the hood the extraction is Praat's autocorrelation pitch tracker
via parselmouth. See src/vocal_analyzer/engine/pitch.py in the
repository for the implementation.
Caveats and measurement warnings
- Creak and falsetto outliers. The
minandmaxfields are single-frame extremes. If they disagree with the distribution percentiles by a lot, one or two outlier frames are dominating the range. The distribution tile is a more reliable read. - Percentage voiced. A recording with very little voiced content produces a less stable median. The tool does not gate on a minimum voiced percentage — it only gates on absolute duration — so a 30-second recording with 10% voicing can still produce a median value that should be treated cautiously.
- Source-filter coupling at high F0. When F0 approaches the first formant (F1), pitch and formant measurements become coupled and the analyzer emits a warning. Pitch itself is the least affected of the acoustic measurements, but the formant-side measurements that depend on accurate F0 can drift.
- SNR. Noisy recordings degrade voice quality and HNR first, formants next, pitch last. Pitch is the most robust cue under noisy conditions, and the dashboard's SNR warning explicitly says so.
References
- Leung, Y., Oates, J., & Chan, S. P. (2018). Voice, articulation, and prosody contribute to listener perceptions of speaker gender. Journal of Speech, Language, and Hearing Research 61(2), 266–280.
- Gelfer, M. P., & Bennett, Q. E. (2013). Speaking fundamental frequency and vowel formant frequencies: effects on perception of gender. Journal of Voice 27(5), 556–566.
- Davies, S., Papp, V. G., & Antoni, C. (2015). Voice and communication change for gender nonconforming individuals: giving voice to the person inside. International Journal of Transgenderism 16(3), 117–159.
- Holmberg, E. B., Oates, J., Dacakis, G., & Grant, C. (2010). Phonetograms, aerodynamic measurements, self-evaluations, and auditory perceptual ratings of biologically female voices converted to male. Journal of Voice 24(5), 511–522.
- De Looze, C., & Hirst, D. J. (2008). Detecting changes in key and range for the automatic modelling and coding of intonation. Speech Prosody 2008.
- Pépiot, E. (2014). Male and female speech: a study of mean F0, F0 range, phonation type and speech rate in Parisian French and American English speakers. Speech Prosody 2014.
- Leyns, C., Adriaansen, A., Daelman, J., Tomassen, P., Corthals, P., T'Sjoen, G., & D'haeseleer, E. (2024). Voice assessment pre- and post-gender affirming hormone therapy: a systematic scoping review. Journal of Voice.