Vocal Analyzer Docs

← Back to app

Voice and gender perception

This page is the background explanation layer under everything else in these docs. It summarises why the cues Vocal Analyzer measures matter perceptually, how they relate to each other, and which of them are strong vs weak predictors of listener judgement. The per-metric reference pages (pitch, formants, voice quality, etc.) treat the same material in depth; this page is the map.

What "gendered voice" means here

When researchers talk about gender perception in voice, they mean one specific thing: the judgement a listener makes about the apparent gender of an unknown speaker based on acoustic input alone. That is a perception, not a property of the speaker — listeners differ, and self-perception and listener perception are independent dimensions (Dacakis et al. 2017).

Vocal Analyzer works at the acoustic layer. It measures cues that listeners are known to use. It does not decide whether a voice sounds any particular way to any particular listener, and it never equates an acoustic measurement with a speaker's gender, identity, or self-perception.

The cue hierarchy

Research on listener judgement generally ranks the cues this way:

  1. Pitch (fundamental frequency, F0). The dominant single cue. Leung et al. (2018) report that F0 accounts for roughly 42% of listener-judgement variance in their English-corpus study. Gelfer & Bennett (2013), Davies et al. (2015), and Holmberg et al. (2010) calibrate the boundary region around 155–165 Hz where listeners stop confidently classifying voices as either masculine or feminine.
  2. Resonance / formant patterns. Formant frequencies reflect the acoustic resonances of the vocal tract. Shorter tracts shift formants up; longer tracts shift them down. This cue family is perceptually salient but has weaker direct validation than F0 (Pisanski & Rendall 2011). In Vocal Analyzer this shows up as individual formant values (F1, F2, F3, F4 medians), a whole- recording formant-dispersion resonance proxy, and per-vowel aggregates.
  3. Voice quality. Spectral tilt (H1-H2 and its corrections), CPPS, and the LTAS complements (alpha ratio, Hammarberg index, spectral slope) describe whether a voice sounds breathy, pressed, or neutral. These cues correlate with listener judgements but more weakly than F0 or resonance, and their measurement is sensitive to F0 above ~175 Hz (Klatt & Klatt 1990; Iseli & Alwan 2004). Vocal Analyzer treats voice quality descriptively, without a gender band overlay.
  4. Prosody and intonation contour. Pitch variability (expressed as standard deviation in semitones, per Pépiot 2014), pitch dynamism quotient, and pitch range. Weaker single-cue predictor than F0 median, but part of the perceptual picture.
  5. Sibilants. /s/-spectrum shape (CoG, spread, skewness, kurtosis) has been discussed as a gendered cue in some studies. Vocal Analyzer reports it at moderate confidence because the detector is heuristic and the acoustic shape is sensitive to the recording chain.

Everything below F0 is an additional cue, not a replacement: a voice with a feminine F0 median but masculine formants typically lands in the perceptual overlap zone, and listeners disagree more about it.

Reference bands, not targets

The bands Vocal Analyzer overlays on pitch and resonance are reference bands: they describe published central tendencies for perceived-masculine and perceived-feminine adult English speech. They are not clinical cutoffs:

For the exact boundaries used in the app, see Pitch and Formants.

The descriptive framing

Vocal Analyzer deliberately frames every result descriptively:

That framing exists because the research literature itself is descriptive. Published acoustic studies correlate cues with listener judgement; they do not prescribe target values for individual speakers. If you are doing active voice training, a voice teacher can translate those measurements into next steps — the app cannot. See Beyond the tool.

Further reading, in order of depth

References