Perceptually Grounded Modeling and Modification of Speaker Identity
Robert Netzorg
EECS Department, University of California, Berkeley
Technical Report No. UCB/EECS-2026-198
May 15, 2026
http://www2.eecs.berkeley.edu/Pubs/TechRpts/2026/EECS-2026-198.pdf
Speech technology operates on representations of speaker identity that are either too high-level to be actionable (e.g., "masculine" or "old") or too high-dimensional to be interpretable (e.g., learned embeddings). How can we bridge this gap to enable interpretable analysis and modification of what makes a voice sound the way it does? I present research addressing this question in three stages. First, I establish that perceptual voice qualities, which are standardized descriptors of voice from speech pathology, speech forensics, and gender-affirming voice training, are reliably perceivable by non-experts and exhibit disentangled correlations with acoustic and physical measurements of the vocal tract. Second, I present the PQ-Encoder, a model that predicts these qualities from speech, and demonstrate that a compact 9-dimensional perceptual vector captures meaningful information about speaker identity, gender, and age. Finally, I introduce StyleMod, a conditional flow matching model that modifies speaker identity along perceptual voice quality axes, enabling interpretable and continuous manipulation of perceived gender and age while preserving non-targeted speaker attributes. Together, this work demonstrates that grounding speech technology in human perception yields representations that are both scientifically meaningful and practically useful for controllable voice generation.
Advisors: Bin Yu
BibTeX citation:
@phdthesis{Netzorg:EECS-2026-198,
Author= {Netzorg, Robert},
Title= {Perceptually Grounded Modeling and Modification of Speaker Identity},
School= {EECS Department, University of California, Berkeley},
Year= {2026},
Month= {May},
Url= {http://www2.eecs.berkeley.edu/Pubs/TechRpts/2026/EECS-2026-198.html},
Number= {UCB/EECS-2026-198},
Abstract= {Speech technology operates on representations of speaker identity that are either too high-level to be actionable (e.g., "masculine" or "old") or too high-dimensional to be interpretable (e.g., learned embeddings). How can we bridge this gap to enable interpretable analysis and modification of what makes a voice sound the way it does? I present research addressing this question in three stages. First, I establish that perceptual voice qualities, which are standardized descriptors of voice from speech pathology, speech forensics, and gender-affirming voice training, are reliably perceivable by non-experts and exhibit disentangled correlations with acoustic and physical measurements of the vocal tract. Second, I present the PQ-Encoder, a model that predicts these qualities from speech, and demonstrate that a compact 9-dimensional perceptual vector captures meaningful information about speaker identity, gender, and age. Finally, I introduce StyleMod, a conditional flow matching model that modifies speaker identity along perceptual voice quality axes, enabling interpretable and continuous manipulation of perceived gender and age while preserving non-targeted speaker attributes. Together, this work demonstrates that grounding speech technology in human perception yields representations that are both scientifically meaningful and practically useful for controllable voice generation.},
}
EndNote citation:
%0 Thesis %A Netzorg, Robert %T Perceptually Grounded Modeling and Modification of Speaker Identity %I EECS Department, University of California, Berkeley %D 2026 %8 May 15 %@ UCB/EECS-2026-198 %U http://www2.eecs.berkeley.edu/Pubs/TechRpts/2026/EECS-2026-198.html %F Netzorg:EECS-2026-198