Robert Netzorg

EECS Department, University of California, Berkeley

Technical Report No. UCB/EECS-2026-198

May 15, 2026

http://www2.eecs.berkeley.edu/Pubs/TechRpts/2026/EECS-2026-198.pdf

Speech technology operates on representations of speaker identity that are either too high-level to be actionable (e.g., "masculine" or "old") or too high-dimensional to be interpretable (e.g., learned embeddings). How can we bridge this gap to enable interpretable analysis and modification of what makes a voice sound the way it does? I present research addressing this question in three stages. First, I establish that perceptual voice qualities, which are standardized descriptors of voice from speech pathology, speech forensics, and gender-affirming voice training, are reliably perceivable by non-experts and exhibit disentangled correlations with acoustic and physical measurements of the vocal tract. Second, I present the PQ-Encoder, a model that predicts these qualities from speech, and demonstrate that a compact 9-dimensional perceptual vector captures meaningful information about speaker identity, gender, and age. Finally, I introduce StyleMod, a conditional flow matching model that modifies speaker identity along perceptual voice quality axes, enabling interpretable and continuous manipulation of perceived gender and age while preserving non-targeted speaker attributes. Together, this work demonstrates that grounding speech technology in human perception yields representations that are both scientifically meaningful and practically useful for controllable voice generation.

Advisors: Bin Yu


BibTeX citation:

@phdthesis{Netzorg:EECS-2026-198,
    Author= {Netzorg, Robert},
    Title= {Perceptually Grounded Modeling and Modification of Speaker Identity},
    School= {EECS Department, University of California, Berkeley},
    Year= {2026},
    Month= {May},
    Url= {http://www2.eecs.berkeley.edu/Pubs/TechRpts/2026/EECS-2026-198.html},
    Number= {UCB/EECS-2026-198},
    Abstract= {Speech technology operates on representations of speaker identity that are either too high-level to be actionable (e.g., "masculine" or "old") or too high-dimensional to be interpretable (e.g., learned embeddings). How can we bridge this gap to enable interpretable analysis and modification of what makes a voice sound the way it does? I present research addressing this question in three stages. First, I establish that perceptual voice qualities, which are standardized descriptors of voice from speech pathology, speech forensics, and gender-affirming voice training, are reliably perceivable by non-experts and exhibit disentangled correlations with acoustic and physical measurements of the vocal tract. Second, I present the PQ-Encoder, a model that predicts these qualities from speech, and demonstrate that a compact 9-dimensional perceptual vector captures meaningful information about speaker identity, gender, and age. Finally, I introduce StyleMod, a conditional flow matching model that modifies speaker identity along perceptual voice quality axes, enabling interpretable and continuous manipulation of perceived gender and age while preserving non-targeted speaker attributes. Together, this work demonstrates that grounding speech technology in human perception yields representations that are both scientifically meaningful and practically useful for controllable voice generation.},
}

EndNote citation:

%0 Thesis
%A Netzorg, Robert 
%T Perceptually Grounded Modeling and Modification of Speaker Identity
%I EECS Department, University of California, Berkeley
%D 2026
%8 May 15
%@ UCB/EECS-2026-198
%U http://www2.eecs.berkeley.edu/Pubs/TechRpts/2026/EECS-2026-198.html
%F Netzorg:EECS-2026-198