Speaker Identity in Non-Verbal Vocalizations: Conditional Distillation and Mixture of Experts Approach

Interspeech

Tzu-Chieh Wei*, Yi-Cheng Lin*, Huang-Cheng Chou, Kuan-Yu Chen, Hsin-Yen Sung, Shrikanth Narayanan, Hung-yi Lee

Proc. Interspeech 2026 , 2026

Abstract

As expressive text-to-speech (TTS) and voice conversion (VC) systems increasingly generate non-verbal vocalizations (NVVs) to enhance naturalness, reliable speaker verification (SV) becomes essential to objectively assess identity consistency across both verbal and non-verbal segments. Yet current SV systems generalize poorly to NVVs, and fine-tuning on NVV data causes catastrophic forgetting of speech performance. We present the first systematic study across 10 NVV types and propose a framework combining frozen Data2Vec self-supervised features with ECAPA-TDNN, enhanced by a Mixture of Experts (MoE) module with learned domain-aware routing. A conditional distillation loss on speech inputs via a pretrained teacher retains speech-to-speech accuracy, while a contrastive loss bridges the speech-NVV domain gap. Our method reduces speech-NVV EER from 38.93% to 22.66% over a pretrained baseline, and improves speech EER from 13.17% to 9.24% via distillation.

BibTeX

@inproceedings{wei2026speaker,
  title = {Speaker Identity in Non-Verbal Vocalizations: Conditional Distillation and Mixture of Experts Approach},
  author = {Wei, Tzu-Chieh and Lin, Yi-Cheng and Chou, Huang-Cheng and Chen, Kuan-Yu and Sung, Hsin-Yen and Narayanan, Shrikanth and Lee, Hung-yi},
  booktitle = {Proc. Interspeech 2026},
  year = {2026},
}

← All publications