Toward Human-Aligned Judgement of Speech Emotion Similarity

arXiv

Yun-Shao Tsai, Yi-Cheng Lin, Chih-Kai Yang, Ho-Jung Cheng, Tsun-Yi Chang, Sheng-Wei Wu, Yi-Shan Chen, Hsiang-Chun Chang, Liang-Chieh Lee, Hung-yi Lee

arXiv preprint arXiv:2609.32504 , 2026

Abstract

Evaluating emotion preservation in expressive speech generation involves assessing how closely generated speech matches a reference in emotion. Human listening tests assess this similarity, but their cost motivates automatic measures aligned with human judgments. To support the development and evaluation of such measures, we introduce SES-Bench, a speech emotion similarity benchmark built from human comparisons of two candidate utterances against a shared reference. These comparisons record which candidate listeners find emotionally closer to the reference and the strength of their preference. Using these annotations, we train SES-Judge to score emotion similarity between two utterances. SES-Judge significantly outperforms embedding cosine similarity and prompted large audio-language models in preference accuracy and correlation with human ratings that capture both preference direction and strength.

BibTeX

@article{tsai2026human,
  title = {Toward Human-Aligned Judgement of Speech Emotion Similarity},
  author = {Tsai, Yun-Shao and Lin, Yi-Cheng and Yang, Chih-Kai and Cheng, Ho-Jung and Chang, Tsun-Yi and Wu, Sheng-Wei and Chen, Yi-Shan and Chang, Hsiang-Chun and Lee, Liang-Chieh and Lee, Hung-yi},
  journal = {arXiv preprint arXiv:2609.32504},
  year = {2026},
}

← All publications