VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Human Speech

SLT

Yi-Cheng Lin, Yusuke Hirota, Sung-Feng Huang, Hung-yi Lee

2026 IEEE Spoken Language Technology Workshop (SLT) , 2026

Abstract

Large Audio-Language Models (LALMs) are increasingly integrated into daily applications, yet their generative biases remain underexplored. Existing speech fairness benchmarks rely on synthetic speech and Multiple-Choice Questions (MCQs), both offering a fragmented view of fairness. We propose VIBE, a framework that evaluates generative bias through open-ended tasks such as personalized recommendations, using human-recorded speech. Unlike MCQs, our method allows stereotypical associations to manifest organically without predefined options, making it easily extensible to new tasks. Evaluating 12 state-of-the-art LALMs reveals systematic biases in realistic scenarios. Both gender and accent cues trigger statistically significant distributional shifts, and bias magnitude is strongly task-dependent.

BibTeX

@inproceedings{lin2026vibe,
  title = {VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Human Speech},
  author = {Lin, Yi-Cheng and Hirota, Yusuke and Huang, Sung-Feng and Lee, Hung-yi},
  booktitle = {2026 IEEE Spoken Language Technology Workshop (SLT)},
  year = {2026},
}

← All publications