Training-Free Instruction TTS Gender Bias Calibration Using Model-Adaptive Steering

arXiv

Kuan-Yu Chen, Yi-Cheng Lin, Jeng-Lin Li, Jian-Jiun Ding

arXiv preprint arXiv:2610.09831 , 2026

Abstract

Instruction text-to-speech (ITTS) systems systematically encode acoustic gender skews from descriptive style prompts, such as occupations or personas, even when demographic attributes are left unspecified. Calibrating the implicit gender distribution of the synthesized voices, across heterogeneous architectures and without model retraining or overriding explicit user prompts, is an open problem. In this work, we propose model-adaptive steering, a training-free bias calibration method that steers post-encoder conditioning representations using a group deviation vector paired with a coarse-to-fine strength search on a development set. A deterministic lexical gate bypasses intervention whenever explicit gender keywords are detected, preserving intended prompt semantics. Evaluated on a 12,800-prompt held-out benchmark across four ITTS models (three architectures), the proposed method reduces aggregate calibration error from 11.5–28.1 to 0.8–5.9 percentage points (e.g., shifting Parler-TTS Mini from 78.1% and PromptTTS++ from 27.3% female to 50.8–52.4%). Output rates stay within 0.9 pp across anchor set sizes at fixed operating points, while UTMOS decreases by at most 0.08 and WER degrades by at most 3.3 pp.

BibTeX

@article{chen2026training,
  title = {Training-Free Instruction TTS Gender Bias Calibration Using Model-Adaptive Steering},
  author = {Chen, Kuan-Yu and Lin, Yi-Cheng and Li, Jeng-Lin and Ding, Jian-Jiun},
  journal = {arXiv preprint arXiv:2610.09831},
  year = {2026},
}

← All publications