publications

55 papers on fairness and social bias in speech models, speech emotion recognition, large audio-language model evaluation, speech quality assessment, and speech-music representation learning. Ordered by year. † denotes equal contribution.

2026

  1. SLT
    How Auditory Knowledge in LLM Backbones Shapes Audio Language Models: A Holistic Evaluation
    Ke-Han Lu, Szu-Wei Fu, Chao-Han Huck Yang, and 13 more authors
    In 2026 IEEE Spoken Language Technology Workshop (SLT), 2026
    Details
  2. SLT
    How Contrastive Decoding Enhances Large Audio Language Models?
    Tzu-Quan Lin, Wei-Ping Huang, Yi-Cheng Lin, and 1 more author
    In 2026 IEEE Spoken Language Technology Workshop (SLT), 2026
    Details
  3. SLT
    Listen, Critique, and Refine: RL-Based Self-Refinement for Instruction-Following Speech Synthesis
    Chee-En Yu, Yi-Cheng Lin, Sung-Feng Huang, and 4 more authors
    In 2026 IEEE Spoken Language Technology Workshop (SLT), 2026
    Details
  4. SLT
    Hearing Like Humans? Sound Symbolism and Perceptual Alignment in Speech Language Models
    Yun-Shao Tsai*, Chun-Wei Chen*, Chee-En Yu*, and 2 more authors
    In 2026 IEEE Spoken Language Technology Workshop (SLT), 2026
    Details
  5. SLT
    AudioICL-Bench: A Benchmark for Large Audio Language Model In-Context Learning
    Jia-Hung Chen, Yi-Cheng Lin, Kai-Wei Chang, and 2 more authors
    In 2026 IEEE Spoken Language Technology Workshop (SLT), 2026
    Details
  6. SLT
    AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition
    Yuqi Li*, Yi-Cheng Lin*, Xianglong Wang, and 5 more authors
    In 2026 IEEE Spoken Language Technology Workshop (SLT), 2026
    Details
  7. SLT
    VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Human Speech
    Yi-Cheng Lin, Yusuke Hirota, Sung-Feng Huang, and 1 more author
    In 2026 IEEE Spoken Language Technology Workshop (SLT), 2026
    Details
  8. TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling
    Hao-Hui Xie, Ho-Lam Chung, Yi-Cheng Lin, and 4 more authors
    In Proc. ISCSLP 2026, 2026
    Details
  9. ACL
    On the Fallacy of Global Token Perplexity in Spoken Language Model Evaluation
    Chan-Jan Hsu, Liang-Hsuan Tseng, Yi-Cheng Lin, and 5 more authors
    In Findings of the Association for Computational Linguistics: ACL 2026, 2026
    Details
  10. CSL
    State Space and Self-Attention Collaborative Network with Feature Aggregation for DOA Estimation
    Qi You, Qinghua Huang, and Yi-Cheng Lin
    Computer Speech & Language, 2026
    Details
  11. ACL
    Pseudo2Real: Task Arithmetic for Pseudo-Label Correction in Automatic Speech Recognition
    Yi-Cheng Lin*, Yu-Hsuan Li Liang*, Hsuan Su, and 4 more authors
    In Findings of the Association for Computational Linguistics: ACL 2026, 2026
    Details
  12. Bloodroot: When Watermarking Turns Poisonous For Stealthy Backdoor
    Kuan-Yu Chen, Yi-Cheng Lin, Jeng-Lin Li, and 1 more author
    In ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2026
    Details
  13. WaveSP-Net: Learnable Wavelet-Domain Sparse Prompt Tuning for Speech Deepfake Detection
    Xi Xuan, Xuechen Liu, Wenxin Zhang, and 3 more authors
    In ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2026
    Details
  14. TAU: A Benchmark for Cultural Sound Understanding Beyond Semantics
    Yi-Cheng Lin, Yu-Hua Chen, Jia-Kai Dong, and 12 more authors
    In ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2026
    Details
  15. MI-Fuse: Label Fusion for Unsupervised Domain Adaptation with Closed-Source Large-Audio Language Model
    Hsiao-Ying Huang*, Yi-Cheng Lin*, and Hung-yi Lee
    In ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2026
    Details
  16. How Does Instrumental Music Help SingFake Detection?
    Xuanjun Chen, Chia-Yu Hu, I-Ming Lin, and 8 more authors
    In ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2026
    Details
  17. Do You Hear What I Mean? Quantifying the Instruction-Perception Gap in Instruction-Guided Expressive Text-To-Speech Systems
    Yi-Cheng Lin, Huang-Cheng Chou, Tzu-Chieh Wei, and 2 more authors
    In ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2026
    Details
  18. DeSTA2.5-Audio: Toward General-Purpose Large Audio Language Model with Self-Generated Cross-Modal Alignment
    Ke-Han Lu, Zhehuai Chen, Szu-Wei Fu, and 25 more authors
    IEEE Transactions on Audio, Speech and Language Processing, 2026
    Details
  19. Speaker Identity in Non-Verbal Vocalizations: Conditional Distillation and Mixture of Experts Approach
    Tzu-Chieh Wei*, Yi-Cheng Lin*, Huang-Cheng Chou, and 4 more authors
    In Proc. Interspeech 2026, 2026
    Details
  20. The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation
    Yun-Shao Tsai, Yi-Cheng Lin, Huang-Cheng Chou, and 5 more authors
    In Proc. Interspeech 2026, 2026
    Details
  21. TaigiSpeech: A Low-Resource Real-World Speech Intent Dataset and Preliminary Results with Scalable Data Mining In-the-Wild
    Kai-Wei Chang, Yi-Cheng Lin, Huang-Cheng Chou, and 9 more authors
    In Proc. Interspeech 2026, 2026
    Details
  22. Latent-Mark: An Audio Watermark Robust to Neural Codec Compression
    Yen-Shan Chen, Shih-Yu Lai, Ying-Jung Tsou, and 5 more authors
    In Proc. Interspeech 2026, 2026
    Details
  23. Findings of the First Teaching Monster Challenge: A Benchmark of Pedagogical Content Knowledge in AI Agents
    Yi-Cheng Lin, Yu-Kai Guo, Szu-Chi Chen, and 15 more authors
    arXiv preprint arXiv:2608.08852, 2026
    Details
  24. Escaping the Procrustean Bed: Groupwise Orthogonal Connectors for Audio-Language Models
    Ho-Lam Chung, Ke-Han Lu, Yi-Cheng Lin, and 3 more authors
    arXiv preprint arXiv:2607.06014, 2026
    Details
  25. Toward Fair Speech Technologies: A Comprehensive Survey of Bias and Fairness in Speech AI
    Yi-Cheng Lin, Yun-Shao Tsai, Kuan-Yu Chen, and 6 more authors
    arXiv preprint arXiv:2605.01597, 2026
    Details
  26. MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation
    Szu-Chi Chen*, I-Ning Tsai*, Yi-Cheng Lin*, and 2 more authors
    In Proc. Interspeech 2026, 2026
    Details
  27. The Binding Effect: Analyzing How Multi-Dimensional Cues Form Gender Bias in Instruction TTS
    Kuan-Yu Chen, Yi-Cheng Lin, Po-Chung Hsieh, and 5 more authors
    In Proc. Interspeech 2026, 2026
    Details
  28. MOS-Bias: From Hidden Gender Bias to Gender-Aware Speech Quality Assessment
    Wenze Ren*, Yi-Cheng Lin*, Wen-Chin Huang, and 5 more authors
    In Proc. Interspeech 2026, 2026
    Details
  29. EduPanel: A Three-Agent LLM Judge for Teaching Videos – Reliability, Complementarity, and Human Trust Calibration
    Jia-Kai Dong, Yi-Cheng Lin, and Hung-yi Lee
    arXiv preprint arXiv:2607.18529, 2026
    Details

2025

  1. BreezyVoice: Adapting TTS for Taiwanese Mandarin with Enhanced Polyphone Disambiguation – Challenges and Insights
    Chan-Jan Hsu, Yi-Cheng Lin, Chia-Chun Lin, and 10 more authors
    In Proc. ISCSLP 2026, 2025
    Details
  2. Fake-Mamba: Real-Time Speech Deepfake Detection Using Bidirectional Mamba as Self-Attention’s Alternative
    Xi Xuan, Zimo Zhu, Wenxin Zhang, and 2 more authors
    In 2025 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2025
    Details
  3. ASTAR-NTU solution to AudioMOS Challenge 2025 Track1
    Fabian Ritter-Gutierrez*, Yi-Cheng Lin*, Jui-Chiang Wei*, and 3 more authors
    In 2025 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2025
    Details
  4. MMMOS: Multi-domain Multi-axis Audio Quality Assessment
    Yi-Cheng Lin*, Jia-Hung Chen*, and Hung-yi Lee
    In 2025 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2025
    Details
  5. HighRateMOS: Sampling-Rate Aware Modeling for Speech Quality Assessment
    Wenze Ren*, Yi-Cheng Lin*, Wen-Chin Huang, and 9 more authors
    In 2025 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2025
    Details
  6. A correlation-permutation approach for speech-music encoders model merging
    Fabian Ritter-Gutierrez, Yi-Cheng Lin, Jeremy H. M Wong, and 3 more authors
    In 2025 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2025
    Details
  7. Multi-Distillation from Speech and Music Representation Models
    Jui-Chiang Wei*, Yi-Cheng Lin*, Fabian Ritter-Gutierrez, and 1 more author
    In 2025 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2025
    Details
  8. CO-VADA: A Confidence-Oriented Voice Augmentation Debiasing Approach for Fair Speech Emotion Recognition
    Yun-Shao Tsai*, Yi-Cheng Lin*, Huang-Cheng Chou, and 1 more author
    In 2025 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2025
    Details
  9. EMO-Debias: Benchmarking Gender Debiasing Techniques in Multi-Label Speech Emotion Recognition
    Yi-Cheng Lin, Huang-Cheng Chou, Yu-Hsuan Li Liang, and 1 more author
    In 2025 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2025
    Details
  10. Creativity in LLM-based Multi-Agent Systems: A Survey
    Yi-Cheng Lin*, Kang-Chieh Chen*, Zhe-Yan Li*, and 5 more authors
    In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025
    Details
  11. Meta-PerSER: Few-Shot Listener Personalized Speech Emotion Recognition via Meta-learning
    Liang-Yeh Shen, Shi-Xin Fang, Yi-Cheng Lin, and 2 more authors
    In Proc. Interspeech 2025, 2025
    Details
  12. ToxicTone: A Mandarin Audio Dataset Annotated for Toxicity and Toxic Utterance Tonality
    Yu-Xiang Luo*, Yi-Cheng Lin*, Ming-To Chuang*, and 9 more authors
    In Proc. Interspeech 2025, 2025
    Details
  13. Mitigating Subgroup Disparities in Multi-Label Speech Emotion Recognition: A Pseudo-Labeling and Unsupervised Learning Approach
    Yi-Cheng Lin, Huang-Cheng Chou, and Hung-yi Lee
    In Proc. Interspeech 2025, 2025
    Details
  14. Distilling a speech and music encoder with task arithmetic
    Fabian Ritter-Gutierrez*, Yi-Cheng Lin*, Jui-Chiang Wei, and 4 more authors
    In Proc. Interspeech 2025, 2025
    Details
  15. Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks
    Chien-yu Huang, Wei-Chih Chen, Shu-wen Yang, and 77 more authors
    In The Thirteenth International Conference on Learning Representations (ICLR), 2025
    Details
  16. Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data Selection
    Hsi-Che Lin*, Yi-Cheng Lin*, Huang-Cheng Chou, and 1 more author
    In ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2025
    Details
  17. Leveraging Joint Spectral and Spatial Learning with MAMBA for Multichannel Speech Enhancement
    Wenze Ren, Haibin Wu, Yi-Cheng Lin, and 7 more authors
    In ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2025
    Details
  18. CantoASR: Prosody-Aware ASR-LALM Collaboration for Low-Resource Cantonese
    Dazhong Chen, Yi-Cheng Lin, Yuchen Huang, and 4 more authors
    arXiv preprint arXiv:2511.04139, 2025
    Details

2024

  1. SLT
    Efficient Training of Self-Supervised Speech Foundation Models on a Compute Budget
    Andy T. Liu, Yi-Cheng Lin, Haibin Wu, and 2 more authors
    In 2024 IEEE Spoken Language Technology Workshop (SLT), 2024
    Details
  2. SLT
    Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models
    Haibin Wu, Xuanjun Chen, Yi-Cheng Lin, and 13 more authors
    In 2024 IEEE Spoken Language Technology Workshop (SLT), 2024
    Details
  3. SLT
    Spoken Stereoset: On Evaluating Social Bias Toward Speaker in Speech Large Language Models
    Yi-Cheng Lin*, Wei-Chih Chen*, and Hung-yi Lee
    In 2024 IEEE Spoken Language Technology Workshop (SLT), 2024
    Details
  4. EMO-Codec: An In-Depth Look at Emotion Preservation capacity of Legacy and Neural Codec Models With Subjective and Objective Evaluations
    Wenze Ren*, Yi-Cheng Lin*, Huang-Cheng Chou, and 6 more authors
    In 2024 Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), 2024
    Details
  5. SLT
    Listen and Speak Fairly: A Study on Semantic Gender Bias in Speech Integrated Large Language Models
    Yi-Cheng Lin, Tzu-Quan Lin, Chih-Kai Yang, and 4 more authors
    In 2024 IEEE Spoken Language Technology Workshop (SLT), 2024
    Details
  6. Emo-bias: A Large Scale Evaluation of Social Bias on Speech Emotion Recognition
    Yi-Cheng Lin, Haibin Wu, Huang-Cheng Chou, and 2 more authors
    In Proc. Interspeech 2024, 2024
    Details
  7. On the social bias of speech self-supervised models
    Yi-Cheng Lin, Tzu-Quan Lin, Hsi-Che Lin, and 2 more authors
    In Proc. Interspeech 2024, 2024
    Details
  8. ACL
    Codec-SUPERB: An In-Depth Analysis of Sound Codec Models
    Haibin Wu, Ho-Lam Chung, Yi-Cheng Lin, and 7 more authors
    In Findings of the Association for Computational Linguistics: ACL 2024, 2024
    Details
  9. Building a Taiwanese Mandarin Spoken Language Model: A First Attempt
    Chih-Kai Yang, Yu-Kuan Fu, Chen-An Li, and 18 more authors
    arXiv preprint arXiv:2411.07111, 2024
    Details
  10. Towards audio language modeling – an overview
    Haibin Wu, Xuanjun Chen, Yi-Cheng Lin, and 4 more authors
    arXiv preprint arXiv:2402.13236, 2024
    Details