Yi-Cheng Lin

PhD student, Graduate Institute of Communication Engineering, National Taiwan University. Speech Processing and Machine Learning Lab, advised by Hung-yi Lee. Student Researcher at Google Cloud AI.

prof_pic.jpg

Taipei, Taiwan

f12942075 [at] ntu.edu.tw

Hi! I’m Yi-Cheng Lin (林羿成), a Ph.D. student at National Taiwan University, advised by Prof. Hung-yi Lee in the Speech Processing and Machine Learning Lab. I’m also a Student Researcher at Google Cloud AI.

My research is on large audio-language models and fairness in speech processing, covering social bias in speech models, speech emotion recognition, and speech quality assessment. I have published 30+ papers at Interspeech, ICASSP, ASRU, SLT, ACL, EMNLP and IEEE TASLP with over 700 citations, and received Best Paper Runner-Up in the Responsible Speech Foundation Models Special Session at Interspeech 2024 and the NSTC Graduate Research Fellowship in 2025.

news

Sep 09, 2026 Seven papers accepted to SLT 2026: Auditory Knowledge in LLM Backbones, Contrastive Decoding in Audio LLMs, Listen, Critique, and Refine, Hearing Like Humans?, AudioICL-Bench, AMRD, VIBE.
Aug 31, 2026 🥇 Our team placed first in Track 3 of the VoiceMOS Challenge 2026 (predicting speaker and accent similarity for codec-based speech synthesis).
Aug 17, 2026 Two papers accepted to ISCSLP 2026: BreezyVoice, TW-Sound580K.
Aug 09, 2026 Preprint released: Teaching Monster Challenge (arXiv:2608.08852).
Jul 20, 2026 Started as a Student Researcher at Google Cloud AI.
Jul 20, 2026 Two preprints released: EduPanel, Escaping the Procrustean Bed.
Jul 07, 2026 Presented two Findings papers at ACL 2026 in San Diego 🇺🇸, and received a Foundation for the Advancement of Outstanding Scholarship Travel Grant.
May 03, 2026 Six papers at ICASSP 2026 in Barcelona 🇪🇸: Do You Hear What I Mean?, MI-Fuse, TAU, Instrumental Music in SingFake Detection, WaveSP-Net, Bloodroot.
May 02, 2026 Preprint released: Toward Fair Speech Technologies (arXiv:2605.01597).
Apr 30, 2026 FAS-Conformer accepted to Computer Speech & Language — an efficient Conformer with feature aggregation for direction-of-arrival estimation. My second journal article.
Apr 01, 2026 Seven papers accepted to INTERSPEECH 2026: Speaker Identity in Non-Verbal Vocalizations, The False Resonance, TaigiSpeech, Latent-Mark, MoVE, The Binding Effect, MOS-Bias.
Feb 01, 2026 Received an ICASSP 2026 Travel Grant.
Jan 15, 2026 DeSTA2.5-Audio published in IEEE Transactions on Audio, Speech and Language Processing — my first journal article.
Jan 10, 2026 Two papers accepted to Findings of ACL 2026: Pseudo2Real, Global Token Perplexity.
Dec 06, 2025 Eight papers at ASRU 2025 in Honolulu 🇺🇸: EMO-Debias, CO-VADA, Multi-Distillation, Correlation-Permutation Model Merging, MMMOS, HighRateMOS, ASTAR-NTU at AudioMOS Challenge 2025, Fake-Mamba.
Nov 06, 2025 Preprint released: CantoASR (arXiv:2511.04139).
Nov 05, 2025 Presented Creativity in LLM-based Multi-Agent Systems at EMNLP 2025 in Suzhou 🇨🇳, and received a Foundation for the Advancement of Outstanding Scholarship Travel Grant.
Aug 17, 2025 Four papers at INTERSPEECH 2025 in Rotterdam 🇳🇱: Mitigating Subgroup Disparities, Distilling Speech and Music Encoders, Meta-PerSER, ToxicTone. Received an INTERSPEECH 2025 Travel Grant.
Jul 31, 2025 🥇 Our team placed first in Track 1 (MOS prediction for text-to-music systems) and Track 3 (MOS prediction for speech at high sampling rates) of the ASRU AudioMOS Challenge 2025.
Apr 24, 2025 Dynamic-SUPERB Phase-2 accepted to ICLR 2025 — a collaboratively expanding benchmark with 180 tasks for spoken language models.
Apr 06, 2025 Two papers at ICASSP 2025 in Hyderabad 🇮🇳: Speech Emotion Recognition in Under-Resourced Languages, MAMBA for Multichannel Speech Enhancement.
Feb 01, 2025 Transferred to the PhD program in Communication Engineering at National Taiwan University, and received the NSTC Graduate Research Fellowship.
Dec 03, 2024 EMO-Codec presented at APSIPA ASC 2024, examining how well legacy and neural codecs preserve emotion.
Dec 02, 2024 Four papers at SLT 2024 in Macao 🇲🇴: Listen and Speak Fairly, Spoken Stereoset, Codec-SUPERB @ SLT 2024, Speech Foundation Models on a Compute Budget.
Nov 11, 2024 Preprint released: Building a Taiwanese Mandarin Spoken Language Model (arXiv:2411.07111).
Sep 05, 2024 On the social bias of speech self-supervised models received Best Paper Runner-Up in the Responsible Speech Foundation Models Special Session at INTERSPEECH 2024 in Kos Island 🇬🇷. Emo-bias was presented at the same conference.
Sep 01, 2024 Received the Garmin Scholarship Award.
Aug 11, 2024 Codec-SUPERB accepted to Findings of ACL 2024 — an in-depth analysis of sound codec models.
Feb 20, 2024 Preprint released: Towards audio language modeling (arXiv:2402.13236).
Sep 01, 2023 Started the M.S. program in Communication Engineering at National Taiwan University, joining the Speech Processing and Machine Learning Lab under Hung-yi Lee.
Jun 30, 2023 Graduated with a B.S. in Electrical Engineering from National Taiwan University, with a double major in the Trans-disciplinary Bachelor Program and a minor in Physics.
Jan 01, 2023 Received the Irving T. Ho Memorial Scholarship.
Sep 01, 2022 Started as a Research & Development Intern at Microsoft Taiwan, upgrading the geolocation prediction system from rule-based to ML-based.

selected publications

  1. Do You Hear What I Mean? Quantifying the Instruction-Perception Gap in Instruction-Guided Expressive Text-To-Speech Systems
    Yi-Cheng Lin, Huang-Cheng Chou, Tzu-Chieh Wei, and 2 more authors
    In ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2026
    Details
  2. DeSTA2.5-Audio: Toward General-Purpose Large Audio Language Model with Self-Generated Cross-Modal Alignment
    Ke-Han Lu, Zhehuai Chen, Szu-Wei Fu, and 25 more authors
    IEEE Transactions on Audio, Speech and Language Processing, 2026
    Details
  3. EMO-Debias: Benchmarking Gender Debiasing Techniques in Multi-Label Speech Emotion Recognition
    Yi-Cheng Lin, Huang-Cheng Chou, Yu-Hsuan Li Liang, and 1 more author
    In 2025 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2025
    Details
  4. Creativity in LLM-based Multi-Agent Systems: A Survey
    Yi-Cheng Lin*, Kang-Chieh Chen*, Zhe-Yan Li*, and 5 more authors
    In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025
    Details
  5. SLT
    Spoken Stereoset: On Evaluating Social Bias Toward Speaker in Speech Large Language Models
    Yi-Cheng Lin*, Wei-Chih Chen*, and Hung-yi Lee
    In 2024 IEEE Spoken Language Technology Workshop (SLT), 2024
    Details
  6. SLT
    Listen and Speak Fairly: A Study on Semantic Gender Bias in Speech Integrated Large Language Models
    Yi-Cheng Lin, Tzu-Quan Lin, Chih-Kai Yang, and 4 more authors
    In 2024 IEEE Spoken Language Technology Workshop (SLT), 2024
    Details
  7. On the social bias of speech self-supervised models
    Yi-Cheng Lin, Tzu-Quan Lin, Hsi-Che Lin, and 2 more authors
    In Proc. Interspeech 2024, 2024
    Details
  8. ACL
    Codec-SUPERB: An In-Depth Analysis of Sound Codec Models
    Haibin Wu, Ho-Lam Chung, Yi-Cheng Lin, and 7 more authors
    In Findings of the Association for Computational Linguistics: ACL 2024, 2024
    Details
  9. Toward Fair Speech Technologies: A Comprehensive Survey of Bias and Fairness in Speech AI
    Yi-Cheng Lin, Yun-Shao Tsai, Kuan-Yu Chen, and 6 more authors
    arXiv preprint arXiv:2605.01597, 2026
    Details
  10. Towards audio language modeling – an overview
    Haibin Wu, Xuanjun Chen, Yi-Cheng Lin, and 4 more authors
    arXiv preprint arXiv:2402.13236, 2024
    Details