Yi-Cheng Lin
PhD student, Graduate Institute of Communication Engineering, National Taiwan University. Speech Processing and Machine Learning Lab, advised by Hung-yi Lee. Student Researcher at Google Cloud AI.
Taipei, Taiwan
f12942075 [at] ntu.edu.tw
Hi! I’m Yi-Cheng Lin (林羿成), a Ph.D. student at National Taiwan University, advised by Prof. Hung-yi Lee in the Speech Processing and Machine Learning Lab. I’m also a Student Researcher at Google Cloud AI.
My research is on large audio-language models and fairness in speech processing, covering social bias in speech models, speech emotion recognition, and speech quality assessment. I have published 30+ papers at Interspeech, ICASSP, ASRU, SLT, ACL, EMNLP and IEEE TASLP with over 700 citations, and received Best Paper Runner-Up in the Responsible Speech Foundation Models Special Session at Interspeech 2024 and the NSTC Graduate Research Fellowship in 2025.
news
| Sep 09, 2026 | Seven papers accepted to SLT 2026: Auditory Knowledge in LLM Backbones, Contrastive Decoding in Audio LLMs, Listen, Critique, and Refine, Hearing Like Humans?, AudioICL-Bench, AMRD, VIBE. |
|---|---|
| Aug 31, 2026 | 🥇 Our team placed first in Track 3 of the VoiceMOS Challenge 2026 (predicting speaker and accent similarity for codec-based speech synthesis). |
| Aug 17, 2026 | Two papers accepted to ISCSLP 2026: BreezyVoice, TW-Sound580K. |
| Aug 09, 2026 | Preprint released: Teaching Monster Challenge (arXiv:2608.08852). |
| Jul 20, 2026 | Started as a Student Researcher at Google Cloud AI. |
| Jul 20, 2026 | Two preprints released: EduPanel, Escaping the Procrustean Bed. |
| Jul 07, 2026 | Presented two Findings papers at ACL 2026 in San Diego 🇺🇸, and received a Foundation for the Advancement of Outstanding Scholarship Travel Grant. |
| May 03, 2026 | Six papers at ICASSP 2026 in Barcelona 🇪🇸: Do You Hear What I Mean?, MI-Fuse, TAU, Instrumental Music in SingFake Detection, WaveSP-Net, Bloodroot. |
| May 02, 2026 | Preprint released: Toward Fair Speech Technologies (arXiv:2605.01597). |
| Apr 30, 2026 | FAS-Conformer accepted to Computer Speech & Language — an efficient Conformer with feature aggregation for direction-of-arrival estimation. My second journal article. |
| Apr 01, 2026 | Seven papers accepted to INTERSPEECH 2026: Speaker Identity in Non-Verbal Vocalizations, The False Resonance, TaigiSpeech, Latent-Mark, MoVE, The Binding Effect, MOS-Bias. |
| Feb 01, 2026 | Received an ICASSP 2026 Travel Grant. |
| Jan 15, 2026 | DeSTA2.5-Audio published in IEEE Transactions on Audio, Speech and Language Processing — my first journal article. |
| Jan 10, 2026 | Two papers accepted to Findings of ACL 2026: Pseudo2Real, Global Token Perplexity. |
| Dec 06, 2025 | Eight papers at ASRU 2025 in Honolulu 🇺🇸: EMO-Debias, CO-VADA, Multi-Distillation, Correlation-Permutation Model Merging, MMMOS, HighRateMOS, ASTAR-NTU at AudioMOS Challenge 2025, Fake-Mamba. |
| Nov 06, 2025 | Preprint released: CantoASR (arXiv:2511.04139). |
| Nov 05, 2025 | Presented Creativity in LLM-based Multi-Agent Systems at EMNLP 2025 in Suzhou 🇨🇳, and received a Foundation for the Advancement of Outstanding Scholarship Travel Grant. |
| Aug 17, 2025 | Four papers at INTERSPEECH 2025 in Rotterdam 🇳🇱: Mitigating Subgroup Disparities, Distilling Speech and Music Encoders, Meta-PerSER, ToxicTone. Received an INTERSPEECH 2025 Travel Grant. |
| Jul 31, 2025 | 🥇 Our team placed first in Track 1 (MOS prediction for text-to-music systems) and Track 3 (MOS prediction for speech at high sampling rates) of the ASRU AudioMOS Challenge 2025. |
| Apr 24, 2025 | Dynamic-SUPERB Phase-2 accepted to ICLR 2025 — a collaboratively expanding benchmark with 180 tasks for spoken language models. |
| Apr 06, 2025 | Two papers at ICASSP 2025 in Hyderabad 🇮🇳: Speech Emotion Recognition in Under-Resourced Languages, MAMBA for Multichannel Speech Enhancement. |
| Feb 01, 2025 | Transferred to the PhD program in Communication Engineering at National Taiwan University, and received the NSTC Graduate Research Fellowship. |
| Dec 03, 2024 | EMO-Codec presented at APSIPA ASC 2024, examining how well legacy and neural codecs preserve emotion. |
| Dec 02, 2024 | Four papers at SLT 2024 in Macao 🇲🇴: Listen and Speak Fairly, Spoken Stereoset, Codec-SUPERB @ SLT 2024, Speech Foundation Models on a Compute Budget. |
| Nov 11, 2024 | Preprint released: Building a Taiwanese Mandarin Spoken Language Model (arXiv:2411.07111). |
| Sep 05, 2024 | On the social bias of speech self-supervised models received Best Paper Runner-Up in the Responsible Speech Foundation Models Special Session at INTERSPEECH 2024 in Kos Island 🇬🇷. Emo-bias was presented at the same conference. |
| Sep 01, 2024 | Received the Garmin Scholarship Award. |
| Aug 11, 2024 | Codec-SUPERB accepted to Findings of ACL 2024 — an in-depth analysis of sound codec models. |
| Feb 20, 2024 | Preprint released: Towards audio language modeling (arXiv:2402.13236). |
| Sep 01, 2023 | Started the M.S. program in Communication Engineering at National Taiwan University, joining the Speech Processing and Machine Learning Lab under Hung-yi Lee. |
| Jun 30, 2023 | Graduated with a B.S. in Electrical Engineering from National Taiwan University, with a double major in the Trans-disciplinary Bachelor Program and a minor in Physics. |
| Jan 01, 2023 | Received the Irving T. Ho Memorial Scholarship. |
| Sep 01, 2022 | Started as a Research & Development Intern at Microsoft Taiwan, upgrading the geolocation prediction system from rule-based to ML-based. |