Building a Taiwanese Mandarin Spoken Language Model: A First Attempt
arXiv preprint arXiv:2411.07111 , 2024
Abstract
This technical report presents our initial attempt to build a spoken large language model (LLM) for Taiwanese Mandarin, specifically tailored to enable real-time, speech-to-speech interaction in multi-turn conversations. Our end-to-end model incorporates a decoder-only transformer architecture and aims to achieve seamless interaction while preserving the conversational flow, including full-duplex capabilities allowing simultaneous speaking and listening. The paper also details the training process, including data preparation with synthesized dialogues and adjustments for real-time interaction. We also developed a platform to evaluate conversational fluency and response coherence in multi-turn dialogues. We hope the release of the report can contribute to the future development of spoken LLMs in Taiwanese Mandarin.
BibTeX
@article{yang2024building,
title = {Building a Taiwanese Mandarin Spoken Language Model: A First Attempt},
author = {Yang, Chih-Kai and Fu, Yu-Kuan and Li, Chen-An and Lin, Yi-Cheng and Lin, Yu-Xiang and Chen, Wei-Chih and Chung, Ho Lam and Kuan, Chun-Yi and Huang, Wei-Ping and Lu, Ke-Han and Lin, Tzu-Quan and Wang, Hsiu-Hsuan and Hu, En-Pei and Hsu, Chan-Jan and Tseng, Liang-Hsuan and Chiu, I-Hsiang and Sanga, Ulin and Chen, Xuanjun and Hsu, Po-chun and Yang, Shu-wen and Lee, Hung-yi},
journal = {arXiv preprint arXiv:2411.07111},
year = {2024},
}