AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition
2026 IEEE Spoken Language Technology Workshop (SLT) , 2026
Abstract
On-device speech emotion recognition (SER) is critical for real-time applications, yet large self-supervised models that excel at SER are too costly for edge devices. Multi-teacher knowledge distillation can compress them into a lightweight student, but two challenges remain: teacher reliability varies across batches, and logit-level distillation ignores inter-sample relational structure. We propose Adaptive Multi-teacher Relational Distillation (AMRD) to address both. A one-class SVM on each teacher’s logit similarity matrix assigns per-batch weights favoring more coherent teachers. A relational distillation loss aligns teacher and student similarity matrices, capturing structure that logit matching misses. On IEMOCAP and CREMA-D datasets across four student architectures, AMRD outperforms single-teacher distillation baselines in most settings, and ablations confirm both components yield complementary gains.
BibTeX
@inproceedings{li2026amrd,
title = {AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition},
author = {Li, Yuqi and Lin, Yi-Cheng and Wang, Xianglong and Yang, Kuo and Feng, Xiaoqin and Wang, Yixuan and Duan, Huiran and Tian, Yingli},
booktitle = {2026 IEEE Spoken Language Technology Workshop (SLT)},
year = {2026},
}