arXiv cs.CL
7/14/2026

Efficiently Adapting Spoken Language Models for the Singaporean Context
Short summary
Researchers adapted an open-source spoken language model for Singapore's Home Team across five speech tasks in four official languages using LoRA fine-tuning and a multi-task objective. They built HTD-multilingual-QA, a 504K-sample multilingual dataset, and the resulting 5B-parameter HT-Moonstone matches or outperforms SLMs up to 7x its size while retaining original speech QA ability. The approach combines surrogate text-QA data to prevent catastrophic forgetting with CoBa reweighting adapted for speech.
- •5B SLM matches 7x larger models on Singaporean multilingual speech tasks
- •LoRA fine-tuning plus surrogate text-QA prevents catastrophic forgetting
- •New 504K-sample multilingual QA dataset released
Generated with AI, which can make mistakes.
Is this a good recommendation for you?
