Back to feed
arXiv cs.CL
arXiv cs.CL
7/14/2026
Efficiently Adapting Spoken Language Models for the Singaporean Context

Efficiently Adapting Spoken Language Models for the Singaporean Context

Short summary

Researchers adapted an open-source spoken language model for Singapore's Home Team across five speech tasks in four official languages using LoRA fine-tuning and a multi-task objective. They built HTD-multilingual-QA, a 504K-sample multilingual dataset, and the resulting 5B-parameter HT-Moonstone matches or outperforms SLMs up to 7x its size while retaining original speech QA ability. The approach combines surrogate text-QA data to prevent catastrophic forgetting with CoBa reweighting adapted for speech.

  • 5B SLM matches 7x larger models on Singaporean multilingual speech tasks
  • LoRA fine-tuning plus surrogate text-QA prevents catastrophic forgetting
  • New 504K-sample multilingual QA dataset released

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more