Back to feed
arXiv cs.CL
arXiv cs.CL
7/16/2026
Do LLMs Need Architectural Changes for Simultaneous Speech Translation? A Prefix-to-Prefix Data Driven Approach

Do LLMs Need Architectural Changes for Simultaneous Speech Translation? A Prefix-to-Prefix Data Driven Approach

Short summary

This paper proposes CSSEL-P2P, a data-driven approach to simultaneous speech translation that avoids architectural changes to decoder-only LLMs. It uses fixed-length chunks with cumulative streaming decoding and teacher-labeled prefix-to-prefix targets for fine-tuning. In conversational speech evaluation, CSSEL-P2P improves streaming quality by +1.54 COMETKiwi over the baseline at comparable latency (+0.15s Average Lagging), demonstrating effective SimulST without brittle read/write policies or model architecture modifications.

  • CSSEL-P2P achieves +1.54 COMETKiwi improvement over streaming baseline at comparable latency
  • Data-driven prefix-to-prefix supervision replaces architectural changes for simultaneous speech translation
  • Fixed-length chunked decoding with rewind-based committed prefix handles ambiguous segmentation boundaries

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more