arXiv cs.CL
7/14/2026

Workload-Driven Optimization for On-Device Real-Time Subtitle Translation
Short summary
This report optimizes on-device English-to-Traditional-Chinese subtitle translation for Taiwan under strict latency and privacy constraints. Replacing a 151k-token vocabulary with a 64k subtitle-domain tokenizer yields 1.63x speedup on Apple M2 Metal. The resulting LocalSubs model achieves 59.2% win rate against Google Translate on short subtitles, though performance declines on longer cues.
- •64k subtitle-domain tokenizer replaces 151k vocabulary for 1.63x speedup on Apple M2 Metal
- •LocalSubs achieves 59.2% tie-excluded win rate vs Google Translate on short subtitles
- •Vocabulary projection becomes dominant decode cost after GGUF quantization of Transformer blocks
Generated with AI, which can make mistakes.
Is this a good recommendation for you?
