arXiv cs.CL
7/9/2026

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs
Short summary
A gradient-based method extracts word-level temporal alignment from any differentiable ASR model by computing per-frame saliency from teacher-forced token log-probability gradients and decoding boundaries via dynamic programming. It requires no training, no model modification, and works across CTC, transducer, AED, and speech LLM families. Evaluated on 16 models, it yields usable alignment everywhere—slightly behind strong native aligners but superior where native alignment is weak, such as streaming models.
- •Gradient-based alignment works for any differentiable ASR model including speech LLMs
- •No training or model modification needed; requires one backward pass per token
- •Evaluated on 16 models across 4 families on TIMIT and Buckeye datasets
Generated with AI, which can make mistakes.
Is this a good recommendation for you?
