Back to feed
arXiv cs.CL
arXiv cs.CL
7/9/2026
Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs

Short summary

A gradient-based method extracts word-level temporal alignment from any differentiable ASR model by computing per-frame saliency from teacher-forced token log-probability gradients and decoding boundaries via dynamic programming. It requires no training, no model modification, and works across CTC, transducer, AED, and speech LLM families. Evaluated on 16 models, it yields usable alignment everywhere—slightly behind strong native aligners but superior where native alignment is weak, such as streaming models.

  • Gradient-based alignment works for any differentiable ASR model including speech LLMs
  • No training or model modification needed; requires one backward pass per token
  • Evaluated on 16 models across 4 families on TIMIT and Buckeye datasets

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more