arXiv cs.LG
7/1/2026

Why Do Few-Step Text Latents Fail When Image Latents Work? Non-Commitment at Sharp Categorical Readouts
Short summary
Research explains why deterministic few-step generation fails on text latents but succeeds on image latents: sharp categorical readouts in text decoders prevent boundary resolution geometrically, not due to training deficiency. The paper proves this mathematically and identifies two escapes—categorical commitment in autoregressive decoding and stochastic re-injection—that sidestep the deterministic bound.
- •Few-step deterministic generation fails on text because decoder sharpness prevents resolving discrete branch choices before sharp readouts
- •Text decoders amplify boundary-aligned perturbations 100x–10,000x more than image decoders (DABI > 10^5 vs ≈1)
- •Autoregressive and stochastic mechanisms escape the continuous deterministic limit, enabling better few-step performance
Generated with AI, which can make mistakes.
Is this a good recommendation for you?