arXiv cs.CL
7/14/2026

Silent Failures in Quantized LLM Reasoning: A Taxonomy-Based Analysis of Hollow Convergence and Failure Mode Shifts
Short summary
Post-training quantization can silently alter LLM reasoning even when accuracy is preserved, this study shows. Using a six-category failure taxonomy across 30,000 chain-of-thought outputs from five LLMs (3B-14B), the authors find Hollow Convergence—correct answers via incomplete reasoning—shifts significantly under NF4 for small models. Shortcut Collapse rises from 44% to 78% of wrong answers in LLaMA 3.2-3B under NF4, invisible to standard accuracy metrics.
- •Six-category failure taxonomy applied to 30K chain-of-thought outputs across 5 LLMs and 3 quantization precisions
- •Hollow Convergence (correct answers via flawed reasoning) shifts size-dependently under NF4, dropping for small models but stable at 12B+
- •Surface-level text features cannot reliably detect Hollow Convergence (best F1=0.53), making it a deployment-relevant risk
Generated with AI, which can make mistakes.
Is this a good recommendation for you?