AR
arXiv CS.AI
7/31/2026

Probing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs. SFT Fine-Tuned Models
Short summary
This study investigates why RL-trained reasoning models outperform SFT counterparts on math tasks by examining internal representations. Linear probes reveal RL models have more structured, linearly separable hidden states, while ablation studies show RL models develop hierarchical layer importance versus SFT's uniform distribution. Token-count variability analysis suggests adaptive compute allocation depends on the full training pipeline rather than RL vs SFT alone.
- •RL models show more linearly separable and structured internal representations than SFT models
- •RL training creates hierarchical layer importance; SFT distributes importance uniformly
- •Token-allocation variability depends on overall training pipeline, not just RL vs SFT
Generated with AI, which can make mistakes.
Is this a good recommendation for you?