AR
arXiv CS.AI
6/30/2026

Data and Evaluation Closed-Loop for Model Capability Enhancement
Short summary
Researchers propose a 'capability slice' framework that closes the gap between evaluation findings and data improvements in LLM training. By decomposing benchmark failures into specific task types and solving operations, engineers can make targeted data interventions. Two case studies show the method can either fix training bugs without changing data or boost math reasoning from 6.67% to 26.67% via targeted sampling.
- •Framework uses 'capability slices' (groups of evaluation samples sharing background, task, operation, and constraint) to map benchmark failures to specific data fixes
- •Case study 1: Fixed a training bug (masked EOS loss) recovering BBH to 66.44% without data changes; Case study 2: Boosted AIME Pass@128 from 6.67%/0.00% to 26.67% via math-targeted sampling
- •Makes evaluation-to-data inference routine, auditable, and experimentally validated rather than intuitive
Generated with AI, which can make mistakes.
Is this a good recommendation for you?

