arXiv cs.CL
7/1/2026

When Calibration Rankings Reverse: Accuracy-Controlled Evaluation for Fair Comparison of LLMs
Short summary
Researchers show that standard calibration metrics for comparing LLMs are confounded by accuracy differences, leading to unfair rankings. They propose ACE (Accuracy-Controlled Evaluation) with three aligned evaluation views. Testing reveals many models ranked favorably by raw metrics rank differently when accuracy is controlled.
- •Standard global calibration metrics (ECE, Brier Score) unfairly rank LLMs when model accuracy differs
- •ACE framework proposes three accuracy-controlled evaluation views for fair cross-model comparison
- •Frequent ranking reversals occur—models favored by raw metrics often rank lower after accuracy control
Generated with AI, which can make mistakes.
Is this a good recommendation for you?