AR
arXiv CS.AI
7/31/2026

Position: Evaluation Scores Are Perishable Knowledge Claims
Short summary
This position paper argues that LLM evaluation scores are perishable epistemic claims that should carry metadata for formality tier, scope, and expiration date. It identifies 'trust inflation' when averaging heterogeneous signals and proposes weakest-link aggregation as a conservative alternative. On the HELM leaderboard, the top-five models ranked by mean score versus weakest-link are completely disjoint, demonstrating that aggregation method materially changes conclusions.
- •Evaluation scores should be treated as epistemic claims with formality, scope, and validity windows
- •Mean aggregation causes trust inflation; weakest-link aggregation is the conservative alternative
- •Top-5 HELM models differ completely between mean-score and weakest-link rankings
Generated with AI, which can make mistakes.
Is this a good recommendation for you?