arXiv cs.LG
6/26/2026

Necessary but Not Sufficient: Temperature Control and Reproducibility in LLM-as-Judge Safety Evaluations
Short summary
Research reveals temperature control is necessary but insufficient for reproducible LLM-as-judge evaluations. Testing 690 API calls across two providers and three model tiers found borderline verdicts flip between pass/fail across identical runs (up to 50% disagreement). Claude Opus 4.7/4.8 deprecated temperature entirely, negating the primary mitigation. Recommendation: treat grader disagreement as a first-class health metric alongside scores to surface evaluation noise.
- •Temperature control alone is insufficient for reproducible LLM-as-judge evaluations
- •Borderline verdicts flip pass/fail across runs even with forced greedy decoding (up to 50% disagreement)
- •Evaluation harnesses should track grader disagreement as a first-class health metric
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



