Back to feed
arXiv cs.LG
arXiv cs.LG
6/26/2026
Necessary but Not Sufficient: Temperature Control and Reproducibility in LLM-as-Judge Safety Evaluations

Necessary but Not Sufficient: Temperature Control and Reproducibility in LLM-as-Judge Safety Evaluations

Short summary

Research reveals temperature control is necessary but insufficient for reproducible LLM-as-judge evaluations. Testing 690 API calls across two providers and three model tiers found borderline verdicts flip between pass/fail across identical runs (up to 50% disagreement). Claude Opus 4.7/4.8 deprecated temperature entirely, negating the primary mitigation. Recommendation: treat grader disagreement as a first-class health metric alongside scores to surface evaluation noise.

  • Temperature control alone is insufficient for reproducible LLM-as-judge evaluations
  • Borderline verdicts flip pass/fail across runs even with forced greedy decoding (up to 50% disagreement)
  • Evaluation harnesses should track grader disagreement as a first-class health metric

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more