Back to feed
arXiv cs.CL
arXiv cs.CL
8/5/2026
JudgeArena: A Unified Framework for Reproducible LLM-Judge Evaluation

JudgeArena: A Unified Framework for Reproducible LLM-Judge Evaluation

Short summary

JudgeArena is an open-source framework that unifies major LLM-judge benchmarks (AlpacaEval, Arena-Hard, MT-Bench, m-Arena-Hard) under a single interface with swappable judges and metadata logging. It enables systematic study of how benchmark, judge model, prompt, and inference backend affect conclusions about model quality. The framework ships tuned configurations for open models that match or outperform closed-model judges and can simulate LMArena Elo scores at low cost.

  • Unifies four major LLM-judge benchmarks under one open-source interface with swappable judges
  • Tuned open-model judge configs match or outperform closed-model judges on human preference datasets
  • Can simulate LMArena Elo scores by combining existing human annotations with LLM-judge evaluations

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more