arXiv cs.CL
8/5/2026

JudgeArena: A Unified Framework for Reproducible LLM-Judge Evaluation
Short summary
JudgeArena is an open-source framework that unifies major LLM-judge benchmarks (AlpacaEval, Arena-Hard, MT-Bench, m-Arena-Hard) under a single interface with swappable judges and metadata logging. It enables systematic study of how benchmark, judge model, prompt, and inference backend affect conclusions about model quality. The framework ships tuned configurations for open models that match or outperform closed-model judges and can simulate LMArena Elo scores at low cost.
- •Unifies four major LLM-judge benchmarks under one open-source interface with swappable judges
- •Tuned open-model judge configs match or outperform closed-model judges on human preference datasets
- •Can simulate LMArena Elo scores by combining existing human annotations with LLM-judge evaluations
Generated with AI, which can make mistakes.
Is this a good recommendation for you?
