arXiv cs.CL
7/27/2026

A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models
Short summary
This paper introduces a consensus-based evaluation framework where a panel of diverse LLMs ranks anonymized candidate responses to measure relative preference rather than absolute correctness. Using five state-of-the-art models across programming, knowledge, safety, reasoning, and math, responses are aggregated into a Relative Intelligence Index (RII). The authors find consistent inter-model preference patterns but caution these reflect model alignment, not human judgment or objective correctness.
- •Proposes consensus-based LLM evaluation using inter-model agreement as a proxy for response quality
- •Introduces Relative Intelligence Index (RII) from structured peer voting across five SOTA models
- •Results show consistent preference patterns but are not directly aligned with human evaluation
Generated with AI, which can make mistakes.
Is this a good recommendation for you?