AR
arXiv CS.AI
7/28/2026

Codifying the Judge: Scalable Evaluation via Program Distillation
Short summary
PAJAMA replaces LLM-as-a-judge with program distillation, synthesizing a committee of programmatic judges that score candidates directly without per-sample API costs. A fallback mechanism escalates low-confidence cases to an LLM. Across five datasets and four model families, programmatic judges match a 13B LLM judge's performance while producing reward signals that outperform proprietary LLM labels at two orders of magnitude lower cost.
- •Introduces program distillation as a cost-effective alternative to LLM-as-a-judge
- •Programmatic judges match 13B LLM judge performance with transparent, editable scoring logic
- •Reward models distilled from program verdicts outperform proprietary LLM labels at 100x lower API cost
Generated with AI, which can make mistakes.
Is this a good recommendation for you?


