Back to feed
AR
arXiv CS.AI
7/28/2026
Codifying the Judge: Scalable Evaluation via Program Distillation

Codifying the Judge: Scalable Evaluation via Program Distillation

Short summary

PAJAMA replaces LLM-as-a-judge with program distillation, synthesizing a committee of programmatic judges that score candidates directly without per-sample API costs. A fallback mechanism escalates low-confidence cases to an LLM. Across five datasets and four model families, programmatic judges match a 13B LLM judge's performance while producing reward signals that outperform proprietary LLM labels at two orders of magnitude lower cost.

  • Introduces program distillation as a cost-effective alternative to LLM-as-a-judge
  • Programmatic judges match 13B LLM judge performance with transparent, editable scoring logic
  • Reward models distilled from program verdicts outperform proprietary LLM labels at 100x lower API cost

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more