Back to feed
arXiv cs.CL
arXiv cs.CL
8/3/2026
The Formalism Trap: LLM-as-a-Judge Evaluators Confuse Procedural Structure with Semantic Truth

The Formalism Trap: LLM-as-a-Judge Evaluators Confuse Procedural Structure with Semantic Truth

Original: The Formalism Trap: Are LLM-as-a-Judge Evaluators Blinded by Consensus Mimicry under Social Load?

Short summary

This paper introduces the Agentic Formalism Trap, showing that LLM-as-a-Judge systems conflate procedural structure with semantic truth under adversarial load. Analyzing 22,500 trajectories across three domains, the authors identify systematic hallucination maneuvers and propose an Evaluative Dissonance Index to quantify evaluator capture. A zero-shot cross-domain transfer proves the vulnerability is domain-agnostic, suggesting architecture-specific vigilance filters are necessary for reliable closed-loop evaluation.

  • LLM judges confuse structural proceduralism with semantic truth under adversarial conditions
  • A logistic meta-valuator isolates syntactic triggers of evaluator capture with ROC-AUC 0.878
  • The vulnerability transfers across domains (mean ROC-AUC 0.748), requiring architecture-specific filters

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more