Back to feed
arXiv cs.CL
arXiv cs.CL
7/31/2026
LayerRAG-Bench: A Cross-Layer Reliability Benchmark for Agentic Retrieval-Augmented Generation

LayerRAG-Bench: A Cross-Layer Reliability Benchmark for Agentic Retrieval-Augmented Generation

Short summary

LayerRAG-Bench introduces a cross-layer reliability benchmark for agentic RAG systems, testing 240 tasks across 8 enterprise domains with 9 fault scenarios and nearly 39K records from OpenAI, Anthropic, and Gemini models. Schema normalization fixes schema-drift but fails to recover stale evidence, missing tool output, denied permissions, or wrong-session context. The key finding: reliability interventions should be evaluated layer-specifically, not treated as universal fixes.

  • Benchmark covers 8 enterprise domains, 240 tasks, 38,880 records across 9 models
  • Schema normalization raises schema-drift success from 0 to 0.913 but doesn't fix other fault layers
  • Layer-specific evaluation principle: credit interventions for their target layer only

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more