arXiv cs.CL
7/10/2026

Hallucination Self-Play: Bootstrapping Reinforced Detector via Evolved Generator
Short summary
Researchers introduce Hallucination Self-Play (HSP), a novel framework for training detectors to identify false claims in LLM outputs. The system pairs a detector trained on human-labeled data with an evolved generator that produces increasingly subtle hallucinations via reinforcement learning from AI feedback. Experiments on RAGTruth show small LLMs can match or exceed larger models in hallucination detection without external supervision.
- •Novel framework uses paired detector (human-labeled) and evolved generator (RLAIF) for improved hallucination detection
- •Small models achieve parity with larger LLMs in detection accuracy without external supervision
- •Code available; validated on RAGTruth benchmark across two model families
Generated with AI, which can make mistakes.
Is this a good recommendation for you?