Back to feed
arXiv cs.CL
arXiv cs.CL
7/10/2026
Hallucination Self-Play: Bootstrapping Reinforced Detector via Evolved Generator

Hallucination Self-Play: Bootstrapping Reinforced Detector via Evolved Generator

Short summary

Researchers introduce Hallucination Self-Play (HSP), a novel framework for training detectors to identify false claims in LLM outputs. The system pairs a detector trained on human-labeled data with an evolved generator that produces increasingly subtle hallucinations via reinforcement learning from AI feedback. Experiments on RAGTruth show small LLMs can match or exceed larger models in hallucination detection without external supervision.

  • Novel framework uses paired detector (human-labeled) and evolved generator (RLAIF) for improved hallucination detection
  • Small models achieve parity with larger LLMs in detection accuracy without external supervision
  • Code available; validated on RAGTruth benchmark across two model families

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more