AR
arXiv CS.AI
7/2/2026

A Contextual-Bandit Oversight Game with Two-Sided Informational Asymmetry
Short summary
Researchers model human oversight of AI agents where both sides hold private information: humans don't know action quality, AI doesn't know reward functions. Using contextual-bandit theory, they characterize the gap between optimal oversight and myopic behavior, identifying where AI can propose harm undetected. The work connects game theory with practical safeguards for autonomous systems.
- •Analyzes human oversight with two-sided private information (asymmetric knowledge)
- •Identifies 'avoidable harm' gap where AI can propose harmful actions undetected
- •Applies contextual-bandit and game theory to improve AI safety mechanisms
Generated with AI, which can make mistakes.
Is this a good recommendation for you?

