Back to feed
AR
arXiv CS.AI
6/30/2026
The Two Genie Game: Adoption and Welfare in Audit-Grounded AI Governance

The Two Genie Game: Adoption and Welfare in Audit-Grounded AI Governance

Short summary

Applies evolutionary game theory (Moran-Fermi dynamics) to analyze when harm-minimizing AI agents outcompete approval-seeking ones in competitive markets, deriving critical adoption thresholds and irreversibility conditions. Proves that self-audited agents cannot prevent community harm without perfect alignment between agent audits and community values, with results formally verified in Lean 4. Warns that successful adoption creates an absorbing trap where policies lock in deferred harm, even under alignment.

  • Applies game theory to model competition between harm-minimizing and approval-seeking AI agents in resource-constrained markets
  • Derives critical thresholds where adoption becomes irreversible, with fixation depending on community feedback characteristics
  • Demonstrates self-auditing is insufficient for harm prevention; alignment with community values is necessary but not sufficient

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more