Back to feed
arXiv cs.LG
arXiv cs.LG
7/7/2026
Out-of-Distribution Generalization of Risk Aversion in Language Models

Out-of-Distribution Generalization of Risk Aversion in Language Models

Short summary

Researchers introduce RiskAverseOOD, a benchmark testing whether risk aversion trained at low stakes generalizes to astronomical stakes—critical for AI alignment. Testing Qwen3, Gemma, and Llama models, they show risk aversion generalizes across 98 orders of magnitude, though consistency remains insufficient for deployment as a reliable safety failsafe.

  • New benchmark (RiskAverseOOD) measures risk aversion generalization across 98 orders of magnitude in stake size
  • Risk aversion generalizes with 70% cooperation rate (SFT), 52% (DPO), 39% (activation steering) across model families
  • Results promising for alignment research but not yet consistent enough for production-grade AI safety guarantees

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more