arXiv cs.LG
7/7/2026

Out-of-Distribution Generalization of Risk Aversion in Language Models
Short summary
Researchers introduce RiskAverseOOD, a benchmark testing whether risk aversion trained at low stakes generalizes to astronomical stakes—critical for AI alignment. Testing Qwen3, Gemma, and Llama models, they show risk aversion generalizes across 98 orders of magnitude, though consistency remains insufficient for deployment as a reliable safety failsafe.
- •New benchmark (RiskAverseOOD) measures risk aversion generalization across 98 orders of magnitude in stake size
- •Risk aversion generalizes with 70% cooperation rate (SFT), 52% (DPO), 39% (activation steering) across model families
- •Results promising for alignment research but not yet consistent enough for production-grade AI safety guarantees
Generated with AI, which can make mistakes.
Is this a good recommendation for you?
