arXiv cs.CL
8/4/2026

Role Steering of Language Models for Social Simulations
Short summary
This paper introduces an activation-steering screening workflow for role-conditioned language model agents in social simulations. Applied to OLMo-3-7B-Instruct across 275 roles, role-specific steering directions achieve higher alignment scores (63.2 vs 41.1) than prior persona-vector approaches while preserving lexical diversity. The key practical finding: 38 of 275 roles decline with increased steering, so coefficients should be tuned per role rather than set uniformly.
- •Activation-steering workflow for screening role-conditioned agents before simulation deployment
- •Role-specific directions outperform persona-vector controls on alignment (63.2 vs 41.1) while preserving diversity
- •38 of 275 roles decline with stronger steering, requiring per-role coefficient tuning
- •Code and evaluation artifacts publicly available
Generated with AI, which can make mistakes.
Is this a good recommendation for you?