arXiv cs.LG
7/21/2026

From Weights to Words: Expressing and Editing Preference Model Inferences in Natural Language
Short summary
"Weights to Words" is a method that automatically discovers domain-relevant preference dimensions described in natural language and paired with model vectors, enabling users to inspect and edit preference model inferences. Validated across moral dilemmas, movies, wines, and LLM responses with two pre-registered experiments (N=450, N=449), it shows that regularizing toward learned dimensions and incorporating user edits both improve prediction accuracy. Participants preferred its inferred profiles and endorsed its predictions as more accurate in head-to-head comparisons.
- •Method discovers natural-language preference dimensions paired with model vectors for interpretability
- •Two pre-registered experiments confirm accuracy gains from regularization and user edits
- •Demonstrated across four domains including moral dilemmas, movies, wines, and LLM responses
Generated with AI, which can make mistakes.
Is this a good recommendation for you?