Back to feed
arXiv cs.LG
arXiv cs.LG
7/21/2026
From Weights to Words: Expressing and Editing Preference Model Inferences in Natural Language

From Weights to Words: Expressing and Editing Preference Model Inferences in Natural Language

Short summary

"Weights to Words" is a method that automatically discovers domain-relevant preference dimensions described in natural language and paired with model vectors, enabling users to inspect and edit preference model inferences. Validated across moral dilemmas, movies, wines, and LLM responses with two pre-registered experiments (N=450, N=449), it shows that regularizing toward learned dimensions and incorporating user edits both improve prediction accuracy. Participants preferred its inferred profiles and endorsed its predictions as more accurate in head-to-head comparisons.

  • Method discovers natural-language preference dimensions paired with model vectors for interpretability
  • Two pre-registered experiments confirm accuracy gains from regularization and user edits
  • Demonstrated across four domains including moral dilemmas, movies, wines, and LLM responses

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more