Back to feed
Alignment Forum
Alignment Forum
7/7/2026
Data filtering works a lot worse than you would expect

Data filtering works a lot worse than you would expect

Short summary

Training data filtering fails to remove most undesired LLM behaviors in OLMo-3 fine-tuning—LLM judges, probes, and gradient-based attribution methods perform no better than random deletion. Only refusal behavior is effectively filterable, suggesting most unwanted traits arise from the model adopting an assistant persona rather than from specific training documents. Implies controllability via data curation is limited for broad behaviors.

  • Data filtering ineffective for most behaviors (political bias, formatting, feeling-validation) despite using multiple attribution methods
  • Only refusal behavior is consistently filterable using probes and LLM judges
  • Behaviors likely elicited as part of assistant-mode shift rather than taught by targeted data points

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more