Alignment Forum
7/7/2026

Data filtering works a lot worse than you would expect
Short summary
Training data filtering fails to remove most undesired LLM behaviors in OLMo-3 fine-tuning—LLM judges, probes, and gradient-based attribution methods perform no better than random deletion. Only refusal behavior is effectively filterable, suggesting most unwanted traits arise from the model adopting an assistant persona rather than from specific training documents. Implies controllability via data curation is limited for broad behaviors.
- •Data filtering ineffective for most behaviors (political bias, formatting, feeling-validation) despite using multiple attribution methods
- •Only refusal behavior is consistently filterable using probes and LLM judges
- •Behaviors likely elicited as part of assistant-mode shift rather than taught by targeted data points
Generated with AI, which can make mistakes.
Is this a good recommendation for you?


