arXiv cs.CL
6/17/2026

RepSelect: Robust LLM Unlearning via Representation Selectivity
Short summary
RepSelect proposes a method for making LLMs permanently forget specific knowledge (biohazards, abuse patterns) without degrading general capabilities by isolating forget-set-specific representations through gradient principal-component collapse. Tested on Llama, Qwen, Gemma, and DeepSeek models, it achieved 4-50x stronger unlearning than existing baselines while resisting fine-tuning and few-shot re-learning attacks.
- •New method (RepSelect) makes LLM forgetting robust to adversarial re-learning attacks
- •Collapses gradient principal components to isolate forget-specific representations
- •4-50x stronger results than 5 baselines across 4 model families
Generated with AI, which can make mistakes.
Is this a good recommendation for you?