Back to feed
arXiv cs.CL
arXiv cs.CL
6/17/2026
RepSelect: Robust LLM Unlearning via Representation Selectivity

RepSelect: Robust LLM Unlearning via Representation Selectivity

Short summary

RepSelect proposes a method for making LLMs permanently forget specific knowledge (biohazards, abuse patterns) without degrading general capabilities by isolating forget-set-specific representations through gradient principal-component collapse. Tested on Llama, Qwen, Gemma, and DeepSeek models, it achieved 4-50x stronger unlearning than existing baselines while resisting fine-tuning and few-shot re-learning attacks.

  • New method (RepSelect) makes LLM forgetting robust to adversarial re-learning attacks
  • Collapses gradient principal components to isolate forget-specific representations
  • 4-50x stronger results than 5 baselines across 4 model families

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more