arXiv cs.LG
7/1/2026

Revocable Learned State via Process Sidecars
Short summary
Researchers introduce 'process sidecars,' a novel mathematical technique for revoking specific learned information from language models after safety training has been applied. The innovation accounts for how downstream safety training reshapes learned feature directions, achieving second-order accuracy improvements over naive parameter subtraction. Experimental validation across three models demonstrates the sidecar method improves safety refusal behavior compared to standard task arithmetic.
- •Introduces 'process sidecars' for revocable model memory after safety training
- •Mathematically accounts for how safety training reshapes learned features
- •Outperforms naive task-arithmetic editing in refusal tests across three models
Generated with AI, which can make mistakes.
Is this a good recommendation for you?

