Alignment Forum
7/9/2026
Modular Pretraining Enables Access Control
Short summary
Anthropic researchers introduce Gradient Routed Auxiliary Modules (GRAM), a method for isolating dangerous knowledge to specific modules in language models that can be toggled on or off. A single model trained with GRAM approximates the performance of multiple data-filtered models without the cost of training them separately, tested from 50M to 5B parameters. This enables fine-grained capability access control to restrict sensitive knowledge while maintaining general performance.
- •GRAM isolates dangerous model knowledge to switchable modules for capability access control
- •Single model approximates performance of multiple data-filtered models without separate training runs
- •Tested at scale (50M-5B parameters) with isolated virology, cybersecurity, and nuclear physics knowledge
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



