Back to feed
Alignment Forum
Alignment Forum
7/9/2026
Modular Pretraining Enables Access Control

Modular Pretraining Enables Access Control

Short summary

Anthropic researchers introduce Gradient Routed Auxiliary Modules (GRAM), a method for isolating dangerous knowledge to specific modules in language models that can be toggled on or off. A single model trained with GRAM approximates the performance of multiple data-filtered models without the cost of training them separately, tested from 50M to 5B parameters. This enables fine-grained capability access control to restrict sensitive knowledge while maintaining general performance.

  • GRAM isolates dangerous model knowledge to switchable modules for capability access control
  • Single model approximates performance of multiple data-filtered models without separate training runs
  • Tested at scale (50M-5B parameters) with isolated virology, cybersecurity, and nuclear physics knowledge

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more