flâneur — a map of the web's best reading

Modular Pretraining Enables Access Control

alignment.anthropic.com · 2,716 words · saved by 1 readers

Frontier AI models have knowledge that could be misused for nefarious purposes. To address this risk, we introduce Gradient Routed Auxiliary Modules (GRAM), a method for isolating dangerous knowledge to specific modules within a language model. These modules can be switched on or off to control what the model knows, making it possible to restrict or extend access to the most sensitive model capabilities based on user need and trust. In our experiments, we find evidence that a single model trained in this way can approximate multiple models, each trained with a different category of dangerous data filtered out, and this ability holds for models ranging from 50M to 5B parameters. This research is preliminary and has not been applied to production models at Anthropic. 📄 Paper, 💻 Code This work was done at AE Studio, in collaboration with Anthropic. One of the major threats from frontier AI models is the misuse of legitimately helpful knowledge for harmful tasks, such as creating biologi

Modular Pretraining Enables Access Control Alignment Science Blog Modular Pretraining Enables Access Control Ethan Roland¹*, Murat Cubuktepe¹*, Erick Martinez¹* July 8, 2026 Stijn Servaes¹, Keenan Pepper¹, Mike Vaiana¹, Diogo Schwerz de Lucena¹, Judd Rosenblatt¹ Addie Foote² Cem Anil³, Alex Cloud³ ¹ AE Studio; ² Independent; ³ Anthropic; *Equal contribution tl;dr Frontier AI models have knowledge that could be misused for nefarious purposes. To address this risk, we introduce Gradient Routed Auxiliary Modules (GRAM), a method for isolating dangerous knowledge to specific modules within a langu

Explore this link on the map →

related reading