✳flâneur — a map of the web's best reading
Preventing model exfiltration with upload limits
redwoodresearch.substack.com · 4,772 words · saved by 1 readers
Unlike most files you might want to secure, model weights are extremely big. This might make them much easier to secure.
Preventing model exfiltration with upload limits Unlike most files you might want to secure, model weights are extremely big. This might make them much easier to secure. Ryan Greenblatt May 08, 2024 3 Share At some point in the future, AI developers will need to ensure that when they train sufficiently capable models, the weights of these models do not leave the developer’s control. Ensuring that weights are not exfiltrated seems crucial for preventing threat models related to both misalignment and misuse. The challenge of defending model weights has previously been discussed in a RAND report
Explore this link on the map →related reading
- My picture of the present in AI — LessWronglesswrong.com
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- Self-exfiltration is a key dangerous capabilityaligned.substack.com
- A basic systems architecture for AI agents that do autonomous research — LessWronglesswrong.com
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- What I learned this week - Can distillation be stopped, Mythos and the cybersecurity equilibrium, Pipeline RLdwarkesh.com
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- arxiv.org/pdf/2505.24832arxiv.org
- Interpretability Infrastructure at Frontier Scale: Harvesting Activations from a Trillion-Parameter Modelgoodfire.ai
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- Why I'm Not a Security Doomer - by Miles Brundagemilesbrundage.substack.com
- What will GPT-2030 look like? — AI Alignment Forumalignmentforum.org