A Safe Path to Open Weights - Thinking Machines Lab
thinkingmachines.ai · 2,309 words · saved by 9 readers
Strategic openness can strengthen AI safety and support broader access as defenses and safety science mature.
Abstract: Safe open-weight models are public goods, as they put AI development and safety work in many hands and make training choices inspectable. Open models also carry real misuse risks, and release is irreversible. To release safely, we must consider both the model and the ecosystem it enters. For the model, we conduct robust safety testing and research whether dangerous capabilities can be decoupled from general intelligence. For the ecosystem, we invest in its readiness through staged releases, supporting defenders, and collaboration with safety researchers. By iteratively choosing the m
saved by
- Elizabeth Qiu
- Brandon Wang
- Yixiong Hao
- Lydia Nottingham
- Yushan Li
- Vincent Cheng
- Florent Tavernier
- Kaustubh Kislay
- Arunim Agarwal
related reading
- Announcing Safety Research Grantsthinkingmachines.ai
- Inkling: Our Open-Weights Model - Thinking Machines Labthinkingmachines.ai
- Our position on open-weights models \ Anthropicanthropic.com
- [2608.07514] Open Technical Problems in Open-Weight AI Model Risk Managementarxiv.org
- Why open-weight models without guardrails are a AI safety risk : NPRnpr.org
- AI in 2025: gestalt — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- AI safety is not a model propertysubstack.com
- Open-Weights-and-American-AI-Leadership.pdfimages.nvidia.com
- GLM-5.2 Risk Evaluation Report – SaferAIsafer-ai.org
- Common Elements of Frontier AI Safety Policiesmetr.org
- The Myth of unsafe Open Source AIflorianbrand.com