Evaluating DeepSeek v4 Pro for Frontier Risks · Neo Research
An independent safety evaluation of DSv4 Pro across CBRN, cyber, harmful manipulation, and loss-of-control. Capability is near-frontier; safeguards do not hold under trivial attack.
← All research Contents Executive summary 1 Cyber 2 CBRN 3 Harmful manipulation 4 Loss of control 5 Adversarial robustness 6 Tracking evaluation awareness 7 Testing whether judge origin affects safety scoring 8 Conclusion ← All research Summary Executive summary # We ran an independent safety evaluation of DeepSeek v4 Pro (DSv4 Pro), the open-weight preview released on 24 April 2026. We worked only from the public weights and API, with no developer cooperation and no privileged model access. The evaluation covers the four EU AI Act Code of Practice systemic-risk areas (CBRN, cyber, harmful man
Explore this link on the map →saved by
related reading
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- gpt-4.pdfcdn.openai.com
- Claude 4 System Cardwww-cdn.anthropic.com
- Claude Sonnet 4.5 System Cardassets.anthropic.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- [2403.13793] Evaluating Frontier Models for Dangerous Capabilitiesarxiv.org
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- GLM-5.2 Risk Evaluation Report – SaferAIsafer-ai.org
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- 2312.06942arxiv.org
- Claude Opus 4.8: The System Card - by Zvi Mowshowitzthezvi.substack.com