Evaluating DeepSeek v4 Pro for Frontier Risks · Neo Research
An independent safety evaluation of DSv4 Pro across CBRN, cyber, harmful manipulation, and loss-of-control. Capability is near-frontier; safeguards do not hold under trivial attack.
← All research Contents Executive summary 1 Cyber 2 CBRN 3 Harmful manipulation 4 Loss of control 5 Adversarial robustness 6 Tracking evaluation awareness 7 Testing whether judge origin affects safety scoring 8 Conclusion ← All research Summary Executive summary # We ran an independent safety evaluation of DeepSeek v4 Pro (DSv4 Pro), the open-weight preview released on 24 April 2026. We worked only from the public weights and API, with no developer cooperation and no privileged model access. The evaluation covers the four EU AI Act Code of Practice systemic-risk areas (CBRN, cyber, harmful man
saved by
related reading
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- gpt-4.pdfcdn.openai.com
- Claude 4 System Cardwww-cdn.anthropic.com
- GLM-5.2 Risk Evaluation Report – SaferAIsafer-ai.org
- Claude Sonnet 4.5 System Cardassets.anthropic.com
- Claude Opus 4.5: Model Card, Alignment and Safetythezvi.substack.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- CAIS AI Dashboarddashboard.safe.ai
- An alignment assessment of recent cybersecurity incidentsanthropic.com
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- [2403.13793] Evaluating Frontier Models for Dangerous Capabilitiesarxiv.org
- Agentic Misalignment in Summer 2026alignment.anthropic.com