GLM-5.2 Risk Evaluation Report – SaferAI
SaferAI's independent evaluation of GLM-5.2, the first in Europe, tests Zhipu AI's open-weight flagship across the four systemic risk areas in the EU Code of Practice and finds frontier-level capability on cyber and biology benchmarks without the safeguards frontier developers apply.
GLM-5.2 Risk Evaluation Report – SaferAI We use cookies to make your experience on this website better. Agree Disagree GLM-5.2 Risk Evaluation Report Publication date August 2, 2026 authors Chinmayi Dixit*, Jacob Davies*, Jasmine Li, Ben Snodin, Jack Kengott, Henry Papadatos share Abstract SaferAI's independent evaluation of GLM-5.2, the first in Europe, tests Zhipu AI's open-weight flagship across the four systemic risk areas in the EU Code of Practice and finds frontier-level capability on cyber and biology benchmarks without the safeguards frontier developers apply. read full paper Ex
Explore this link on the map →saved by
related reading
- AI in 2025: gestalt — LessWronglesswrong.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- A Safe Path to Open Weights - Thinking Machines Labthinkingmachines.ai
- gpt-4.pdfcdn.openai.com
- Claude Fable 5 & Claude Mythos 5 — AI System Cardsmalob.github.io
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Summary of METR's predeployment evaluation of GPT-5.6 Solmetr.org
- GLM-5.2 is the step change for open agentsinterconnects.ai
- [2403.13793] Evaluating Frontier Models for Dangerous Capabilitiesarxiv.org
- Evaluating DeepSeek v4 Pro for Frontier Risks · Neo Researchneoresearch.ai
- Security incident disclosure — July 2026huggingface.co
- Incident Report: unsanctioned agent behaviour during cyber testing | AISI Workaisi.gov.uk