CAIS AI Dashboard
dashboard.safe.ai · 158 words · saved by 6 readers
Evaluating frontier AI systems on capability and safety benchmarks
Evaluating frontier AI systems on capability and safety benchmarks Text Capabilities Index Size: StandardMini Model| GPT-6 Astra 63.953.674.2 Fable 5.1 55.654.656.6 Opus 5 52.251.053.3 Gemini 3.1 Pro 45.845.945.8 Grok 4.6 43.539.747.3 Kimi K3 40.541.140.0 Muse Spark 1.1 37.249.424.9 GLM 5.3 35.639.132.0 Grok 4.5 32.339.325.4 DeepSeek 4 Pro 26.432.420.3 Vision Capabilities Index Size: StandardMini Model| GPT-6 Astra 83.082.392.793.177.469.7 Fable 5.1 68.970.384.577.967.744.2 Opus 5 68.070.884.375.865.143.9 Gemini 3.1 Pro 62.174.284.166.153.632.4 Kimi K3…
saved by
related reading
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Goodfire AIgoodfire.ai
- GLM-5.2 Risk Evaluation Report – SaferAIsafer-ai.org
- Conceptual Reasoning Indexconceptualreasoning.ai
- Datacurve | The data engine for frontier AIdatacurve.ai
- Anthropic/values-in-the-wild · Datasets at Hugging Facehuggingface.co
- There's An AI For That® — The front page of AItheresanaiforthat.com
- RSI Simulator | Paradigmparadigm.xyz
- The Shape of AI | UX Patterns for Artificial Intelligence Designshapeof.ai
- Parsed | Custom, interpretable AI systems that continuously learnparsed.com
- Intelligenceintelligence.ai
- GPT-6 Astra System Card - OpenAI Deployment Safety Hubdeploymentsafety.openai.com