benchmarks.bio — Agentic AI benchmarks on messy, real-world biological data
Can frontier AI agents reason about real, messy biological data? SpatialBench (159 evals, 5 spatial transcriptomics platforms) is live, with deterministic graders that verify the key biological result. scBench for single-cell RNA-seq coming soon. By LatchBio.
Model-level refusal is a single safeguard within a larger system. DNA synthesis screening, institutional biosafety review, controlled access to reagents and equipment, and the practical difficulty of physical work all remain in place around it. Real misuse is unlikely to come from a single prompt; like legitimate research, it would demand sustained, multi-step effort informed by laboratory feedback. We measure refusal because it is a meaningful and newly important layer—not because it is the only one.
Explore this link on the map →saved by
related reading
- AI Benchmark Leaderboards & Model Evals | BenchmarkListbenchmarklist.com
- Cookbookcookbook.openai.com
- AlphaProteo generates novel proteins for biology and health research — Google DeepMinddeepmind.google
- Elicit: AI for scientific researchelicit.org
- Elicit: AI for scientific researchelicit.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Parsed | Custom, interpretable AI systems that continuously learnparsed.com
- There's An AI For That® — The front page of AItheresanaiforthat.com
- Paving the way for agents in biology \ Anthropicanthropic.com
- AI Agent Benchmark for Real-World Professional Workflowsagents-last-exam.org
- ML Contestsmlcontests.com
- GitHub - karpathy/autoresearch: AI agents running research on single-GPU nanochat training automatically · GitHubgithub.com