benchmarks.bio — Agentic AI benchmarks on messy, real-world biological data
Can frontier AI agents reason about real, messy biological data? SpatialBench (159 evals, 5 spatial transcriptomics platforms) is live, with deterministic graders that verify the key biological result. scBench for single-cell RNA-seq coming soon. By LatchBio.
Agentic benchmarks on messy, real-world tasks Each problem includes a snapshot of real experimental data taken immediately prior to a target decision or analysis step, a description of the task through a high-level scientific lens, and a deterministic grader (e.g., Jaccard similarity of sets) that evaluates recovery of the key biological result in a verifiable manner. The benchmark is designed to test durable biological reasoning rather than method-specific implementation details and require empirical interaction with the data. Metric Balanced score — Trial-weighted harmonic mean of…
saved by
related reading
- BenchmarkList: Track the Frontier of AI Capabilitiesbenchmarklist.com
- PostTrainBenchposttrainbench.com
- Quantifying infrastructure noise in agentic coding evals \ Anthropicanthropic.com
- Paving the way for agents in biology \ Anthropicanthropic.com
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- GitHub - harbor-framework/terminal-bench-science: Terminal-Bench-Science: Evaluating AI agents on research workflows across scientific domainsgithub.com
- Giovanni D'Antoniogiovannidantonio.com
- f316275b44ee2de533102913828a8107-Paper-Datasets_and_Benchmarks_Track.pdfproceedings.neurips.cc
- Coding Agents Are Changing the Biosecurity Risk Landscape | GovAIgovernance.ai
- Center for Responsible, Decentralized Intelligence at Berkeleyrdi.berkeley.edu
- Import AIjack-clark.net