flâneur

benchmarks.bio — Agentic AI benchmarks on messy, real-world biological data

benchmarks.bio · 964 words · saved by 1 readers

Can frontier AI agents reason about real, messy biological data? SpatialBench (159 evals, 5 spatial transcriptomics platforms) is live, with deterministic graders that verify the key biological result. scBench for single-cell RNA-seq coming soon. By LatchBio.

Agentic benchmarks on messy, real-world tasks Each problem includes a snapshot of real experimental data taken immediately prior to a target decision or analysis step, a description of the task through a high-level scientific lens, and a deterministic grader (e.g., Jaccard similarity of sets) that evaluates recovery of the key biological result in a verifiable manner. The benchmark is designed to test durable biological reasoning rather than method-specific implementation details and require empirical interaction with the data. Metric Balanced score — Trial-weighted harmonic mean of…

saved by

related reading