flâneur — a map of the web's best reading

benchmarks.bio — Agentic AI benchmarks on messy, real-world biological data

benchmarks.bio · 76 words · saved by 1 readers

Can frontier AI agents reason about real, messy biological data? SpatialBench (159 evals, 5 spatial transcriptomics platforms) is live, with deterministic graders that verify the key biological result. scBench for single-cell RNA-seq coming soon. By LatchBio.

Model-level refusal is a single safeguard within a larger system. DNA synthesis screening, institutional biosafety review, controlled access to reagents and equipment, and the practical difficulty of physical work all remain in place around it. Real misuse is unlikely to come from a single prompt; like legitimate research, it would demand sustained, multi-step effort informed by laboratory feedback. We measure refusal because it is a meaningful and newly important layer—not because it is the only one.

Explore this link on the map →

saved by

related reading