flâneur — a map of the web's best reading

Six things to keep in mind while reading biology ML papers

owlposting.com · 1,177 words · saved by 1 readers

This is a small compilation of things I have picked up as ‘things to watch out for’ while reading biology machine learning papers over the last few years. I’ll use ‘I’ for the entirety of this. However, this post is co-written with Nathan C. Frey, a scientist at Prescient Design who similarly writes on Substack, check him out! Established benchmarks are rarely reflective of the real world. Naively accepting strong results on benchmarks (e.g MoleculeNet, FLIP,) is a bad idea. Keep in mind that benchmarks in this field are incredibly challenging to create; the distribution shifts created when moving across datasets are massive. Most of the used benchmarks are a concession to standardization, not something that people actually agree on! Excellent performance on one dataset often doesn't translate to another, even within the same problem domain. Folks at inductive.bio made a very compelling explanation of how this affects small molecule datasets, and, unlike me, have offered a potential fi

Misc Five things to keep in mind while reading biology ML papers 1.1k words, 6 minutes reading time Abhishaike Mahajan and Nathan C. Frey Jun 24, 2024 27 2 2 Share This is a small compilation of things I have picked up as ‘ things to watch out for ’ while reading biology machine learning papers over the last few years. I’ll use ‘I’ for the entirety of this. However, this post is co-written with Nathan C. Frey , a scientist at Prescient Design who similarly writes on Substack , check him out! Established benchmarks are rarely reflective of the real world. Naively accepting strong results on ben

Explore this link on the map →

saved by

related reading