Curie: Toward Rigorous and Automated Scientific Experimentation with AI Agents | alphaXiv
View recent discussion. Abstract: Scientific experimentation, a cornerstone of human progress, demands rigor in reliability, methodical control, and interpretability to yield meaningful results. Despite the growing capabilities of large language models (LLMs) in automating different aspects of the scientific process, automating rigorous experimentation remains a significant challenge. To address this gap, we propose Curie, an AI agent framework designed to embed rigor into the experimentation process through three key components: an intra-agent rigor module to enhance reliability, an inter-agent rigor module to maintain methodical control, and an experiment knowledge module to enhance interpretability. To evaluate Curie, we design a novel experimental benchmark composed of 46 questions across four computer science domains, derived from influential research papers, and widely adopted open-source projects. Compared to the strongest baseline tested, we achieve a 3.4$\times$ improvement in correctly answering experimental questions. Curie is open-sourced at this https URL
Submitted 26 Feb 2025 University of MichiganCISCO Systems PT Patrick Tser Jern KonJiachen Liu QD Qiuyi DingYiming Qiu ZY Zhenning Yang YH Yibo Huang JS Jayanth SrinivasaMyungjin Lee +2 moreShow less Abstract Scientific experimentation, a cornerstone of human progress, demands rigor in reliability, methodical control, and interpretability to yield meaningful results. Despite the growing capabilities of large language models (LLMs) in automating different aspects of the scientific process, automating rigorous experimentation remains a significant challenge. To address this gap,…
saved by
related reading
- When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Researchalphaxiv.org
- The Need for Verification in AI-Driven Scientific Discoveryalphaxiv.org
- The machines are fine. I'm worried about us.ergosphere.blog
- Automated Weak-to-Strong Researcheralignment.anthropic.com
- Tech | LILAlila.ai
- Meet the Humans Building AI Scientists - Asimov Pressasimov.press
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discoverysakana.ai
- Danger, AI Scientist, Danger - by Zvi Mowshowitzthezvi.substack.com
- CausaLab — Can LLM Agents Discover Causal Mechanisms by Experiment?dylanzsz.github.io
- Can AI automate computational reproducibility?normaltech.ai
- GitHub - harbor-framework/terminal-bench-science: Terminal-Bench-Science: Evaluating AI agents on research workflows across scientific domainsgithub.com
- Research Robots: When AIs Experiment on Ustheaidigest.org