To Err Is Human: Systematic Quantification of Errors in Published AI Papers via LLM Analysis
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. How many mistakes do published AI papers contain? Peer-reviewed publications form the foundation upon which new research and knowledge are built. Errors that persist in the literature can propagate unnoticed, creating confusion in follow-up studies and complicating reproducibility. The accelerating pace of research and the increasing demands on the peer-review system make such mistakes harder to detect and avoid. To address this, we developed a Paper Correctness Checker based on GPT-5 to systematically identify mistakes in papers previously published at top AI conferences and journals. Our analysis focuses on objective mistakes—e.g., errors in formulas, derivations, calculations, figures, and tables—that have a clearly verifiable ground truth. We intenti
To Err Is Human: Systematic Quantification of Errors in Published AI Papers via LLM Analysis Federico Bianchi † 1 , Yongchan Kwon † 1 , Zachary Izzo † 2 , Linjun Zhang 3 , James Zou 1,4 † Equal Contribution. 1 Together AI, 2 NEC Labs America, 3 Rutgers University, 4 Stanford University. Abstract How many mistakes do published AI papers contain? Peer-reviewed publications form the foundation upon which new research and knowledge are built. Errors that persist in the literature can propagate unnoticed, creating confusion in follow-up studies and complicating reproducibility. The accelerating pac
Explore this link on the map →related reading
- [2606.28277] Towards Automating Scientific Review with Google's Paper Assistant Toolarxiv.org
- The machines are fine. I'm worried about us.ergosphere.blog
- Refine - AI-Powered Research Assistantrefine.ink
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Mathematics in the Library of Babel - Daniel Littdaniellitt.com
- Danger, AI Scientist, Danger - by Zvi Mowshowitzthezvi.substack.com
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discoverysakana.ai
- [2603.26524] Mathematical methods and human thought in the age of AIarxiv.org
- Import AIjack-clark.net
- Templates for machine learning research papers | Neel Guhaneelguha.github.io
- The bitter lesson of LLM evalsparsed.com