To Err Is Human: Systematic Quantification of Errors in Published AI Papers via LLM Analysis
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. How many mistakes do published AI papers contain? Peer-reviewed publications form the foundation upon which new research and knowledge are built. Errors that persist in the literature can propagate unnoticed, creating confusion in follow-up studies and complicating reproducibility. The accelerating pace of research and the increasing demands on the peer-review system make such mistakes harder to detect and avoid. To address this, we developed a Paper Correctness Checker based on GPT-5 to systematically identify mistakes in papers previously published at top AI conferences and journals. Our analysis focuses on objective mistakes—e.g., errors in formulas, derivations, calculations, figures, and tables—that have a clearly verifiable ground truth. We intenti
To Err Is Human: Systematic Quantification of Errors in Published AI Papers via LLM Analysis Federico Bianchi † 1 , Yongchan Kwon † 1 , Zachary Izzo † 2 , Linjun Zhang 3 , James Zou 1,4 † Equal Contribution. 1 Together AI, 2 NEC Labs America, 3 Rutgers University, 4 Stanford University. Abstract How many mistakes do published AI papers contain? Peer-reviewed publications form the foundation upon which new research and knowledge are built. Errors that persist in the literature can propagate unnoticed, creating confusion in follow-up studies and complicating reproducibility. The accelerating pac
related reading
- [2606.28277] Towards Automating Scientific Review with Google's Paper Assistant Toolarxiv.org
- When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Researchalphaxiv.org
- AI-Generated Papers in the NeurIPS 2026 Position Paper Trackblog.neurips.cc
- How much science is verifiable? Results from replicating ICML 2026 oral paperssai.science
- Refine - AI-Powered Research Assistantrefine.ink
- The machines are fine. I'm worried about us.ergosphere.blog
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Mathematics in the Library of Babel - Daniel Littdaniellitt.com
- Is AI Reasoning Right for the Wrong Reasons? | Quanta Magazinequantamagazine.org
- Scientists should use AI as a tool, not an oraclenormaltech.ai
- Predicting Empirical AI Research Outcomes with Language Modelsarxiv.org