flâneur — a map of the web's best reading

To Err Is Human: Systematic Quantification of Errors in Published AI Papers via LLM Analysis

arxiv.org · 4,411 words · saved by 1 readers

This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. How many mistakes do published AI papers contain? Peer-reviewed publications form the foundation upon which new research and knowledge are built. Errors that persist in the literature can propagate unnoticed, creating confusion in follow-up studies and complicating reproducibility. The accelerating pace of research and the increasing demands on the peer-review system make such mistakes harder to detect and avoid. To address this, we developed a Paper Correctness Checker based on GPT-5 to systematically identify mistakes in papers previously published at top AI conferences and journals. Our analysis focuses on objective mistakes—e.g., errors in formulas, derivations, calculations, figures, and tables—that have a clearly verifiable ground truth. We intenti

To Err Is Human: Systematic Quantification of Errors in Published AI Papers via LLM Analysis Federico Bianchi † 1 , Yongchan Kwon † 1 , Zachary Izzo † 2 , Linjun Zhang 3 , James Zou 1,4 † Equal Contribution. 1 Together AI, 2 NEC Labs America, 3 Rutgers University, 4 Stanford University. Abstract How many mistakes do published AI papers contain? Peer-reviewed publications form the foundation upon which new research and knowledge are built. Errors that persist in the literature can propagate unnoticed, creating confusion in follow-up studies and complicating reproducibility. The accelerating pac

Explore this link on the map →

related reading