Generating Medical Errors: GenAI and Erroneous Medical References
A new study finds that large language models used widely for medical assessments cannot back up claims. ISTOCK/ARMAND BURGER Large language models (LLMs) are infiltrating the medical field. One in 10 doctors already use ChatGPT in day-to-day work, and patients have taken to ChatGPT to diagnose themselves. The Today Show featured the story of a 4-year-old boy, Alex, whose chronic illness was diagnosed by ChatGPT after over a dozen doctors failed to do so. This rapid adoption to much fanfare is in spite of substantial uncertainties about the safety, effectiveness, and risk of generative AI (GenAI). U.S. Food and Drug Administration Commissioner Robert Califf has publicly stated that the agency is "struggling" to regulate GenAI. The reason is that GenAI sits in a gray area between two existing forms of technology. On one hand, sites like WebMD that strictly report known medical information from credible sources are not regulated by the FDA. On the other hand, medical devices that interpre
iStock/Armand Burger A new study finds that large language models used widely for medical assessments cannot back up claims. Large language models (LLMs) are infiltrating the medical field. One in 10 doctors already use ChatGPT in day-to-day work, and patients have taken to ChatGPT to diagnose themselves. The Today Show featured the story of a 4-year-old boy, Alex, whose chronic illness was diagnosed by ChatGPT after over a dozen doctors failed to do so. This rapid adoption to much fanfare is in spite of substantial uncertainties about the safety, effectiveness, and risk of generative AI (GenA
related reading
- Peter Lee and the Impact of GPT-4 + Large Language AI Models in Medicineerictopol.substack.com
- Introducing HealthBench | OpenAIopenai.com
- gpt-4.pdfcdn.openai.com
- Evaluation and mitigation of the limitations of large language models in clinical decision-making | Nature Medicinenature.com
- 2309.07430.pdfarxiv.org
- What We’ve Learned From A Year of Building with LLMs – Applied LLMsapplied-llms.org
- 2025.acl-long.896.pdfaclanthology.org
- When Patient Questions Are Answered With Higher Quality and Empathy by ChatGPT than Physicianserictopol.substack.com
- Mediuminflecthealth.medium.com
- Google’s AMIE Validates The Clinician Cockpit | Counsel Healthcounselhealth.com
- Chain-of-Verification Reduces Hallucination in Large Language Modelsarxiv.org
- OpenAI Sued Over ChatGPT’s ‘Dangerous’ Health Advicenytimes.com