Generating Medical Errors: GenAI and Erroneous Medical References
A new study finds that large language models used widely for medical assessments cannot back up claims. ISTOCK/ARMAND BURGER Large language models (LLMs) are infiltrating the medical field. One in 10 doctors already use ChatGPT in day-to-day work, and patients have taken to ChatGPT to diagnose themselves. The Today Show featured the story of a 4-year-old boy, Alex, whose chronic illness was diagnosed by ChatGPT after over a dozen doctors failed to do so. This rapid adoption to much fanfare is in spite of substantial uncertainties about the safety, effectiveness, and risk of generative AI (GenAI). U.S. Food and Drug Administration Commissioner Robert Califf has publicly stated that the agency is "struggling" to regulate GenAI. The reason is that GenAI sits in a gray area between two existing forms of technology. On one hand, sites like WebMD that strictly report known medical information from credible sources are not regulated by the FDA. On the other hand, medical devices that interpre
iStock/Armand Burger A new study finds that large language models used widely for medical assessments cannot back up claims. Large language models (LLMs) are infiltrating the medical field. One in 10 doctors already use ChatGPT in day-to-day work, and patients have taken to ChatGPT to diagnose themselves. The Today Show featured the story of a 4-year-old boy, Alex, whose chronic illness was diagnosed by ChatGPT after over a dozen doctors failed to do so. This rapid adoption to much fanfare is in spite of substantial uncertainties about the safety, effectiveness, and risk of generative AI (GenA
Explore this link on the map →related reading
- Peter Lee and the Impact of GPT-4 + Large Language AI Models in Medicineerictopol.substack.com
- Introducing HealthBench | OpenAIopenai.com
- gpt-4.pdfcdn.openai.com
- Evaluation and mitigation of the limitations of large language models in clinical decision-making | Nature Medicinenature.com
- 2025.acl-long.896.pdfaclanthology.org
- Mediuminflecthealth.medium.com
- Google’s AMIE Validates The Clinician Cockpit | Counsel Healthcounselhealth.com
- When Patient Questions Are Answered With Higher Quality and Empathy by ChatGPT than Physicianserictopol.substack.com
- Medical AI Advances, Chatbots Work the Drive-Thru, and moredeeplearning.ai
- Google AI has better bedside manner than human doctors — and makes better diagnosesnature.com
- The Path to Medical Superintelligence | Microsoft AImicrosoft.ai
- AI on Trial: Legal Models Hallucinate in 1 out of 6 (or More) Benchmarking Queries | Stanford HAIhai.stanford.edu