Evaluation and mitigation of the limitations of large language models in clinical decision-making | Nature Medicine
Thank you for visiting nature.com. You are using a browser version with limited support for CSS. To obtain the best experience, we recommend you use a more up to date browser (or turn off compatibility mode in Internet Explorer). In the meantime, to ensure continued support, we are displaying the site without styles and JavaScript. Advertisement Nature Medicine (2024)Cite this article 11k Accesses 1 Citations 169 Altmetric Metrics details Clinical decision-making is one of the most impactful parts of a physician’s responsibilities and stands to benefit greatly from artificial intelligence solutions and large language models (LLMs) in particular. However, while LLMs have achieved excellent performance on medical licensing exams, these tests fail to assess many skills necessary for deployment in a realistic clinical decision-making environment, including gathering information, adhering to guidelines, and integrating into clinical workflows. Here we hav
Download PDF Subjects Diagnosis Health care economics Translational research Abstract Clinical decision-making is one of the most impactful parts of a physician’s responsibilities and stands to benefit greatly from artificial intelligence solutions and large language models (LLMs) in particular. However, while LLMs have achieved excellent performance on medical licensing exams, these tests fail to assess many skills necessary for deployment in a realistic clinical decision-making environment, including gathering information, adhering to guidelines, and integrating into clinical workflows. Here
Explore this link on the map →related reading
- Google’s AMIE Validates The Clinician Cockpit | Counsel Healthcounselhealth.com
- The Path to Medical Superintelligence | Microsoft AImicrosoft.ai
- [2401.05654] Towards Conversational Diagnostic AIarxiv.org
- Introducing HealthBench | OpenAIopenai.com
- LLM Evaluation doesn't need to be complicatedphilschmid.de
- The bitter lesson of LLM evalsparsed.com
- Generating Medical Errors: GenAI and Erroneous Medical References | Stanford HAIhai.stanford.edu
- [2604.15597] LLMs Corrupt Your Documents When You Delegatearxiv.org
- [2602.16703] Measuring Mid-2025 LLM-Assistance on Novice Performance in Biologyarxiv.org
- [2406.02061] Alice in Wonderland: Simple Tasks Showing Complete Reasoning Breakdown in State-Of-the-Art Large Language Modelsarxiv.org
- Clinical Text Summarization: Adapting Large Language Models Can Outperform Human Experts - PMCncbi.nlm.nih.gov
- Things we learned about LLMs in 2024simonwillison.net