Introducing HealthBench | OpenAI
Improving human health will be one of the defining impacts of AGI. If developed and deployed effectively, large language models have the potential to expand access to health information, support clinicians in delivering high-quality care, and help people advocate for their health and that of their communities. To get there, we need to ensure models are useful and safe. Evaluations are essential to understanding how models perform in health settings. Significant efforts have already been made across academia and industry, yet many existing evaluations do not reflect realistic scenarios, lack rigorous validation against expert medical opinion, or leave no room for state-of-the-art models to improve. Today, we’re introducing HealthBench: a new benchmark designed to better measure capabilities of AI systems for health. Built in partnership with 262 physicians who have practiced in 60 countries, HealthBench includes 5,000 realistic health conversations, each with a custom physician-created
May 12, 2025 Publication Introducing HealthBench An evaluation for AI systems and human health. Read paper (opens in a new window) View code (opens in a new window) Loading… Share Improving human health will be one of the defining impacts of AGI. If developed and deployed effectively, large language models have the potential to expand access to health information, support clinicians in delivering high-quality care, and help people advocate for their health and that of their communities. To get there, we need to ensure models are useful and safe. Evaluations are essential to understanding how m
Explore this link on the map →saved by
related reading
- AI Triage Safety: HealthBench & Emergency Escalation | Counsel Healthcounselhealth.com
- GPT-4openai.com
- gpt-4.pdfcdn.openai.com
- Generating Medical Errors: GenAI and Erroneous Medical References | Stanford HAIhai.stanford.edu
- Peter Lee and the Impact of GPT-4 + Large Language AI Models in Medicineerictopol.substack.com
- Evaluation and mitigation of the limitations of large language models in clinical decision-making | Nature Medicinenature.com
- Hippocratic is building a large language model for healthcare | TechCrunchtechcrunch.com
- The Path to Medical Superintelligence | Microsoft AImicrosoft.ai
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Advancing Claude in healthcare and the life sciences \ Anthropicanthropic.com
- The Medical AI Manifesto - Introducing Sophont – Dr. Tanishq Abrahamtanishq.ai
- PostTrainBenchposttrainbench.com