Introducing HealthBench | OpenAI
Improving human health will be one of the defining impacts of AGI. If developed and deployed effectively, large language models have the potential to expand access to health information, support clinicians in delivering high-quality care, and help people advocate for their health and that of their communities. To get there, we need to ensure models are useful and safe. Evaluations are essential to understanding how models perform in health settings. Significant efforts have already been made across academia and industry, yet many existing evaluations do not reflect realistic scenarios, lack rigorous validation against expert medical opinion, or leave no room for state-of-the-art models to improve. Today, we’re introducing HealthBench: a new benchmark designed to better measure capabilities of AI systems for health. Built in partnership with 262 physicians who have practiced in 60 countries, HealthBench includes 5,000 realistic health conversations, each with a custom physician-created
May 12, 2025 Publication Introducing HealthBench An evaluation for AI systems and human health. Read paper (opens in a new window) View code (opens in a new window) Loading… Share Improving human health will be one of the defining impacts of AGI. If developed and deployed effectively, large language models have the potential to expand access to health information, support clinicians in delivering high-quality care, and help people advocate for their health and that of their communities. To get there, we need to ensure models are useful and safe. Evaluations are essential to understanding how m
saved by
related reading
- AI Triage Safety: HealthBench & Emergency Escalation | Counsel Healthcounselhealth.com
- gpt-4.pdfcdn.openai.com
- Generating Medical Errors: GenAI and Erroneous Medical References | Stanford HAIhai.stanford.edu
- Peter Lee and the Impact of GPT-4 + Large Language AI Models in Medicineerictopol.substack.com
- PostTrainBenchposttrainbench.com
- 2309.07430.pdfarxiv.org
- Evaluation and mitigation of the limitations of large language models in clinical decision-making | Nature Medicinenature.com
- The Path to Medical Superintelligence | Microsoft AImicrosoft.ai
- Hippocratic is building a large language model for healthcare | TechCrunchtechcrunch.com
- Advancing Claude in healthcare and the life sciences \ Anthropicanthropic.com
- GitHub - open-compass/opencompass: OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.github.com
- The Medical AI Manifesto - Introducing Sophont – Dr. Tanishq Abrahamtanishq.ai