Evidence / original sources

A claim should lead somewhere.

Follow the references behind Health Benchmark. Each entry explains its role, links to the original publication, and identifies the guides that use it.

Content updated September 28, 2026. Inclusion does not imply endorsement.

01

Rahul K. Arora and colleagues · OpenAI

HealthBench: Evaluating Large Language Models Towards Improved Human Health ↗

Original benchmark construction, coverage, scoring and historical evaluation variability.

Version: arXiv v1 · 2025-05-13

Evidence locator: Sections 2–5 and 8; Tables 2, 3 and 7; Appendix D

arxiv.org