Rahul K. Arora and colleagues · OpenAI
HealthBench: Evaluating Large Language Models Towards Improved Human Health ↗
Original benchmark construction, coverage, scoring and historical evaluation variability.
Version: arXiv v1 · 2025-05-13
Evidence locator: Sections 2–5 and 8; Tables 2, 3 and 7; Appendix D
arxiv.org