{"publication":"Health Benchmark","url":"https://healthbenchmark.ai","publisher":"Arcophos","updated":"2026-09-28","provenance":"Independent analytical publication. Benchmark creation and experimental results belong to their cited authors. Reported results are source-version snapshots, not new Arcophos runs or a live leaderboard.","benchmarks":[{"slug":"healthbench","name":"HealthBench","shortName":"HealthBench","version":"Original HealthBench; May 2025 paper v1","creators":"Rahul K. Arora and colleagues · OpenAI","paperDate":"2025-05-13","headline":"A rubric score is a balance of rewarded and penalized behavior.","summary":"HealthBench evaluates open-ended health conversations through physician-written rubrics. Its score can look like a percentage without being the percentage of answers that are fully correct. We make the calculation inspectable: signed points, positive-only denominators, equal conversation weighting and clipping after aggregation. Our coverage analysis also separates conversation themes from rubric axes, since those charts count different units. The original paper provides the historical results and distributions; our contribution is the score workbench and interpretation. This page centers original HealthBench, while Hard and Professional are distinct evaluation configurations.","task":{"input":"A health conversation ending at the next response to be evaluated.","output":"An open-ended assistant response.","unit":"Conversation with an example-specific rubric","setting":"Original 5,000-conversation benchmark; physician criteria graded by a model."},"dataOrigin":"Most conversations were generated synthetically through a physician-informed pipeline; physicians authored evaluation criteria. Do not describe the set as real patient records.","facts":[{"label":"Evaluation conversations","value":"5,000","detail":"The original benchmark denominator.","sourceIds":["hb-paper"],"locator":"Abstract; Section 2"},{"label":"Physician contributors","value":"262","detail":"Rubric development and annotation cohort.","sourceIds":["hb-paper"],"locator":"Abstract; Section 4.1"},{"label":"Unique criteria","value":"48,562","detail":"Distinct criterion text, not total repeated assignments.","sourceIds":["hb-paper"],"locator":"Abstract; Sections 2–3"},{"label":"Criterion assignments","value":"57,237","detail":"Axis-table denominator including repeated criteria.","sourceIds":["hb-paper"],"locator":"Table 3"},{"label":"Conversation themes","value":"7","detail":"Mutually exclusive partition of the 5,000 examples.","sourceIds":["hb-paper"],"locator":"Table 2"},{"label":"Rubric axes","value":"5","detail":"Behavior categories applied to criteria.","sourceIds":["hb-paper"],"locator":"Table 3"}],"metric":{"name":"Mean normalized rubric score","description":"For each conversation, divide met signed points by all available positive points. Average conversation scores and clip the mean to [0,1]; display times 100.","formula":"sᵢ = Σ(points × met) / Σ(positive points); score = 100 × clip(mean(sᵢ), 0, 1)","direction":"Higher means more rewarded and fewer penalized rubric behaviors under this evaluator.","comparability":"Keep rubric version, grader, sampling and subset fixed. A score of 60 is not 60% fully correct conversations.","sourceIds":["hb-paper","hb-code"]},"workflow":[{"label":"Generate the next response","detail":"Give the evaluated system the conversation context.","sourceIds":["hb-paper"]},{"label":"Grade each criterion","detail":"The original paper uses GPT-4.1 to decide whether each rubric criterion is met.","sourceIds":["hb-paper"]},{"label":"Normalize signed points","detail":"Keep negative points and divide by possible positive points for that example.","sourceIds":["hb-code"]},{"label":"Aggregate conversations","detail":"Average normalized scores, then clip the overall mean.","sourceIds":["hb-paper","hb-code"]}],"slices":[{"label":"Global health","value":1097,"unit":"conversations","detail":"Published mutually exclusive theme","sourceIds":["hb-paper"]},{"label":"Responding under uncertainty","value":1071,"unit":"conversations","detail":"Published mutually exclusive theme","sourceIds":["hb-paper"]},{"label":"Expertise-tailored communication","value":919,"unit":"conversations","detail":"Published mutually exclusive theme","sourceIds":["hb-paper"]},{"label":"Context seeking","value":594,"unit":"conversations","detail":"Published mutually exclusive theme","sourceIds":["hb-paper"]},{"label":"Emergency referrals","value":482,"unit":"conversations","detail":"Published mutually exclusive theme","sourceIds":["hb-paper"]},{"label":"Health data tasks","value":477,"unit":"conversations","detail":"Published mutually exclusive theme","sourceIds":["hb-paper"]},{"label":"Response depth","value":360,"unit":"conversations","detail":"Published mutually exclusive theme","sourceIds":["hb-paper"]}],"sliceTitle":"Seven conversation themes","sliceNote":"Paper Table 2. These counts sum to 5,000; rubric-axis counts use a different denominator.","results":[{"id":"healthbench-repeatability","title":"Historical overall means across repeated runs","metric":"Rubric score × 100","unit":"points","lower":0,"upper":100,"scope":"May 2025 paper, original HealthBench; mean across 16 evaluation runs in Table 7.","sourceIds":["hb-paper"],"locator":"Table 7","rows":[{"label":"o3","value":59.9,"display":"59.9","detail":"Reported run SD 0.16 points; run range 59.51–60.14"},{"label":"GPT-4.1","value":47.78,"display":"47.78","detail":"Reported run SD 0.22 points; run range 47.42–48.15"},{"label":"o1","value":42,"display":"42","detail":"Reported run SD 0.22 points; run range 41.53–42.30"},{"label":"GPT-4o (Aug 2024)","value":32.33,"display":"32.33","detail":"Reported run SD 0.20 points; run range 31.88–32.57"}],"note":"Paper-reported historical means. Standard deviations describe repeated-run variability, not confidence intervals for real-world clinical outcomes."}],"analysis":[{"heading":"The denominator is local to each conversation.","evidence":"Rubric points are normalized by available positive points before averaging.","interpretation":"A conversation with more criteria does not automatically receive more overall weight. The unit of aggregation stays the conversation.","sourceIds":["hb-paper","hb-code"]},{"heading":"Clipping too early changes the result.","evidence":"Individual scores may be negative; the final mean is clipped.","interpretation":"Replacing negative example scores with zero would remove some penalties before they can affect the aggregate.","sourceIds":["hb-paper","hb-code"]},{"heading":"Two coverage charts can count different objects.","evidence":"Themes count conversations; axes count criterion assignments.","interpretation":"A larger axis share is not evidence that the same share of patients or conversations has a particular condition.","sourceIds":["hb-paper"]},{"heading":"Grader validation has a defined scope.","evidence":"The paper’s meta-evaluation uses consensus criteria.","interpretation":"Do not extend that evidence into a claim that every individually authored criterion was independently validated by several physicians.","sourceIds":["hb-paper"]}],"limitations":[{"title":"Predominantly synthetic conversations","detail":"Realistic scenario design is not the same as observing outcomes in an actual care population.","sourceIds":["hb-paper"]},{"title":"Rubrics are incomplete descriptions","detail":"Criterion satisfaction may omit clinically relevant qualities; the paper acknowledges disagreement and incomplete coverage.","sourceIds":["hb-paper"]},{"title":"Model grader dependence","detail":"Changing the grader or grading prompt can change measurement; document both.","sourceIds":["hb-paper","hb-code"]},{"title":"Historical model configurations","detail":"The displayed study is from May 2025 and cannot identify today’s best-performing system.","sourceIds":["hb-paper"]}],"access":{"status":"Official data and reference implementation are public.","license":"Official dataset-card metadata: MIT.","restrictions":"Keep evaluation examples out of public explanatory pages to reduce contamination; use aggregate metadata and abstract illustrations.","url":"https://huggingface.co/datasets/openai/healthbench","sourceIds":["hb-card","hb-release"]},"sourceIds":["hb-paper","hb-release","hb-code","hb-card"]}],"explorer":{"kind":"weighted-rubric","title":"Build a HealthBench-style score","intro":"Toggle abstract illustrative criteria to see how signed points and the positive-point denominator interact. The example weights are invented for arithmetic; no benchmark questions or medical recommendations are reproduced.","caution":"This illustrates the published score rule. It does not perform model grading, represent an actual HealthBench example, or turn the result into percent accuracy. Negative example scores must be retained until aggregation.","sourceIds":["hb-paper","hb-code"],"rows":[],"parameters":[]},"references":[{"id":"hb-paper","title":"HealthBench: Evaluating Large Language Models Towards Improved Human Health","organization":"Rahul K. Arora and colleagues · OpenAI","url":"https://arxiv.org/html/2505.08775v1","note":"Original benchmark construction, coverage, scoring and historical evaluation variability.","locator":"Sections 2–5 and 8; Tables 2, 3 and 7; Appendix D","version":"arXiv v1 · 2025-05-13"},{"id":"hb-release","title":"Introducing HealthBench","organization":"OpenAI","url":"https://openai.com/index/healthbench/","note":"Original release and explanation of physician-authored evaluation rubrics.","locator":"Benchmark release","version":"Published 2025-05-12"},{"id":"hb-code","title":"HealthBench reference evaluator","organization":"OpenAI","url":"https://github.com/openai/simple-evals/blob/652c89d0ca9df547706735883097e9537d40dc47/healthbench_eval.py","note":"Pinned reference scoring implementation, including positive-point normalization.","locator":"calculate_score; calculate_length_adjusted_score; aggregate results","version":"Commit 652c89d0ca9df547706735883097e9537d40dc47"},{"id":"hb-card","title":"HealthBench official dataset card","organization":"OpenAI","url":"https://huggingface.co/datasets/openai/healthbench","note":"Official dataset access and MIT license metadata.","locator":"Dataset card","version":"Accessed 2026-09-28"}]}