Independent benchmark analysis / Original HealthBench; May 2025 paper v1

HealthBench

A rubric score is a balance of rewarded and penalized behavior.

HealthBench evaluates open-ended health conversations through physician-written rubrics. Its score can look like a percentage without being the percentage of answers that are fully correct. We make the calculation inspectable: signed points, positive-only denominators, equal conversation weighting and clipping after aggregation. Our coverage analysis also separates conversation themes from rubric axes, since those charts count different units. The original paper provides the historical results and distributions; our contribution is the score workbench and interpretation. This page centers original HealthBench, while Hard and Professional are distinct evaluation configurations.

01 / What is being tested?

The task, before the score.

input
A health conversation ending at the next response to be evaluated.
output
An open-ended assistant response.
unit
Conversation with an example-specific rubric
setting
Original 5,000-conversation benchmark; physician criteria graded by a model.

Data origin. Most conversations were generated synthetically through a physician-informed pipeline; physicians authored evaluation criteria. Do not describe the set as real patient records. [1][2][3][4]

Evaluation conversations
5,000

The original benchmark denominator.

Abstract; Section 2 [1]
Physician contributors
262

Rubric development and annotation cohort.

Abstract; Section 4.1 [1]
Unique criteria
48,562

Distinct criterion text, not total repeated assignments.

Abstract; Sections 2–3 [1]
Criterion assignments
57,237

Axis-table denominator including repeated criteria.

Table 3 [1]
Conversation themes
7

Mutually exclusive partition of the 5,000 examples.

Table 2 [1]
Rubric axes
5

Behavior categories applied to criteria.

Table 3 [1]
  1. 01

    Generate the next response

    Give the evaluated system the conversation context.

    [1]
  2. 02

    Grade each criterion

    The original paper uses GPT-4.1 to decide whether each rubric criterion is met.

    [1]
  3. 03

    Normalize signed points

    Keep negative points and divide by possible positive points for that example.

    [3]
  4. 04

    Aggregate conversations

    Average normalized scores, then clip the overall mean.

    [1][3]

Dataset anatomy

Seven conversation themes

Global health

Published mutually exclusive theme

1,097 conversations[1]
Responding under uncertainty

Published mutually exclusive theme

1,071 conversations[1]
Expertise-tailored communication

Published mutually exclusive theme

919 conversations[1]
Context seeking

Published mutually exclusive theme

594 conversations[1]
Emergency referrals

Published mutually exclusive theme

482 conversations[1]
Health data tasks

Published mutually exclusive theme

477 conversations[1]
Response depth

Published mutually exclusive theme

360 conversations[1]

Paper Table 2. These counts sum to 5,000; rubric-axis counts use a different denominator. Bar lengths use the largest listed count as their reference; they are not percentages of a shared population.

02 / Measurement

Mean normalized rubric score

Higher means more rewarded and fewer penalized rubric behaviors under this evaluator.

For each conversation, divide met signed points by all available positive points. Average conversation scores and clip the mean to [0,1]; display times 100.

Scoring definition

sᵢ = Σ(points × met) / Σ(positive points); score = 100 × clip(mean(sᵢ), 0, 1)

Keep rubric version, grader, sampling and subset fixed. A score of 60 is not 60% fully correct conversations. [1][3]

03 / Measured evidence

Results, with their conditions attached.

Paper-reported results / selected rows

Historical overall means across repeated runs

May 2025 paper, original HealthBench; mean across 16 evaluation runs in Table 7.

Rubric score × 100 · points
050100
Reported
o3Reported run SD 0.16 points; run range 59.51–60.14
59.9
GPT-4.1Reported run SD 0.22 points; run range 47.42–48.15
47.78
o1Reported run SD 0.22 points; run range 41.53–42.30
42
GPT-4o (Aug 2024)Reported run SD 0.20 points; run range 31.88–32.57
32.33

Paper-reported historical means. Standard deviations describe repeated-run variability, not confidence intervals for real-world clinical outcomes.

Source: Table 7 [1]

04 / Our original analysis

What follows from the design?

01

The denominator is local to each conversation.

Published evidence

Rubric points are normalized by available positive points before averaging. [1][3]

Our interpretation

A conversation with more criteria does not automatically receive more overall weight. The unit of aggregation stays the conversation.

02

Clipping too early changes the result.

Published evidence

Individual scores may be negative; the final mean is clipped. [1][3]

Our interpretation

Replacing negative example scores with zero would remove some penalties before they can affect the aggregate.

03

Two coverage charts can count different objects.

Published evidence

Themes count conversations; axes count criterion assignments. [1]

Our interpretation

A larger axis share is not evidence that the same share of patients or conversations has a particular condition.

04

Grader validation has a defined scope.

Published evidence

The paper’s meta-evaluation uses consensus criteria. [1]

Our interpretation

Do not extend that evidence into a claim that every individually authored criterion was independently validated by several physicians.

05 / Scope of the evidence

Where this benchmark stops.

Predominantly synthetic conversations

Realistic scenario design is not the same as observing outcomes in an actual care population. [1]

Rubrics are incomplete descriptions

Criterion satisfaction may omit clinically relevant qualities; the paper acknowledges disagreement and incomplete coverage. [1]

Model grader dependence

Changing the grader or grading prompt can change measurement; document both. [1][3]

Historical model configurations

The displayed study is from May 2025 and cannot identify today’s best-performing system. [1]

06 / Working with the benchmark

Access & reuse.

Open the author’s resource ↗
Availability
Official data and reference implementation are public.
License
Official dataset-card metadata: MIT.
Conditions
Keep evaluation examples out of public explanatory pages to reduce contamination; use aggregate metadata and abstract illustrations.
[4][2]

Evidence trail

Read the originals.

  1. HealthBench: Evaluating Large Language Models Towards Improved Human Health ↗

    Rahul K. Arora and colleagues · OpenAI. Original benchmark construction, coverage, scoring and historical evaluation variability.

  2. Introducing HealthBench ↗

    OpenAI. Original release and explanation of physician-authored evaluation rubrics.

  3. HealthBench reference evaluator ↗

    OpenAI. Pinned reference scoring implementation, including positive-point normalization.

  4. HealthBench official dataset card ↗

    OpenAI. Official dataset access and MIT license metadata.

Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.

Explore the assumptions ↗