Independent benchmark analysis / Original HealthBench; May 2025 paper v1
HealthBench
A rubric score is a balance of rewarded and penalized behavior.
HealthBench evaluates open-ended health conversations through physician-written rubrics. Its score can look like a percentage without being the percentage of answers that are fully correct. We make the calculation inspectable: signed points, positive-only denominators, equal conversation weighting and clipping after aggregation. Our coverage analysis also separates conversation themes from rubric axes, since those charts count different units. The original paper provides the historical results and distributions; our contribution is the score workbench and interpretation. This page centers original HealthBench, while Hard and Professional are distinct evaluation configurations.
01 / What is being tested?
The task, before the score.
- input
- A health conversation ending at the next response to be evaluated.
- output
- An open-ended assistant response.
- unit
- Conversation with an example-specific rubric
- setting
- Original 5,000-conversation benchmark; physician criteria graded by a model.
Data origin. Most conversations were generated synthetically through a physician-informed pipeline; physicians authored evaluation criteria. Do not describe the set as real patient records. [1][2][3][4]
- Evaluation conversations
- 5,000
The original benchmark denominator.
Abstract; Section 2 [1] - Physician contributors
- 262
Rubric development and annotation cohort.
Abstract; Section 4.1 [1] - Unique criteria
- 48,562
Distinct criterion text, not total repeated assignments.
Abstract; Sections 2–3 [1] - Criterion assignments
- 57,237
Axis-table denominator including repeated criteria.
Table 3 [1] - Conversation themes
- 7
Mutually exclusive partition of the 5,000 examples.
Table 2 [1] - Rubric axes
- 5
Behavior categories applied to criteria.
Table 3 [1]
- 01
Generate the next response
Give the evaluated system the conversation context.
[1] - 02
Grade each criterion
The original paper uses GPT-4.1 to decide whether each rubric criterion is met.
[1] - 03
Normalize signed points
Keep negative points and divide by possible positive points for that example.
[3] - 04
Aggregate conversations
Average normalized scores, then clip the overall mean.
[1][3]
Dataset anatomy
Seven conversation themes
Published mutually exclusive theme
Published mutually exclusive theme
Published mutually exclusive theme
Published mutually exclusive theme
Published mutually exclusive theme
Published mutually exclusive theme
Published mutually exclusive theme
Paper Table 2. These counts sum to 5,000; rubric-axis counts use a different denominator. Bar lengths use the largest listed count as their reference; they are not percentages of a shared population.
02 / Measurement
Mean normalized rubric score
Higher means more rewarded and fewer penalized rubric behaviors under this evaluator.
For each conversation, divide met signed points by all available positive points. Average conversation scores and clip the mean to [0,1]; display times 100.
sᵢ = Σ(points × met) / Σ(positive points); score = 100 × clip(mean(sᵢ), 0, 1)
Keep rubric version, grader, sampling and subset fixed. A score of 60 is not 60% fully correct conversations. [1][3]
03 / Measured evidence
Results, with their conditions attached.
Paper-reported results / selected rows
Historical overall means across repeated runs
May 2025 paper, original HealthBench; mean across 16 evaluation runs in Table 7.
Paper-reported historical means. Standard deviations describe repeated-run variability, not confidence intervals for real-world clinical outcomes.
Source: Table 7 [1]
04 / Our original analysis
What follows from the design?
Clipping too early changes the result.
Two coverage charts can count different objects.
Themes count conversations; axes count criterion assignments. [1]
A larger axis share is not evidence that the same share of patients or conversations has a particular condition.
Grader validation has a defined scope.
The paper’s meta-evaluation uses consensus criteria. [1]
Do not extend that evidence into a claim that every individually authored criterion was independently validated by several physicians.
05 / Scope of the evidence
Where this benchmark stops.
Predominantly synthetic conversations
Realistic scenario design is not the same as observing outcomes in an actual care population. [1]
Rubrics are incomplete descriptions
Criterion satisfaction may omit clinically relevant qualities; the paper acknowledges disagreement and incomplete coverage. [1]
Model grader dependence
Changing the grader or grading prompt can change measurement; document both. [1][3]
Historical model configurations
The displayed study is from May 2025 and cannot identify today’s best-performing system. [1]
- Availability
- Official data and reference implementation are public.
- License
- Official dataset-card metadata: MIT.
- Conditions
- Keep evaluation examples out of public explanatory pages to reduce contamination; use aggregate metadata and abstract illustrations.
Evidence trail
Read the originals.
- HealthBench: Evaluating Large Language Models Towards Improved Human Health ↗
Rahul K. Arora and colleagues · OpenAI. Original benchmark construction, coverage, scoring and historical evaluation variability.
- Introducing HealthBench ↗
OpenAI. Original release and explanation of physician-authored evaluation rubrics.
- HealthBench reference evaluator ↗
OpenAI. Pinned reference scoring implementation, including positive-point normalization.
- HealthBench official dataset card ↗
OpenAI. Official dataset access and MIT license metadata.
Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.