The short answer
A HealthBench score can be displayed on a 100-point scale without being the percentage of conversations that are completely correct. The benchmark combines graded behaviors inside each conversation, then aggregates across conversations. This guide follows that calculation and uses abstract hypothetical arithmetic to show why the order of operations matters. No actual benchmark examples are reproduced.
Start with a criterion and its signed weight
Each rubric item describes a behavior and assigns points. The grader decides whether the criterion is met. Positive points reward a desired property; negative points penalize an undesirable property. A negative item is therefore scored when the unwanted behavior is present, rather than when a desired behavior is absent.
That distinction is easy to lose in a checklist interface. Our calculator keeps the sign visible and labels its criteria as abstract illustrations. The user supplies the met decisions. It demonstrates arithmetic only; it does not perform the model-grading judgment required to decide whether a real response satisfies a physician’s criterion.
Divide by available positive points
For one conversation, sum the signed points of every met item. Divide that numerator by the sum of all positive point values in the rubric. Negative items do not increase the denominator. This normalization asks how much of the available positive score remains after both earned credit and triggered penalties.
Consider an invented rubric with seven possible positive points. Earning three positive points and triggering a two-point penalty leaves one point in the numerator, giving one seventh of the available score. The example illustrates the formula; its weights and outcomes are not measurements of any model or a released HealthBench conversation.
Keep negative conversation scores
If penalties exceed earned positive points, the conversation score can be negative. That is permitted by the published definition and reference implementation. It preserves the effect of serious penalized behavior when the benchmark later averages across examples. A display that floors every conversation at zero would implement a different metric.
For a purely hypothetical pair of conversation scores, take minus one half and one. Their mean is one quarter. If the first value were prematurely replaced by zero, the mean would become one half. The numerical difference comes entirely from changing the aggregation rule, not from a better response.
Average conversations, then clip the mean
HealthBench averages normalized conversation scores and clips the final mean to the interval from zero to one. Multiplying by one hundred is a presentation step. The conversation remains the unit of aggregate weighting even when different conversations have different numbers of criteria or available positive points.
A rubric with more items does not automatically give its conversation more weight in the overall mean. It can still change what is measured inside that conversation. This is why criterion count and conversation count are both useful metadata but cannot be substituted for each other when explaining the score.
Interpret the result as a rubric measurement
A score of sixty does not identify sixty percent of conversations as fully correct. Different patterns of omissions, rewarded behaviors and penalties can produce the same mean. To understand a system, retain the distribution and the kinds of criteria it misses, not only the final average.
The original paper and code define the benchmark calculation; our workbench makes that calculation easier to inspect. A reproduction also needs the conversation set, rubric version, grader and generation settings. A score becomes meaningful when those choices stay attached to it, and its interpretation remains about the rubric behaviors that were actually evaluated.
References & further reading
These original sources support the methods discussed here. Our suggested planning steps are editorial guidance, not an endorsement by the source authors.
- HealthBench: Evaluating Large Language Models Towards Improved Human Health ↗Rahul K. Arora and colleagues · OpenAI. Original benchmark construction, coverage, scoring and historical evaluation variability.
- HealthBench reference evaluator ↗OpenAI. Pinned reference scoring implementation, including positive-point normalization.
Published by Arcophos. Educational material, not clinical advice or a claim of regulatory compliance. Read our editorial method.