# Health Benchmark > Independent analysis of original HealthBench: rubric scoring, seven themes, five axes, historical results and an interactive signed-points calculator. Original HealthBench evaluates open-ended health conversations with physician-written criteria. Its score combines rewarded behavior, penalties and a conversation-level denominator. This publication makes that measurement inspectable through sourced benchmark facts, a coverage breakdown and an abstract rubric calculator. The guides explain what a score means, how its units differ, and where grader validation applies. Arcophos provides independent analysis; OpenAI and its research collaborators created HealthBench. ## Provenance Independent analysis published by Arcophos. Benchmark creation belongs to the credited authors. Result rows are selected paper-reported measurements with their source versions and evaluation conditions, not new Arcophos runs or a live leaderboard. ## Benchmark dossiers - [HealthBench](https://healthbenchmark.ai/benchmarks/healthbench/): A rubric score is a balance of rewarded and penalized behavior. Source version: Original HealthBench; May 2025 paper v1. ## Original analyses - [How HealthBench scoring works, step by step](https://healthbenchmark.ai/guides/how-healthbench-scoring-works/): Follow signed rubric points, positive-point denominators and aggregate clipping without confusing the result with accuracy. - [Seven themes, five axes: two different HealthBench denominators](https://healthbenchmark.ai/guides/healthbench-themes-axes-and-denominators/): Read conversation coverage and rubric composition without treating criteria as patients or repeated assignments as unique items. - [What HealthBench grader validation does and does not establish](https://healthbenchmark.ai/guides/healthbench-grader-validation-and-repeatability/): Separate agreement with physicians, repeated-run variability and the clinical claims that neither measurement directly tests. ## Inspect the evidence - [Evidence JSON](https://healthbenchmark.ai/evidence.json): Task definitions, dataset facts, scoring rules, source-version results, our interpretations, and reference IDs. - [Sources](https://healthbenchmark.ai/sources/): Original papers and repositories with evidence locators. - [Editorial method](https://healthbenchmark.ai/methodology/): Source reconciliation and interpretation boundaries. - [About](https://healthbenchmark.ai/about/): Ownership and corrections. Analysis updated: 2026-09-28