HealthBench · Inspect eval logs

HealthBench eval logs: 68 runs, 6 models, 8 benches

Start here

Every HealthBench run we have, as Inspect .eval logs, in one place. Six models (GPT-5.5, Claude Opus 4.7, DeepSeek-V4-Pro, PLaMo-3.0-Prime, MedGemma-27B, MedGemma-4B) across HealthBench full, consensus, hard, and Professional with its four use-case slices.

68
eval logs
2.3 GB
total size
23
fresh runs
34
cache replays

To pull a single log straight down, use the .eval button in the tables below, or fetch it directly:

huggingface.co/datasets/kirby44/healthbench-eval-logs/resolve/main/logs/<bench>__<model>__<date>.eval

Filenames carry the config, so professional__gpt-5.5__2026-07-24.eval needs no lookup. A -replay or -cached suffix means the model responses came from Inspect's cache rather than a fresh generation; -FAILED means the run errored and is kept only for provenance.

All 68 logs

Scores are ×100. raw is the HealthBench score, adj is the length-adjusted one. Grouped by bench, sorted by model.

68 logs

HealthBench full 10 logs

5000 samples, the whole open-ended set

ModelDateepnJudgerawadjProvenanceFile
gpt-5.52026-07-091openai/gpt-4o-minifailed.eval
gpt-5.52026-07-091openai/gpt-4o-minifailed.eval
gpt-5.52026-07-0915000openai/gpt-4o-mini48.7fresh.eval
gpt-5.52026-07-1515000gpt-4.156.955.8replay.eval
opus-4.72026-07-0915000openai/gpt-4o-mini47.6fresh.eval
opus-4.72026-07-1515000gpt-4.153.454.3replay.eval
deepseek-v4-pro2026-07-1615000gpt-4.151.441.7fresh.eval
plamo-3.0-prime2026-07-1615000gpt-4.139.432.4fresh.eval
medgemma-27b2026-07-2415000gpt-4.147.233.2fresh.eval
medgemma-4b2026-07-2415000gpt-4.127.018.3cached.eval

HealthBench consensus 8 logs

3671 samples, criteria physicians agreed on

ModelDateepnJudgerawadjProvenanceFile
gpt-5.52026-07-1613671gpt-4o-mini ?82.182.0replay.eval
opus-4.72026-07-1613671gpt-4o-mini ?80.280.2replay.eval
deepseek-v4-pro2026-07-1613671gpt-4o-mini ?79.178.5replay.eval
plamo-3.0-prime2026-07-1613671gpt-4o-mini ?75.274.8replay.eval
medgemma-27b2026-07-2413671gpt-4o-mini ?77.676.6fresh.eval
medgemma-27b2026-08-0513671gpt-4.191.090.1fresh.eval
medgemma-4b2026-07-2413671gpt-4o-mini ?71.470.8fresh.eval
medgemma-4b2026-08-0513671gpt-4.175.875.2fresh.eval

HealthBench hard 13 logs

1000 samples, the hardest slice

ModelDateepnJudgerawadjProvenanceFile
gpt-5.52026-07-091failed.eval
gpt-5.52026-07-1611000gpt-4o-mini ?27.326.0replay.eval
opus-4.72026-07-091failed.eval
opus-4.72026-07-1611000gpt-4o-mini ?26.627.8replay.eval
deepseek-v4-pro2026-07-1611000gpt-4o-mini ?24.813.8replay.eval
plamo-3.0-prime2026-07-1611000gpt-4o-mini ?17.49.6replay.eval
medgemma-27b2026-07-2411000gpt-4o-mini ?21.14.8cached.eval
medgemma-27b2026-08-0511000gpt-4.113.3-4.4fresh.eval
medgemma-27b2026-08-0511000gpt-4.114.1-2.3fresh.eval
medgemma-4b2026-07-2411000gpt-4o-mini ?10.61.3fresh.eval
medgemma-4b2026-08-0511000gpt-4.1-3.1-16.4fresh.eval
medgemma-4b2026-08-0511000gpt-4.1-3.5-12.6fresh.eval
gpt-5-nano2026-07-091failed.eval

HealthBench Professional 7 logs

525 samples, physician-written, has a human baseline

ModelDateepnJudgerawadjProvenanceFile
gpt-5.52026-07-2484200gpt-5.452.947.8fresh.eval
opus-4.72026-07-2484200gpt-5.450.848.0fresh.eval
deepseek-v4-pro2026-07-251525gpt-5.434.327.4cached.eval
deepseek-v4-pro2026-08-061525gpt-5.437.831.0fresh.eval
plamo-3.0-prime2026-07-2484200gpt-5.420.813.7fresh.eval
medgemma-27b2026-07-2484200gpt-5.431.220.0fresh.eval
medgemma-4b2026-07-2584200gpt-5.416.59.0fresh.eval

Professional: consult 6 logs

236 samples

ModelDateepnJudgerawadjProvenanceFile
gpt-5.52026-07-2481888gpt-5.451.048.6replay.eval
opus-4.72026-07-2481888gpt-5.449.147.0replay.eval
deepseek-v4-pro2026-07-251236gpt-5.431.225.6replay.eval
plamo-3.0-prime2026-07-2481888gpt-5.421.815.4replay.eval
medgemma-27b2026-07-2481888gpt-5.428.517.8cached.eval
medgemma-4b2026-07-2581888gpt-5.415.28.2cached.eval

Professional: writing 6 logs

142 samples

ModelDateepnJudgerawadjProvenanceFile
gpt-5.52026-07-2481136gpt-5.440.636.0replay.eval
opus-4.72026-07-2481136gpt-5.439.536.1replay.eval
deepseek-v4-pro2026-07-2481136gpt-5.49.55.0fresh.eval
deepseek-v4-pro2026-07-251142gpt-5.49.75.2replay.eval
plamo-3.0-prime2026-07-2481136gpt-5.4-2.8-4.3replay.eval
medgemma-27b2026-07-2481136gpt-5.418.99.1cached.eval

Professional: research 6 logs

147 samples

ModelDateepnJudgerawadjProvenanceFile
gpt-5.52026-07-2481176gpt-5.468.057.9replay.eval
opus-4.72026-07-2481176gpt-5.464.661.1replay.eval
deepseek-v4-pro2026-07-2481176gpt-5.463.752.9fresh.eval
deepseek-v4-pro2026-07-251147gpt-5.463.051.8replay.eval
plamo-3.0-prime2026-07-2481176gpt-5.441.828.6replay.eval
medgemma-27b2026-07-2481176gpt-5.447.834.4cached.eval

Professional: red-teaming 6 logs

191 samples

ModelDateepnJudgerawadjProvenanceFile
gpt-5.52026-07-2481528gpt-5.429.928.2replay.eval
opus-4.72026-07-2481528gpt-5.428.326.7replay.eval
deepseek-v4-pro2026-07-2481528gpt-5.4-3.7-6.9fresh.eval
deepseek-v4-pro2026-07-251191gpt-5.4-5.3-8.3replay.eval
plamo-3.0-prime2026-07-2481528gpt-5.4-9.9-11.8replay.eval
medgemma-27b2026-07-2481528gpt-5.41.5-6.8replay.eval

Professional: physician baseline 6 logs

the human reference, model-independent

ModelDateepnJudgerawadjProvenanceFile
gpt-5.52026-07-2484200gpt-5.444.343.9baseline.eval
opus-4.72026-07-2484200gpt-5.444.343.9baseline.eval
deepseek-v4-pro2026-07-2484200gpt-5.444.343.9baseline.eval
deepseek-v4-pro2026-07-251525gpt-5.443.342.9baseline.eval
plamo-3.0-prime2026-07-2484200gpt-5.444.343.9baseline.eval
medgemma-27b2026-07-2484200gpt-5.443.943.5baseline.eval

fresh model responses generated in this run   replay every call served from cache   cached under 200 candidate tokens per sample   baseline human responses, no generation by design   failed errored or cancelled
? judge not recorded in task_args, inferred from the scorer default and confirmed against stats.model_usage

Analysis

DocumentWhat it answers
Coverage matrix Which model ran which bench, which cells are comparable, and what is still missing.
Config check v2 Do our numbers reproduce OpenAI's published HealthBench results? Anchored on the physician baseline (ours 43.9 against their 43.7).
Config check v1 The earlier pass over the first four spaces. Superseded by v2, kept for history.

Before you quote a number

Four things will bite you if you take a score straight out of a log.

IssueWhat to do
The judge is not constant. Three graders are in play: gpt-4o-mini (hard, consensus), gpt-4.1 (full, and the Aug-05 MedGemma re-runs), gpt-5.4 (all Professional). Swapping the judge moves a score by up to 14 points, and not always in the same direction. Only compare runs sharing a judge. The Judge column above is the check.
Professional epochs are inconsistent. 8 epochs for most models, 1 for DeepSeek. Check the ep column before putting two Professional rows side by side.
In-log subset metrics are wrong. use_case_*_score, specialty_*_score, difficulty_*_score and source_slice_*_score drop the length adjustment and clip each sample to [0,1] first. Errors run up to +32 points, always upward. Use the standalone professional-* logs above for the four use-case slices. For specialty and difficulty, re-aggregate from per-sample scores yourself.
cache=true on every run. 34 of 68 logs served some or all calls from cache, so an empty stats.model_usage is a replay, not a run. The Provenance column above already classifies this.

The pipeline itself is validated: the physician baseline on Professional lands at 43.9 against OpenAI's published 43.7. Discrepancies in the model numbers are config drift, not a broken harness.

How this was assembled

The runs were executed by Ajay between 2026-07-09 and 2026-08-06 and originally published as ten separate HuggingFace Spaces under ajay-citadel. This Space consolidates all of them into one viewer, renames the logs so the config is legible from the filename, and adds the provenance classification that the raw logs do not carry.

data/log_mapping.csv maps every renamed file back to its original space and filename, so nothing here is a dead end. data/MANIFEST.csv is the full per-run header dump: judge, epochs, token counts, package versions. data/INDEX.md is the short version of the traps list.

The original spaces remain the upstream source. If a number here disagrees with one there, the logs are byte-identical, so the difference is in which run you are reading, not in the data.