Back to Blog

Benchmark scores are an unreliable guide to clinical AI safety: The danger in errors of omission

by

Waymark

Icon

September 9, 2026

Back to Blog

Benchmark scores are an unreliable guide to clinical AI safety: The danger in errors of omission

by

Waymark

September 9, 2026

A published benchmark of 31 general-purpose models found that clinical safety only moderately tracks the general-reasoning and medical-knowledge scores used to rank them. For health plans weighing AI for Medicaid care management, a model’s reputation is an unreliable guide to how safely it handles a clinical case.

In a recent benchmark of general-purpose language models on real clinical consultations, the safest model still carried the potential for severe patient harm in roughly one in every 11 to 12 cases. The worst-performing model reached one in 4.5. Published as “First, do NOHARM,” the benchmark evaluated 31 models on 100 real primary-care-to-specialist consultation cases across 10 specialties, using 12,747 expert annotations for scoring.

The most telling – and concerning – result is the control condition. The researchers included a model that recommended nothing at all, a reassurance-only baseline. It scored a number needed to harm of 3.5, worse than every functioning model in the study. Essentially, the system that did the least produced the most potential for severe harm.

The gaps between models are themselves measurable. Compared with the safest model in the study, the worst-performing model adds 13 cases of severe harm potential per 100. The reassurance-only baseline adds 20. In the interactive visualization below, the difference view isolates this: the safest model sets the floor, and each weaker option shows how the additional harm grows as a system does less and less.

Potential for severe harm per 100 consultation cases. Source: Wu et al., arXiv:2512.01241.

Where the harm comes from

Most of the severe harm in the benchmark (76.6%) came from omission – the model failing to recommend a necessary action rather than recommending a wrong one. Harm to patients accrued directly from what these systems left out.

This follows from how general-purpose models are built. They have no concept of acuity and no triage step, and they’re optimized to produce fluent, complete-sounding answers. The trade-off is that no objectives guarantee a critical action reaches the patient. A model can read as measured and appropriate while omitting the referral or the escalation a case required – and because an omission is never indicated in a response, a human reading a response would be none the wiser. The reassurance-only baseline is that tendency taken to its limit, and its score shows why caution expressed as inaction is itself a clinical hazard.

How standard evaluation misses omissions

In the benchmark, clinical safety was only moderately related to the scores the field uses to rank models. Safety correlated with the general-reasoning and medical-knowledge benchmarks at r = 0.61-0.64, leaving more than half of the variation in safety unexplained by those scores. A model can post strong results on a reasoning test or a medical-knowledge exam and still rank poorly on the safety of its clinical recommendations. The benchmark’s authors reach the same conclusion: clinical safety must be measured directly because performance on existing evaluations does not reliably predict it.

Clinical safety cannot be read off a model’s reputation, its benchmark pedigree, or its score on a medical-knowledge exam. It has to be measured on its own terms, against cases that resemble the population the system will serve.

For example, a patient reports back pain along with new urinary symptoms, a combination that can signal a condition requiring prompt evaluation. A purpose-built clinical system treats that combination as an urgent triage question. However, a general-purpose consumer model, given the same description, returns exercise videos. That is an omission result in practice: a fluent, reasonable-sounding answer that leaves out the action the case required, the exact pattern the benchmark scores as the worst.

What a system designed against this standard looks like

A system built for this setting answers omission structurally. Waymark Compass, our clinical AI assistant for Medicaid patients, separates its safety-critical decisions from its conversation. Emergency detection and triage classification run on a deterministic layer, separate from the language model: the route a case takes is set by predetermined clinical logic rather than predicted by the model. The pathways for urgent action are built in, so a case that meets the criteria is escalated by design. Compass is clinician-supervised, trained and clinically tested on real patient cases, and engineered to detect emergencies using multiple clinical triage frameworks.

The published data sets the standard, and we built Compass to meet – if not exceed – it. Clinical AI for Medicaid care management must be judged by the safety of the actions it recommends and carries out, measured against cases that match the population it serves, because that safety does not follow from general capability. A system designed against that standard places its safety-critical decisions on a layer that does not guess; it acts on them. For a health plan evaluating these systems, the difference between a tool that adds risk to a case and one that does not is clinical AI built to take responsibility for the action a case requires, with safety built into the system itself.

Five Proofs That LLM Hallucination Cannot Be Eliminated, and the Implications for Care Delivery

by

Waymark

by

Read post
Back to Blog
Text Link
Waymark