Two years ago, asking a health system how it governed clinical AI usually produced a blank look or a pointer to the IT security review. Now, that has changed. The leading academic medical centers now run governance programs with named owners, intake processes, evaluation rubrics, and standing committees. The trade press has been slow to credit how far the field has come, still treating AI governance as an open problem after it has become a working discipline.
Stanford, Duke, Mayo, Mass General Brigham, and the broader field have built programs that are strong at two things: evaluating a model before it goes live, and monitoring how it performs across a population once it does. They do not yet do a third thing: check each individual clinical recommendation against the named guideline that governs it, at the moment it is produced, with a record a reviewer could later inspect.
That third capability is the layer the field has not built. Given what the 2026 evidence base has shown about how these models fail, it is also the layer the next phase of clinical AI safety will be organized around.
What “governance” has come to mean
The word “governance” covers more ground than it did a year ago, so it helps to be specific about what a mature program now contains.
Intake and risk classification comes first. Before a model goes near a clinical workflow, it is logged, categorized by risk, and routed to scrutiny proportional to that risk. A scribe that drafts a note a clinician will edit is treated differently from a tool that proposes a triage disposition.
Local validation comes next. A model that performed well in its development setting is re-evaluated against the institution’s own patients and data, because performance rarely transfers cleanly across patient mixes, documentation patterns, and care settings. Much of the field’s hard-won expertise sits here.
Fairness and bias assessment follows. Programs check whether a model performs evenly across the groups the institution serves, and whether deployment would widen existing disparities. This work has produced some of the most rigorous published methodology in the field.
Finally, ongoing monitoring. Once a model is live, programs track its performance over time and watch for the slow degradation that happens as patient populations shift and the conditions it was validated against stop holding.
Each of those four components work at the level of the model and its overall behavior. They ask whether this model, in this setting, performs acceptably across the patients it touches. They were not designed, however, to ask whether the specific recommendation in front of a clinician right now is consistent with the guideline that applies to that patient.
Takeaways from 6 defining programs creating standard AI governance practices
These are the programs that defined the practice, and the same boundary runs through all of them.
- Stanford. The FURM framework, from Stanford Health Care’s data science team and published in NEJM Catalyst in 2024, evaluates candidate models on three properties before deployment: whether they are fair, useful in the workflow they would enter, and reliable. Its insight is that a model’s benefit is inseparable from the workflow it operates in, which is why so many pilots that look promising in the abstract go nowhere. Stanford's 2026 follow-on in NEJM Catalyst extends this into postdeployment monitoring, organizing it around system integrity, performance, and impact across 13 live deployments, which makes it among the most complete model-level programs in the field.
- Duke. Through the Duke Institute for Health Innovation, Duke has done more than perhaps anyone to make governance portable. The Health AI Partnership, launched in 2022, organizes responsible adoption around eight decision points across an AI solution’s full life, from defining the problem through procurement, integration, and ongoing management. Its peer-reviewed work on health equity gives institutions a concrete method for assessing whether a tool will worsen disparities before they adopt it.
- Mayo. Mayo Clinic Platform, under John Halamka, has concentrated on transparency. Halamka’s “nutrition label” for AI algorithms holds that every deployed model should carry a standardized disclosure of the data it was built on and the populations it was validated against, so a clinician or administrator understands what they are relying on. A 2025 paper from Makhni, Halamka, and colleagues extends this into an institution-wide method spanning validation, equity, and lifecycle management.
- Mass General Brigham. MGB’s contribution is organizational, and instructive because of how recent it is. As late as 2024 it governed AI through its existing technology process, building a dedicated AI governance structure only as the volume and stakes grew. That structure now runs an AI steering committee of senior leaders across digital, operations, quality and safety, compliance, and finance, overseen by the institution’s Digital Trust function, a reminder of how young this discipline still is.
- The Coalition for Health AI published an Assurance Standards Guide and, in February 2025, a national registry of model cards for comparing validated tools against a common template. Its September 2025 guidance with the Joint Commission is the first responsible-AI framework from a US accrediting body.
- In March 2026, the National Academy of Medicine launched a two-year “Patient Safety in the Era of AI” initiative whose steering group includes leaders from Mayo and other major systems. Twenty-five years after To Err Is Human made patient safety a national priority, NAM is treating AI safety as the next chapter of that project.
Across all six, the achievement is real and the direction consistent. And across all five, the boundary is the same. Pre-deployment evaluation asks whether a model is fit to deploy. Ongoing monitoring asks whether a deployed model is still behaving as expected across the population it serves. Both are essential, and both work at the level of the model and its behavior in aggregate.
Neither answers the question the 2026 clinical evidence has made impossible to ignore: When a model produces a specific recommendation a clinician is about to act on, has that recommendation been checked against the named guideline that governs this patient’s situation, and is there a record of the check? We’ve written about this previously; essentially, the most common error in clinical AI outputs is omission, and while a model’s accuracy can look acceptable in aggregate, individual outputs can slide quickly away from accuracy standards.
Population-level monitoring surfaces trends across many outputs. It cannot certify the single output in front of a clinician, and the single output is where the harm occurs. This is not a deficiency in the programs above, which were built to do something else and do it well. It is a layer the field must adapt to addressing because until recently the tools did not exist and the evidence demanding it had not been published.
This is the layer we are building at Waymark, in a product called ANCHOR (Auditable Navigation of Clinical Hazards with Oversight and Reasoning). Anchor checks each recommendation a clinical AI produces against the named clinical guidelines that apply, the way a careful reviewer would, by following a fixed library of written clinical rules. Because it follows fixed rules, it returns the same answer every time, and it records which guideline and citation it checked against, so the decision can be reviewed later. It works alongside any model. We are not sharing performance figures here, because that evaluation is under peer review, and the purpose of this piece is the category, not the product. The field needs more than one serious attempt at checking outputs against clinical guidelines, and we would rather see this layer built well by many groups than left as someone else’s problem.
The useful question is no longer whether to govern clinical AI. Most leading systems already do, and a program with intake, local validation, fairness assessment, and ongoing monitoring is operating at the standard the field’s strongest institutions have set.
The next question is narrower. Does governance reach the level of the individual output? When a clinician acts on an AI-assisted recommendation today, is anything in place that checked that specific recommendation against the guideline that governs it, and could the institution reconstruct that check six months later if a regulator or a patient’s attorney asked?
For nearly every program in the field, including the strongest, the honest answer is not yet. That doesn’t mean these programs are failures. It means that we now know which layer is the one to build now.
.png)
.png)