Back to Blog

The Proof Most Medicaid AI Skips: Five Years of Peer-Reviewed Evidence

by

Sanjay Basu

Icon

July 28, 2026

Back to Blog

The Proof Most Medicaid AI Skips: Five Years of Peer-Reviewed Evidence

by

Sanjay Basu

July 28, 2026

When we founded Waymark, we made a commitment that has shaped everything since. Any tool we put into Medicaid care, we would also put in front of independent reviewers, and we would publish what they found. We made that commitment before artificial intelligence was flooding into healthcare. We made it because the people we serve have too often been the testing ground for interventions no one had proven, and we were unwilling to add to that record.

Five years later, the commitment carries more weight, because the stakes have risen. AI is moving into clinical settings faster than anyone is validating it as a care delivery tool. Medicaid, the largest public payer in the country, is among the highest-stakes places for an unproven tool to fail; the people it covers have thin margins for error, long histories of being under-studied, and the least protection when a model gets something wrong.

That raises a question every health system should be asking: How do we actually know an AI tool helps patients receiving Medicaid, beyond a demo, a dashboard, or an organization's word? This is how Waymark is positioned to respond.

Where we started: A borrowed evidence base

In 2021, the case for the kind of care Waymark set out to deliver, community-based and multidisciplinary, built around the people closest to the patient, rested on evidence other researchers had generated.

That evidence was real and it was strong. A pooled analysis of three randomized controlled trials of community health worker programs, following 1,340 patients over 9,398 patient-months, found meaningful reductions in hospital days. A companion return-on-investment analysis found that every dollar invested returned $2.47 to a Medicaid payer within the year. A randomized study of 57,972 adult Medicaid beneficiaries found fewer avoidable hospitalizations when social needs were managed along with physical health.

That same literature came with a warning. One of the most rigorous tests of intensive care coordination ever run found no reduction in readmissions against usual care. The result was null, which was another important finding – that a model that looks obviously right can still fail when applied in the real world.

So two things were still unproven in 2021: whether the care model held up at real-world Medicaid scale, and whether the technology built to deliver it at that scale actually worked. Waymark had inherited the proof, but we had not yet generated any of our own.

What five years of evidence shows

Over the last five years, Waymark has published dozens of peer-reviewed studies on the tools it deploys, in the Nature journals, in NEJM Catalyst, in JAMIA, JMIR AI, and more. Submitting work to independent peer review means holding the company's own tools to a standard it does not control, in front of reviewers with no incentive to be generous, and publishing the result whether it flatters the company or not. That is the mechanism this piece is about, applied to Waymark's own technology, and to the care model it supports.

Everything Waymark builds follows the same arc: predict, prioritize, intervene. This is how our technology works, and it is the sequence the field most often collapses into a single risk score. Each stage has been validated in its own peer-reviewed studies, and each builds on the one before it. At every stage, the published evidence points in the same direction. The tool performs better than what it replaces; it performs equitably across the populations Medicaid serves, and the studies were designed to test for both.

1. Predict risk early

The first published question was the most basic one: Can avoidable acute care be seen before it happens?

In a 2024 study in Nature Scientific Reports, spanning roughly 10 million patients receiving Medicaid benefits, Waymark Signal™ identified patients headed for an avoidable emergency or hospital visit at more than three times the sensitivity of the standard cost-based risk model the field had been using. Two years later, a 2026 study in JAMIA found that an integrated Waymark acuity model achieved 81.3 percent sensitivity and 82.1 percent specificity for predicting 30-day acute care.

That performance held as Waymark extended the model across the populations Medicaid serves. Signal for Quality Improvement reached 84.5 percent accuracy across more than 14 million Medicaid beneficiaries and eliminated Black-white disparities across four quality measures. Signal for Maternity predicted adverse pregnancy outcomes a median of 55 days earlier than traditional clinical indicators. Signal for Duals predicted avoidable hospital and emergency visits at 80 percent accuracy for dual-eligible enrollees.

2. Prioritize who to reach first

Prediction alone does not help anyone. A risk score identifies who is sick, but whether that patient can be helped by outreach is a separate question, and the distance between the two is where most risk-stratification tools fall short.

Consider two high-risk patients. One is on a ventilator, already flagged as high-cost, and unlikely to be moved by an outreach call. The other has a history of falls and was just prescribed a blood thinner at a dose that could put them in the hospital with a brain bleed, and could be redirected to a safer alternative today. A cost-based risk score often ranks the first patient higher. A benefit-based model reaches the second.

Waymark's Most Likely to Benefit model predicts which rising-risk patients will respond to outreach, so teams spend their limited hours where an intervention can change the outcome. In a 2026 study in Health Services Research, prioritizing by predicted benefit reduced acute care visits by 92.4 per 1,000 member-months relative to risk-based prioritization, and by 208.4 per 1,000 member-months among patients who engaged.  By Waymark's own measurement, targeting by benefit delivers a 49 percent additional reduction in acute care events over rising-risk scores alone. This is the step most tools skip.

3. Better predictions become better outcomes

The point of a better prediction is a better outcome for the patient, measured in avoided hospitalizations and emergency visits.

A 2024 study in NEJM Catalyst found that teams guided by Waymark Signal and the company's AI-guided care recommendations reduced all-cause acute care events by 22.9 percent against a matched control group. That is the loop closing: earlier risk identification leads to more equitable outreach and thus reductions in avoidable acute care use for real patients.

The recommendations also improve as they operate. A 2025 reinforcement-learning study in JMIR AI showed Waymark's AI-guided care recommendations learning from deployment and reducing acute care events by 20.7 percent, with no cases of worsened outcomes attributable to following the recommendations.

Across all three stages, the throughline holds up: Waymark's tools perform better than what they replace, they perform equitably across the populations Medicaid serves, and that accuracy translates into fewer avoidable hospitalizations and emergency visits.

Why this matters now

AI is being deployed into healthcare faster than anyone is validating it, and Medicaid is the sharpest edge of that problem. The population has been historically under-studied, the margins for error are thin, and the risk models the field inherited have carried hidden bias for years.

Evidence is the accountability mechanism. An unvalidated tool does not announce that it is under-serving Black patients, and it doesn’t flag its own bias. Surfacing takes a peer-reviewed analysis designed to test for and correct it. In the same vein, building a model is the straightforward part – it’s generating independent proof that it works, and does not harm, that many are choosing to skip over.

That is the standard Waymark can defend in peer review. Our published evidence shows the tools are accurate, and that they close disparities the previous generation of models left in place. The same discipline applies to what Waymark builds next.

Waymark recently announced the launch of Waymark Compass, an SMS-based AI care navigator that helps patients enrolled in Medicaid get assistance with their care, benefits, and community resources. Compass is supervised by Waymark physicians and built on peer-reviewed algorithms that have been rigorously tested by Waymark’s community-based teams.

Our safety architecture underpins this rollout. Waymark ANCHOR (Auditable Navigation of Clinical Hazards with Oversight and Reasoning) is a clinical AI verification layer that checks clinical AI output against formal logic and cited evidence, built so that safety-critical decisions are architecturally separate from the language model.

The standard for any AI tool impacting care delivery and/or patients receiving Medicaid should be published, independent evidence that it works and does not harm. Waymark set that bar for itself five years ago, and intends to keep raising it. The next five years of this work will be measured the way the last five were: by the outcomes our solutions create and the positive impacts on patients’ health outcomes, communities, and lives.

Clinical AI Gets Smarter Every Quarter. It Doesn’t Get Safer.

by

Sanjay Basu

by

Read post
Back to Blog
Text Link
Sanjay Basu