A patient texts a care navigation tool late on a Sunday night. She’s been wheezing and short of breath since the afternoon, has used her rescue inhaler four times with little relief, and mentions that her husband thinks she is overreacting and should just rest. A consumer AI assistant, reading the same messages, agrees with the husband. It acknowledges the breathing trouble and the inhaler use, then recommends monitoring overnight and booking an appointment in the next day or two.
But, as any clinician knows, an asthma exacerbation that has stopped responding to a rescue inhaler does not belong in the wait-and-see category. A rising CO2, accessory muscle use, and difficulty speaking in full sentences are the markers of an exacerbation turning into respiratory failure, and the interval between a patient who is still talking and a patient who is in serious trouble can be short. A clinician reading those same messages would do the opposite of a care navigation tool. They would escalate the issue while discounting the husband’s reassurance – because frightened family members routinely talk a patient down from care they need.
The assistant in that scenario was not broken. It was doing the thing it was built to do. It is also the failure case the evidence finds most often: in a 2026 Nature Medicine evaluation of ChatGPT Health, asthma exacerbation accounted for 28 of the 33 undertriaged emergencies. General-purpose AI assistants are optimized, in large part, to produce responses and answers that feel helpful and reassuring and agreeable to the consumer – and that is not the ideal, or even the safest, optimization for a patient-facing clinical care AI tool.
An objective that rewards agreement
This is an AI industry-wide issue. In April 2025, OpenAI created an update to their GPT-4o model, then pulled it within days after the model became, in the company’s own words, “overly flattering or agreeable,” validating users’ doubts and endorsing decisions it should have pushed back on. OpenAI later explained that the company had leaned too hard on short-term user feedback and had not been testing the model for sycophancy before release. That behavior came straight from the training objective, surfacing in production at the scale of a product millions of people use.
Similarly, most current assistants are tuned with reinforcement learning from human feedback, where people compare candidate responses and a reward model learns what they prefer. People reliably prefer answers that agree with them, and research has found that both human raters and the reward models built from their choices will pick a confident, well-written wrong answer over a correct one a meaningful share of the time. A study in Science documented this sycophancy across eleven leading models and stated that the agreeable behavior that puts a user at risk is the same behavior that keeps them engaged and rating the product highly, which is a contributing factor in why this continues to happen. More capable models do not grow out of this; in fact, a smarter model that is still rewarded for agreement applies that intelligence to agreeing more convincingly.
What this looks like at the point of care
A 2026 study in Nature Medicine ran ChatGPT Health, OpenAI’s consumer health feature, through 60 clinician-authored vignettes across 21 clinical domains. The model performed well in the middle of the acuity range and broke down at both ends of that range. Among gold-standard emergencies, it undertriaged 52 percent, routing patients with conditions like diabetic ketoacidosis toward a 24-to-48-hour appointment instead of the emergency department. Diabetic ketoacidosis is an emergency by definition – and the model recognized it, called it “early” or “mild,” and recommended outpatient management anyway.
The failures did not spread evenly across the cases. Common emergencies presenting with no argument – stroke, anaphylaxis, meningitis, aortic dissection – were triaged correctly every time. The misses clustered in the cases where danger depends on trajectory rather than presentation, essentially, when a patient looks stable in the moment but is moving toward something serious. The asthma exacerbation from the start of this piece is one of those cases, and at 28 of 33, it accounted for the large majority of the undertriaged emergencies in the study.
The same study measured the agreement problem directly. When a vignette included a friend or family member downplaying the symptom, the model shifted toward less urgent care, with an odds ratio of 11.7 for a change in the borderline cases. That reassurance is doing real work in the opening scenario. A minimizing comment from someone close to the patient was the single factor the researchers found most likely to push a recommendation in the wrong direction.
Designing for the case that matters
A system that holds up at the high-acuity edge has to be built against a different objective and that starts with how the system is measured. A model graded on average accuracy or on user satisfaction can post strong headline numbers while still missing half of emergencies, because emergencies are rare in the data and unwelcome in the conversation. For a triage tool, the number that counts is performance on the high-acuity cases specifically, including the ones whose urgency depends on where the patient is heading.
It also depends on where the triage decision gets made. When emergency detection runs inside the same conversational model that has been rewarded for being agreeable, the agreement objective will eventually override the safety call, the way a minimizing family member altered ChatGPT Health’s recommendations. Keeping the escalation logic on a separate track, governed by rules that do not bend to a reassuring patient or a soothing phrasing, is what holds the line.
Finally, it requires honest evaluation. Testing a clinical AI on textbook emergencies reveals very little, since the Nature Medicine results suggest almost any capable model handles those. The case that earns trust is the trajectory-dependent one, where the model has to act on a pattern instead of smoothing it over. A vendor either treats patient safety as a design constraint from the start or does not, and that decision is clear in how the product is evaluated and in how its triage logic functions and evaluates conversations.
For a medical director weighing AI for a Medicaid population, the demo and the benchmark headline matter less than what the system was actually optimized for. The questions that separate a safe tool from a plausible one are centered on the push-pull between evaluation and agreeability:
- How does it perform on emergencies, not on average?
- How does it handle the cases where urgency turns on trajectory rather than first impression?
- Can the part of the system that decides to escalate be talked down by the conversational layer’s pull toward agreement?
A tool built for engagement tends to fail those questions, in exactly the cases where failure has the most detrimental impact for patients and clinicians alike. Optimization is a design decision – before any patient is ever on the other end of the conversation – which is why these questions belong at the start of an evaluation, while the choice of which tool to trust is still open.
.png)
.png)