ESC 2026
AI chatbots cave under pressure on sleep apnea advice

Clinical takeaway: Patients with likely OSA may arrive already reassured by a chatbot that their symptoms can wait. Ask what they've been told, and don't let that reassurance stand in for evaluation.
Patients with symptoms of obstructive sleep apnea (OSA) increasingly put their questions to a free AI chatbot before any clinician sees them. What the chatbot says back can decide whether loud snoring, witnessed pauses in breathing, or dozing at the wheel ever becomes a referral, or quietly remains unaddressed.
Testing of these tools has mostly asked whether they answer clearly worded medical questions correctly, and on that measure they perform well. But real patients are not clearly worded questions. OSA sits at the sharp end of the problem: diagnosis runs through referral for a sleep study, and the patients who need one often minimize their symptoms. A controlled conversation study presented at the European Respiratory Society Congress in Barcelona tested what happens when those two facts collide, holding the medical picture constant while only the simulated patient's attitude changed.
When the simulated patient was cooperative, all five chatbots recommended specialist assessment in every conversation, 350 of 350. When the same medical facts came from that patient minimizing symptoms and resisting referral, that advice survived in 64.3% of conversations.
The collapse was steepest where the stakes were highest. The chatbots gave way most often in the most severe scenarios, including those carrying a driving risk, and in the driving-risk scenarios, the dropped referrals usually came without any mention of the driving danger. The models also differed widely under pushback, with the best holding its advice in 85.7% of those conversations and the worst in 55.7%.
Researchers wrote seven fictional OSA patient cases, all meeting guideline criteria for sleep-study referral. No real patients were involved: each case ran as a scripted, multi-turn conversation against the five most widely used free chatbots, in two versions with identical medical facts, one in which the simulated patient cooperated and one in which the patient minimized symptoms and resisted referral, 700 conversations in all. Whether the chatbot maintained its referral recommendation was scored against a structured rubric.
The models never overestimated risk; every error ran toward agreeing with the patient who wanted permission to stall. And because performance varied so widely, a patient has no way of knowing whether the chatbot they consult is the one that will typically hold onto its initial advice.
"AI chatbots are widely available and we know that people are using them more and more to ask questions about their health. This means that we need to test them out in a realistic way to see how people might use chatbots and whether they respond in helpful or unhelpful ways," said Chi Yan (Io) Hui, PhD, chair of the European Respiratory Society's group on m-health and e-health and honorary fellow in digital health at the University of Edinburgh, who was not involved in the research.
"This research shows that chatbots may give good advice with the ideal 'cooperative' patient, but that they talk themselves out of it when talking to a more realistic, reluctant patient," she continued. "The problem is not what the chatbots know, it is how they handle disagreement; they appear to exhibit a tendency to please the user, a phenomenon known as 'AI sycophancy.'"
Source: Ratneswaran D, et al. (2026 Sep 6) European Respiratory Society Congress 2026. Consumer AI chatbots validate patient denial and abandon specialist referral for obstructive sleep apnoea: a multi-turn controlled conversation trial against five free frontier models