When “it’s safe” isn’t safe
On 17 September, Leila Turner-Scott sat before lawmakers in Washington and described reading more than 700 conversations between her son and a chatbot.
Sam Nelson was 19, a psychology student, and by his mother’s account a normal kid who wanted to help people. He died at home on 31 May 2025 from a combination of Kratom, Xanax, and alcohol. Turner-Scott alleges that hours before his death, ChatGPT told him mixing Kratom and Xanax was safe.
The family sued OpenAI in California in May 2026. The company has called Sam’s death a heartbreaking situation, says the model he used is no longer available, and says it has strengthened its safety systems. No court has ruled on the allegations.
But the sentence at the center of the case deserves attention regardless of how the case ends. “It’s safe” is the most consequential thing a health answer can say. It’s also the easiest thing for a language model to say fluently.
Why “check with AI first” is a risky habit
Most of us now ask a chatbot before we ask anyone else. It's free, instant, and never makes you feel stupid for asking. At 2 a.m., that’s genuinely valuable.
The problem is what the answer looks like. Clean headings and a calm, confident tone read as authoritative. But fluency and accuracy are separate things. A model is optimized to produce text that sounds like a good answer. Nothing guarantees it is one.
So the instinct to double-check rarely fires. There’s nothing visibly wrong to check.
The danger is what’s missing
We picture AI health risk as a chatbot saying something absurd. The real pattern is quieter. Kratom, benzodiazepines, and alcohol all slow down the central nervous system. Ask a doctor whether it’s safe to combine them, and they won’t just answer yes or no. They’ll ask what else you're taking, whether you’ve had problems with substances before, and whether anyone else is at home tonight. With a question like this, the warning is the consultation.
An AI without that reflex answers the question asked, warmly and completely, and omits the part that mattered. This is measurable. In NOHARM, an independent benchmark built by a Stanford- and Harvard-led group of over 50 researchers, more than 80% of severe errors in AI medical recommendations were omissions rather than false statements.
An answer with a missing warning looks identical to a complete one. That’s why you can’t evaluate these systems just by reading them.
Someone has to test them
If a reader can’t tell a safe answer from an unsafe one by reading it, reading it isn’t a safety check. That gap has to be closed by testing, by people qualified to notice what’s absent, before the answer reaches a patient.
That work barely exists. Millions of health questions are answered by AI daily, almost none reviewed by a clinician. These systems are tested for whether users like them, which isn’t the same thing.
How iCliniq evaluates AI health advice
This is the work we’ve taken on alongside our clinical services, built around one finding that keeps repeating.
We score every AI answer twice. Our expert physicians assess medical accuracy, contraindications, safety communication, and sourcing. Separately, ordinary users assess clarity, usefulness, and trust.
The two groups disagree consistently and always in the same direction. Answers users rate as highly credible are routinely the ones our physicians flag for a missing risk warning or an unsourced claim. Users aren’t careless. They’re responding to tone and structure, the things a model is best at, which say nothing about whether advice is safe to act on.
We also retest the same questions over time, because these systems change without announcement. Running one set five months apart, we found a major platform's safety and sourcing performance had shifted substantially.
It’s slow, physician-heavy, manual work. It’s also the only thing standing between a fluent answer and a reader with no way to judge it.
What this means for you
Use AI to understand symptoms, prepare questions, and make sense of what a doctor told you. That’s a good use of it.
Don’t use it to decide. Not about medicines, not about combinations, not about whether something can wait until tomorrow. Run it past someone qualified first.
Turner-Scott asked lawmakers for one thing: that AI shouldn’t give medical advice without oversight. Whatever the courts decide, the principle holds. Someone has to check. Right now, mostly nobody is.