May 22, 2026
The models knew the answers. The humans could not extract them.
For decades, the alarm management conversation focused on the device. Sensitivity. Specificity. Threshold optimization. The bedside monitor was never designed to account for what happens after the alarm fires. It was designed to detect a threshold crossing. The assumption: a better alarm produces a better response.

A randomized controlled trial published earlier this year in Nature Medicine tested what happens when real people use large language models to assess medical scenarios.
1,298 participants. Ten clinical vignettes. Three LLMs. One control group with no access to AI.
The models alone were remarkably capable. GPT-4o identified relevant conditions in 94.7% of cases. Llama 3 hit 99.2%. The accuracy of recommendations about what to do about it ranged from 48.8% to 64.7%.
Then real people used them.
Participants with LLM access correctly identified relevant conditions in fewer than 34.5% of cases. Correct on what to do: fewer than 44.2%. Both no better than the control group; people using Google, their own knowledge, whatever they normally reach for.
The models knew the answers. The humans could not extract them.
This is not a niche finding. One in six American adults already consults AI chatbots for health information at least monthly. These interactions are happening at scale, without clinical supervision, and we have been measuring the wrong thing.
The industry conversation about clinical AI centers on model performance. Benchmarks. Exam scores. Accuracy metrics on structured questions with clean inputs. The assumption: more capable models produce better human outcomes.
This study breaks that assumption with controlled evidence. The bottleneck was never the model’s capability. It was the interaction layer – the space between what the model knows and what the person does with that knowledge.
The model excels in isolation. Performance degrades at the human boundary. Systematically. Measurably. In a way that benchmarks do not predict.
Healthcare AI is being evaluated based on the models’ scores. It should be evaluated on what happens when people use them.
The interaction layer is unmeasured.
That is where safety lives.
The Nature Medicine study did not just show that LLM-assisted participants performed no better than controls. It showed where the interaction failed.
Three breakdowns. All at the human-AI boundary.
First: incomplete information. In 16 of 30 sampled interactions, initial messages contained only partial information. Key details emerged later, or never. In clinical practice, this is called the history. Physicians are trained to elicit it through structured questioning. The model was not given that role. It responded to what it received. What it received was incomplete.
Second: misinterpretation. When users provided information, models sometimes latched onto a single term rather than the full clinical picture. The model treated language as meaning when the user was approximating.
Third: non-adoption. Even when the model mentioned the correct condition during the conversation, users frequently omitted it from their final answers. LLMs suggested an average of 2.21 conditions per interaction. Users listed 1.33 in their final responses. The right answer surfaced, then was filtered out.
The most revealing detail: two participants described nearly identical symptoms of subarachnoid hemorrhage. Terrible headache, stiff neck, light sensitivity. One was told to seek emergency care. The other was advised to lie down in a dark room.
The inputs were semantically similar. The outputs were life-and-death different.
These are not model errors in the traditional sense. They are interaction layer failures. Information loss. Context collapse. Signal degradation at the boundary between human and system. The model performs. The boundary does not.
Benchmarks do not test boundaries.
That is why they do not predict safety.
The researchers tried one more thing.
They replaced real participants with LLM-simulated patients and reran the experiment. If simulated testing could predict real-world failures, it would be a scalable way to evaluate interaction safety.
The simulated patients performed better than real ones. And the results were only weakly predictive of actual human performance. Regression coefficients between simulated and real outcomes: 0.33 for GPT-4o, 0.31 for Llama 3, 0.20 for Command R+.
Benchmarks do not predict interaction safety. Simulations do not predict interaction safety. The only thing that predicts interaction safety is observing real interactions.
Hospitals have already learned this lesson in a different domain.
For decades, the alarm management conversation focused on the device. Sensitivity. Specificity. Threshold optimization. The bedside monitor was never designed to account for what happens after the alarm fires. It was designed to detect a threshold crossing. The assumption: a better alarm produces a better response.
It does not. A nurse responding to the 200th alarm on a shift does not evaluate alarm 201 on its merits. She evaluates it in the context of the channel it arrived through. If that channel has been contaminated by 190 non-actionable alarms that were not clinically relevant, the channel itself has lost credibility. The alarm is correct. The environment it arrives in is degraded.
The fix was never a better alarm. It was infrastructure that separated device state from patient state before the signal reached the clinician.
Make the channel trustworthy, and the signal regains meaning.
The same principle applies to every system where AI outputs reach human decision-makers. The model is capable. The benchmark confirms it. None of that matters if the interaction layer degrades the signal before the human can use it.
The question is not whether the AI knows the answer. It is whether the architecture preserves that knowledge through the boundary where it meets human judgment.
That boundary needs infrastructure. Not better models.
Continue Reading
ICU monitoring was built to catch the crash. There is no equivalent investment in the…
For decades, the alarm management conversation focused on the device. Sensitivity. Specificity. Threshold optimization. The…
The AI does not help. Every output arrives with the same confidence, whether the underlying…
