May 9, 2026

The Cost of Opaque AI Recommendations

The AI does not help. Every output arrives with the same confidence, whether the underlying claim was verified against source data or inferred from context. The operator has no signal to distinguish

The Wharton School just published a study that measures the cost of opaque AI recommendations, and it is eye-opening.

When an AI recommendation arrives confident and unexplained, people accept it roughly 80 percent of the time. Including when the AI is systematically wrong.

Shaw and Nave ran three preregistered experiments. 1,372 participants. 9,593 trials. When the AI was accurate, performance improved by 25 points above baseline. When the AI was faulty, performance dropped 15 points below baseline. Worse than no AI at all. The effect size was large (Cohen’s h of 0.82).

The researchers call it *cognitive surrender*. Not laziness. Not inattention. A measurable shift in how humans process information when an external system presents answers with high confidence and no evidence chain.

Confidence went up even when accuracy went down. People felt more certain while being more wrong.

This is not theoretical. Budzyn and colleagues measured it in clinical practice. Endoscopists exposed to AI-generated recommendations showed measurable erosion of unaided diagnostic performance.

Published in the the Lancet.

Different specialty. Same mechanism.

An opaque recommendation suppresses deliberation. An auditable recommendation invites it. One produces surrender. The other preserves the clinician’s independent judgment.

This is a design choice, not a training problem. You cannot coach your way out of an architecture that was built to make acceptance the path of least resistance.

The architecture of the recommendation is what determines whether the clinician reasons or defers. That is not a training problem. It is an engineering one.

Kahneman gave us two systems. Fast and intuitive. Slow and deliberate.

Shaw and Nave’s 2026 research at Wharton identifies a third. They call it System 1.5. It activates specifically in human-AI collaboration.

The mechanism: when an AI output is fluent and confident, the human experiences the phenomenology of critical evaluation. She reads the output. She feels engaged. She feels like she assessed it. She proceeds.

She did not assess it. Adoption rate regardless of correctness: approximately 80%.

System 1.5 is dangerous because it is invisible from the inside. System 1 feels automatic, and you know it. System 2 feels effortful, and you know it. System 1.5 feels like System 2. You believe you are thinking critically.

You are not.

This maps directly to something hospitals have observed for decades.

A nurse who could face over two thousand alarms in a single shift is not ignoring alarm 200. But the environment she is judging in has been shaped by every alarm she has heard before it, in that same shift. The channel’s credibility has eroded. Not the clinician’s.

The same pattern appears in the use of operational AI. An operator reviewing the 15th AI output in a session is not careless. She is reading, processing, and approving. But the environment she is approving in has been shaped by the 14 correct outputs that preceded it. Her verification threshold has recalibrated. Not because she stopped caring. Because nothing signaled that she should care more.

The AI does not help. Every output arrives with the same confidence, whether the underlying claim was verified against source data or inferred from context. The operator has no signal to distinguish “I checked and this is true” from “I am guessing based on pattern.”

Confidence uniformity is a different kind of problem. In clinical environments, technical and physiological alarms at least present differently. AI outputs do not. Every recommendation arrives with the same fluency, the same confidence, the same formatting. There is no blue light for “I am guessing.”
The countermeasure is not “pay more attention.” If the system cannot signal its own uncertainty, vigilance becomes the bottleneck.

Vigilance does not scale.

Continue Reading

June 2, 2026
Blog & Insights

Recovery is invisible

ICU monitoring was built to catch the crash. There is no equivalent investment in the…