Why clinicians still hesitate at AI decision support
7 min read
A resident opens a chart at 02:40. The CDS banner is already there: sepsis risk elevated, start the bundle, document why if you override. The patient looks borderline. The resident knows the alert fires often on this ward. They also know that an override gets logged, and that logged overrides get reviewed. The clinical question and the organisational question arrive in the same second.
That is the shape of AI clinical decision support in practice. Not a seminar about the future of medicine. A banner, a clock, and a career that has learned to treat the banner as part of the room.
The original puzzle was technical; the lasting one was not
In 1959, Robert Ledley and Lee Lusted published "Reasoning Foundations of Medical Diagnosis" in Science. They translated diagnosis into logical statements and Bayesian probability, and they expected physicians to welcome the collaboration. Adoption did not follow the paper. They had solved a formal problem and underestimated a social one: who owns the decision when a machine has already spoken, and what happens to trust when the machine is wrong in public.
We are watching a version of the same pattern with contemporary AI-backed CDS. The models are better. The deployment context is harder. The reluctance is not a failure to catch up with 1959. It is a response to incentives Ledley and Lusted did not have to design for: EHR-embedded alerts, malpractice culture, performance dashboards, and training data that do not look like every patient who walks in.
What the evidence actually supports, and what it does not
Controlled evaluations do show gains. A 2020 systematic review by Sutton and colleagues in npj Digital Medicine found CDS can improve diagnostic processes and reduce some medication errors under study conditions. Clinicians who use these tools day to day rarely describe them as a second doctor. They describe a noisy colleague who is useful for spotting trends and running calculations, and who collapses the moment the chart is missing the context the model assumed.
The more useful finding is not "AI makes clinicians better." It is that the tool reshapes the decision. Sometimes it sharpens attention. Sometimes it adds a second hypothesis the clinician would have missed. Sometimes it creates a gap between the model's confidence and the patient in the room. Performance that looks strong on one population can fail on another without the model "breaking," because the training set never saw the second population in the first place. Obermeyer and colleagues' 2019 Science paper on a widely used commercial algorithm is the clearest public example: a proxy for need that tracked cost instead of illness systematically under-referred Black patients.
If you only read vendor accuracy slides, you will miss that.
Why the hesitation looks different now
Today's reluctance is not technophobia dressed as professionalism. The concern clinicians name most consistently is autonomy under liability pressure. Once a recommendation sits in the chart, declining it is no longer a private act of judgement. It is an auditable event. That changes behaviour even when the clinician believes the algorithm is wrong for this patient.
Workflow friction is its own veto. A tool that generates alert fatigue or extra documentation gets worked around, however good the underlying model is. Bates and colleagues have documented this pattern for decades in medication alerting; generative and predictive tools inherit the same ward physics. A recommendation that cannot be acted on inside the time available is not a recommendation. It is noise with a confidence score.
There is also an equity awareness Ledley and Lusted's generation did not carry in the same form. A model trained mostly on data from well-resourced, urban, majority populations can encode exactly the care patterns equity-conscious practice is trying not to reproduce. The question is not only "can I trust the machine." It is "whose patterns is this machine encoding, and who is absent from the training set."
The evidence gap that actually matters
Most published evaluations of AI CDS still concentrate on specific disease areas in well-resourced settings. Real-world effectiveness across rural versus urban sites, insured versus uninsured patients, and racial and ethnic groups remains thinner than the launch rhetoric. The communities that stand to gain most from decision support when specialist access is scarce are often the same communities least represented in development and validation.
By the time differential performance is discovered, the tool is frequently already wired into order sets and quality metrics. Blind spots become institutional habits.
A regulatory lag, not a mysterious cultural failure
Existing oversight was built for slower medical devices. Many AI decision tools reach clinical use with limited scrutiny of training data, validation cohorts, or bias mitigation, often classified in ways that minimise formal review. Rajkomar, Dean, and Kohane argued in 2019 (New England Journal of Medicine) that machine learning in medicine will fail patients if equity is treated as an afterthought. Healthcare organisations are inventing governance in the gap: evaluation committees, shadow-mode pilots, post-deployment monitoring. The honest description is that we are often learning on live patients while building the structures that should have preceded go-live.
That is a reason for caution. It is not a reason to reject every tool. It is a reason to refuse the framing that treats clinician hesitation as ignorance.
A vignette worth keeping in mind
Picture a sepsis model with a published AUROC that looks excellent on the vendor slide. On the ward, the false-positive rate is high enough that night staff start documenting overrides with a canned phrase. The metric that survives in the dashboard is still "alerts acknowledged." The metric that does not appear is "minutes of attention stolen from the next patient." Nothing in that story requires the model to be fraudulent. It only requires the organisation to optimise the wrong visible number.
Students who came from clinical practice recognise this pattern immediately, because they have lived inside workflows that punished the honest override. That recognition is useful in an AI committee. Use it.
Questions worth carrying into an adoption meeting
If you are a student or early-career practitioner sitting in a vendor demo or a hospital AI committee, these are more useful than "does it work":
- Which patient-centred outcomes are you measuring besides diagnostic accuracy? Trust, time-to-decision, override rates with reason codes, equity of access.
- Is performance reported as one headline AUC, or stratified by demographic group and care setting?
- Who decided this gets implemented, and were the clinics most exposed to its failure modes in the room?
- What happens when the model is wrong at 02:40: clear override path, or career friction?
- Will the tool still be monitored after go-live, or does evaluation end at the pilot slide deck?
What to do with this if you are the square peg in the room
You do not need to become an ML researcher to participate usefully. You need to recognise when a conversation has collapsed "works on a retrospective dataset" into "safe for this ward." Clinical instinct is not a soft skill in that meeting. It is the check on whether the model's world matches the patient in front of you.
The persistence of physician reluctance is not proof that doctors are behind. It may be proof that this generation of clinicians understands something the pioneers of computerised diagnosis did not have to: a powerful tool requires a powerful safeguard, and "does it work" is incomplete until it is followed by "for whom, under what liability, and at what cost to attention."