Assistencia Labs
For clinicians

What AI actually fixes in clinical admin, and what it does not

The tools work, but not always the way the pitch says. Here is what the evidence shows AI reduces in clinical admin, and where the wins are quieter than promised

By Ajay Bansal··4 min read
What AI actually fixes in clinical admin, and what it does not

Having spent three pieces describing the administrative burden, it would be strange to pretend technology cannot help with it. It can, and in specific places it clearly does. But the gap between what these tools are sold to do and what the evidence shows they do is wide enough to waste real money in, so this piece is about telling the two apart.

The short version: automation reliably helps at two layers, the documentation layer and the transactional layer, but the benefit often shows up as reduced burden rather than reduced minutes. That distinction sounds like a quibble. It is actually the most important thing to understand before you buy anything.

Ambient scribes: real adoption, careful claims

The most talked-about tool in this space is the ambient AI scribe: software that listens to the visit and drafts the clinical note, so the clinician can face the patient instead of the keyboard. The idea directly targets the two-to-one problem we opened the series with.

The best peer-reviewed account so far comes from The Permanente Medical Group, published in NEJM Catalyst in 2024. It is a serious deployment: an ambient scribe was made available to roughly 10,000 physicians and staff, and within the first ten weeks 3,442 physicians had used it across more than 300,000 patient encounters. Clinician feedback was described as favourable, the documentation it produced was judged high quality, and early patient reactions were positive.

Now the honest caveat, because it matters. That 2024 paper is an implementation report. It documents enthusiastic adoption, but it does not contain a rigorous, quantified measurement of hours saved or pyjama time reduced. Larger "hours saved" figures that circulate in press coverage come from later reporting, not from this peer-reviewed study, and they should be treated as vendor-adjacent until independently verified. The trustworthy claim today is that ambient scribes are being adopted at scale and clinicians like using them. The claim that they save a specific number of hours per week is not yet on the same evidentiary footing, and you should ask any vendor to show you their measurement, not their marketing.

Draft replies: the clearest lesson in the whole field

The inbox draft-reply studies are, oddly, the most instructive thing in this literature, precisely because their headline result is a "failure". When Stanford Health Care piloted AI-generated draft replies to patient messages with 162 clinicians in 2024, there was no statistically significant reduction in the time spent reading or answering messages. On a stopwatch, the tool did nothing.

But the clinicians' experience changed markedly. Their measured task load fell by about 14 points on a 100-point scale, and their reported work exhaustion dropped significantly too. The tool did not make the work faster. It made it lighter. Starting from a competent draft removed the cognitive friction of composing every reply from a blank page.

This is the single most useful reframing in the series. Time saved and burden reduced are different outcomes, and they do not always move together. A tool can be well worth having because it makes a draining task tolerable, even if it never shows up as minutes on a productivity dashboard. Conversely, a tool that "saves ten minutes" but adds anxiety, or shifts effort into checking its output, may be a poor trade. When you evaluate any of these products, decide up front which outcome you are buying, and measure that one.

The transactional layer: where the money actually is

Beyond documentation sits the least glamorous and most financially concrete opportunity: the routine transactions between providers and payers. Eligibility checks, claim submissions, claim status enquiries, and prior authorisations are high-volume, rule-bound, and increasingly standardised, which makes them well suited to automation.

As covered in the economics piece, the CAQH Index estimates that shifting the remaining manual and partly manual transactions to fully electronic ones represents an annual opportunity of around 20 billion dollars across the industry, roughly a fifth of what those transactions currently cost. This is where automation delivers hard-dollar returns rather than softer well-being ones. It is also, tellingly, the part of admin that clinicians rarely see and vendors rarely feature in a demo, because it happens in the back office rather than the exam room.

What automation does not fix

Two things, and both are important enough to have their own place in this series.

First, it does not touch the structural administrative complexity that comes from a fragmented payment system. As the economics piece showed, the largest review of healthcare waste could not identify a single intervention with demonstrated savings against administrative complexity. Software can make a broken process faster; it cannot make the process make sense. Expecting a tool to fix that is how disappointment gets budgeted.

Second, automation introduces new work and new risks of its own: output that must be checked, skills that can quietly erode, and confident-sounding errors. That is a large enough problem that the next piece is devoted to it.

How to read a vendor claim

A short checklist you can use in any sales meeting:

  • Is the promise time saved or burden reduced? Both are legitimate. Confusing them is not. Make them tell you which, and how they measured it.
  • Measured where, and by whom? A figure from the vendor's own uncontrolled pilot is weaker than a peer-reviewed study, which is weaker than a result reproduced somewhere like your setting.
  • At what accuracy, and who checks the output? A draft you must read line by line is not free. Ask what the error rate is and whose time absorbs the checking.
  • Does it add to anyone's inbox? Remember that nearly half of in-basket messages are already machine-generated. A tool that generates more notifications can subtract from well-being even as it claims to add efficiency.

Used this way, the good tools survive the questions and the weak ones do not. That is the entire skill. The technology in this field is improving quickly and some of it is genuinely worth adopting today, but the discipline of asking what, exactly, it fixes is what separates a modernised clinic from an expensively disappointed one.

Browse the full Admin Tax series.

References

  1. Tierney AA, Gayre G, Hoberman B, et al. Ambient Artificial Intelligence Scribes to Alleviate the Burden of Clinical Documentation. NEJM Catalyst Innovations in Care Delivery. 2024;5(3). https://catalyst.nejm.org/doi/full/10.1056/CAT.23.0404
  2. Garcia P, Ma SP, Shah S, et al. Artificial Intelligence-Generated Draft Replies to Patient Inbox Messages. JAMA Network Open. 2024;7(3):e243201. https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2816494
  3. CAQH. 2024 CAQH Index Report. https://www.caqh.org/hubfs/Index/2024 Index Report/CAQH_IndexReport_2024_FINAL.pdf
  4. Shrank WH, Rogstad TL, Parekh N. Waste in the US Health Care System: Estimated Costs and Potential for Savings. JAMA. 2019;322(15):1501-1509. https://jamanetwork.com/journals/jama/fullarticle/2752664