The AI Doctor Isn’t the Story—The AI Gatekeeper Is

Everyone Is Asking if AI Will Replace Doctors. They’re Asking the Wrong Question.

What the Courts Are Forcing Insurers to Reveal About the Algorithms Deciding Your Care

From the Craig Bushon Show Media Team

On March 9, 2026, a federal magistrate judge in Minnesota ordered the largest health insurer in the country to open its books on a piece of software. UnitedHealth Group must produce internal records showing how a predictive tool called nH Predict was built and what it was designed to do, documents concerning the company’s internal AI review board along with the identities of the people who sat on it, performance evaluation and compensation records for the post-acute care coordinators and medical directors who applied the tool, and the names of the personnel who issued denials to hundreds of class members. Some categories of documents reach back to January 2017. The court granted or partially granted six of the seven categories the plaintiffs requested. A discovery order is not a verdict, and that distinction deserves to be stated plainly, because nothing here has been proven. What the order establishes is narrower and still consequential: the inner workings of the software that decides whether American patients receive care their doctors ordered are no longer beyond the reach of a courtroom.

The case was brought in November 2023 by the families of two deceased Wisconsin men who held Medicare Advantage plans, and the lead plaintiff is the estate of Gene Lokken, who was 91 when he fell at home in May 2022 and fractured his leg and ankle. He was recovering, and an orthopedic physician had placed him in a removable boot and ordered intensive physical therapy, when the insurer terminated his skilled nursing coverage that July on the grounds that further days were unnecessary. His family appealed and lost repeatedly, then paid between twelve and fourteen thousand dollars a month out of pocket for nearly a year until his death in July 2023. The tool at the center of the case was developed by naviHealth, a company UnitedHealth acquired in 2020 and has since rebranded, and the plaintiffs allege the insurer used it to supplant physicians’ clinical judgment and prematurely refuse coverage for medically necessary care. Optum disputes that characterization, maintaining that nH Predict is a care-support tool rather than a claims adjudication instrument and that medical necessity determinations are made by qualified physicians in accordance with federal guidance. In February 2025 the court dismissed several claims on Medicare Advantage preemption grounds while allowing the breach of contract and good-faith claims to move forward, holding that those claims turned on whether the company honored its own policy language promising that coverage decisions would be made by clinical staff and physicians. UnitedHealth then asked the court to split discovery in two and start with the narrow question of whether nH Predict was actually used on these particular claims. The court refused, finding the factual and legal issues too enmeshed to separate, and ordered class-wide discovery to proceed from the outset. That refusal is what set up the March order, and the case is now moving toward class certification, with declarations on that question due September 14, 2026.

The arithmetic underneath the litigation is the part worth noting. A Senate investigation found that UnitedHealthcare’s denial rate for post-acute care claims more than doubled after the company began using naviHealth and nH Predict in 2019, and the magistrate judge cited that finding in the March ruling as the reason records predating the algorithm’s deployment were relevant to the plaintiffs’ theory of the case. The plaintiffs further allege that roughly nine of every ten people who challenged an nH Predict denial won their appeal, while only about two-tenths of one percent of policyholders ever filed a challenge at all. Those figures are allegations, and the second traces back to an outside estimate rather than the insurer’s own records. But the government’s data points the same direction. An analysis of federal records found that across all of Medicare Advantage in 2024, patients appealed only about 11.5 percent of denied prior-authorization requests, and 80.7 percent of the appeals that were filed got overturned. A parallel suit against Humana, which used the same tool and survived a dismissal bid in August 2025, puts its own appeal rate at roughly two percent. Cigna is defending a separate class action in the Eastern District of California over a system called PxDx, which according to ProPublica’s reporting was used to reject more than 300,000 payment requests across two months in 2022, with medical directors averaging 1.2 seconds per claim and one reportedly denying 60,000 in a single month. Cigna maintains that PxDx is simple sorting technology rather than artificial intelligence and says the reporting mischaracterized it; a judge allowed key claims to proceed in March 2025 after finding the company’s reading of its own plan terms to be an abuse of discretion. The defendants differ, the technologies differ, and the companies dispute the allegations against them. The pattern in the numbers holds regardless: a process that loses most of the fights it picks, aimed at a population that almost never shows up to fight.

Legislators have noticed, and the volume of their response is its own kind of evidence. The Centers for Medicare & Medicaid Services has told Medicare Advantage plans that artificial intelligence may assist in prior authorization but must account for a beneficiary’s individual clinical circumstances and the recommendations of the treating physician rather than leaning on group datasets. The states have moved faster and considerably harder. Washington’s SB 5395, which took effect in June, provides that artificial intelligence shall not be the sole means used to deny, delay, or modify care and requires that only a licensed physician or health professional issue a medical-necessity denial. Alabama enacted a law in April requiring insurers to certify annually that their AI tools do not rely on group datasets and to disclose their use in utilization-review policies. Georgia’s version takes effect in January 2027 and bars an AI system from issuing an adverse determination until a qualified human reviewer working with a clinical peer has conducted the review. Indiana’s addresses AI-driven downcoding of claims, and California’s SB 1120 was among the first laws of its kind in the country. More than a dozen new state statutes governing artificial intelligence in healthcare have been enacted in 2026 alone, with legislation introduced or passed across the large majority of states. Nobody writes that many laws about a hypothetical problem.

None of this is what most people picture when they hear about artificial intelligence in medicine, and that gap is precisely the point. For generations, healthcare has followed the same basic pattern. You don’t feel well, so you schedule an appointment. A physician asks questions, orders tests, reaches a diagnosis, and recommends a treatment. It has always been a fundamentally reactive system, one that waits for disease to announce itself before anyone acts, and the promise attached to artificial intelligence is that it finally breaks that pattern. On the clinical side of the ledger, that promise is real, uneven, and much further along than the public conversation suggests.

The scale is a matter of public record. The Food and Drug Administration maintains a running catalog of AI-enabled medical devices authorized for marketing in the United States, and the version released this past June, covering decisions through the end of March 2026, lists 1,524 of them. Radiology accounts for 1,163, roughly three-quarters of the total. In the first quarter of 2026 alone the agency authorized 92 devices, a 28 percent increase over the final quarter of 2025. GE HealthCare leads all manufacturers with 130 authorizations, followed by Siemens Healthineers at 95 and Philips at 58. The agency cautions that its own list is not a comprehensive inventory. Even so, artificial intelligence has been clearing regulatory review at a pace of roughly one device a day, and it is already reading images alongside your radiologist.

That number is where most coverage of this subject stops, and it is where the harder question starts, because clearance is not proof of benefit. A cross-sectional study published in JAMA Health Forum examined all 691 artificial intelligence and machine-learning devices the FDA cleared between September 1995 and July 2023, and its findings are difficult to reconcile with the enthusiasm surrounding the field. Six of those devices, amounting to 1.6 percent, reported data from a randomized clinical trial, and three of them, fewer than one percent, reported actual patient outcomes rather than analytical measures such as sensitivity and specificity. Nearly half of the decision summaries did not describe the study design at all, more than half omitted the training sample size, and 95.5 percent said nothing about the demographic makeup of the populations the software had been tested on. Premarket safety assessments were documented for 28.2 percent of the devices, while postmarket adverse events, including one reported patient death, were recorded for 36 of them. Almost all of this equipment, 668 devices out of 691, reached the market through the 510(k) pathway, which asks whether a product is substantially equivalent to something already being sold rather than whether it leaves patients measurably better off. The honest read of the FDA’s list is that it documents the growth of an industry rather than delivering a verdict on the medicine.

Where rigorous evidence does exist, it deserves to be taken seriously. Sweden’s MASAI trial, run inside the national breast screening program with more than 105,000 participants, is the first randomized controlled trial of artificial intelligence in cancer screening and the largest study of its kind anywhere. AI-supported screening increased cancer detection by 29 percent without increasing false positives while cutting radiologists’ screen-reading workload by 44 percent. The full results published in The Lancet in January went further, showing a 12 percent reduction in interval cancers, the ones that surface between screening rounds and tend to be the most lethal, along with 27 percent fewer of the aggressive subtypes. Sensitivity rose from 73.8 percent to 80.5 percent at identical specificity. That is not a vendor announcement. That is a randomized trial reporting that the machine found more of the cancers that kill people while the radiologists read fewer images.

The continuous-monitoring future gets described in glossy terms, and it exists today in pieces rather than as a whole. Apple’s atrial fibrillation history feature cleared the FDA in 2022, AliveCor’s mobile ECG products are cleared, and Eko’s software for analyzing heart sounds was cleared back in January 2020. What does not yet exist is the integration, meaning a system that pulls heart rhythm, blood oxygen, sleep, activity, glucose trends, laboratory results, imaging, family history, and genetic data into one continuously updated picture and flags a deterioration weeks before symptoms drive someone to an emergency room. Each individual sensor is real and cleared. The assembled whole remains a projection, and it should be labeled as one.

Drug development has produced the clearest single proof point so far. Insilico Medicine used one AI platform to identify a novel fibrosis target and a separate generative platform to design a molecule against it, producing rentosertib, a candidate treatment for idiopathic pulmonary fibrosis. Phase IIa results published in Nature Medicine in June 2025 covered 71 patients and showed a mean improvement in forced vital capacity of 98.4 milliliters at twelve weeks in the once-daily 60-milligram arm, with manageable safety and tolerability. That publication is widely described as the first peer-reviewed clinical proof of concept for a drug whose target and chemical structure both emerged from artificial intelligence, and the program has since entered a Phase III trial expected to enroll about 320 patients in China. The company reports that its candidates from 2021 through 2024 reached preclinical nomination in twelve to eighteen months against a traditional benchmark of two and a half to four years, though that comparison comes from Insilico itself and should be weighed accordingly. Roughly nine out of ten drugs that enter clinical trials still fail, and compressing the discovery phase does nothing to change the stretch of the pipeline where most candidates die.

For practicing physicians, the nearer-term value is subtraction rather than addition. Doctors spend substantial portions of their day documenting visits, reviewing records, completing insurance paperwork, and hunting for relevant literature. Technology has genuine potential to lift much of that burden, which would return time to the work that actually requires a physician: listening, applying judgment, communicating hard decisions, and weighing the human factors no algorithm grasps. Technology can analyze data. Experience, compassion, ethics, and wisdom remain distinctly human responsibilities, and no amount of computing power transfers them to a machine.

History offers a caution worth keeping in view through all of it. Medical science evolves, and treatments once considered groundbreaking have later been abandoned as better evidence emerged. Physicians once recommended practices that modern medicine has since rejected outright. That record does not make medicine unreliable; it demonstrates that science advances precisely by testing and revising its own conclusions. Artificial intelligence will follow the same path, with some early expectations proving correct and others proving badly overstated. The systems that succeed will be the ones that pair adoption with validation instead of assuming that a new algorithm is automatically better than established practice.

Reading Between the Lines

The question everyone asks is whether artificial intelligence will replace physicians, and the fixation on that question has functioned as an extraordinarily effective distraction. The AI doctor is not what should concern anyone. The AI gatekeeper already does the work that matters, and it has been doing it for years without much public accounting.

Set the two halves of the record side by side and the shape of the problem becomes clear. On the clinical side, the technology is proliferating far faster than the evidence supporting it, with more than fifteen hundred authorized devices and vanishingly few randomized trials behind them, and where rigorous trials do exist, as in Sweden, the results are genuinely good. On the payment side, the same class of technology has been deployed at industrial scale against patients, and we know that only because courts are now compelling insurers to explain themselves while legislatures in dozens of states pass laws to constrain them.

The same capability that catches an aggressive tumor on a mammogram two years early can also generate a denial in 1.2 seconds, and the difference between those two outcomes has nothing to do with the sophistication of the algorithm. It comes down entirely to who owns it, what they are measured on, and whether anybody outside the building is positioned to look inside it. A gatekeeper nobody can audit is not a technology problem. It is an accountability problem wearing a technology costume, and until this year the only people who understood how these systems actually worked were the people being paid to keep them running.

As we often say on The Craig Bushon Show, we don’t just follow the headlines… we read between the lines to get to the bottom line of what’s really going on.


Disclaimer: This editorial is an opinion and analysis piece from the Craig Bushon Show Media Team. It is intended for educational and informational purposes and should not be interpreted as medical, legal, or financial advice. Readers should consult qualified healthcare professionals regarding individual medical decisions. Allegations described in pending litigation are allegations that have not been proven, and the companies named dispute them. AI technologies continue to evolve, and their capabilities, limitations, regulatory oversight, and clinical applications remain the subject of ongoing research and development.

 

Picture of Craig Bushon

Craig Bushon

Leave a Replay

Sign up for our Newsletter

Click edit button to change this text. Lorem ipsum dolor sit amet, consectetur adipiscing elit