The argument — stated once, then earned
In September 2016, MD Anderson switched off an IBM Watson system it had spent $62 million teaching to recommend cancer treatments. The Oncology Expert Advisor had never advised on the treatment of a single patient outside a test. It did not fail for lack of capability. It failed because they had not built the path from the system to the actual practice of medicine.
The consensus explanation for the enterprise AI disappointment — the models aren't good enough, the data isn't ready — describes a minority of the wreckage. The majority of pilots that produce nothing are pilots where the technology worked. The demo was crisp. The accuracy target was hit. Some even shipped, ran in production for a year, and still created no tangible value.
Getting the right answer clears one gate out of four. The system still has to be pointed at a problem that is financially worth solving. It has to land inside a workflow, in front of people willing to change how they work. And then comes the gate almost nobody plans for: the capacity it frees or the revenue it wins is only real when it shows up on the bottom line — and somebody has to put it there. That is the distance between technical success and value delivery. Value leaks at every one of these gates, and each leak traces back to an operating choice made, or not made, at pilot design time — usually months before a model is trained.
Everything that follows turns on what counts as a pilot and what counts as value — two terms management uses freely and rarely defines. A pilot here is any funded, bounded attempt to put AI in front of real work with the intention of expanding it if it works — a proof of conceptA short, small-scale build whose only job is to show that something can work: a sample of data, a handful of users, a few weeks. It establishes feasibility, not value, and it is usually run outside the process that would govern a real system., a limited production rollout, a departmental deployment. Value is something a controller can find: a cost line that fell, a revenue line that rose, a mission outcome measurably better, cash. Hours "saved" don't qualify, and neither do satisfaction scores or adoption numbers. Saved hours are an input to value — and, as the evidence below shows, they are where the trail goes cold in most enterprises.
Executive takeaways — the seven claims this brief proves
- 1The famous "95% of AI pilots fail" number is roughly right about the outcome and badly wrong about the cause — and it is not a measurement of pilots. MIT NANDA's own text — NANDA being a research project at the MIT Media Lab, and the study an interview-and-survey exercise rather than an audit of company accounts — defines success as tools "users or executives have remarked as causing a marked and sustained productivity and/or P&L impact," from 52 interviews and 153 conference surveys, and warns its figures are "directionally accurate based on individual interviews rather than official company reporting." Moderate confidence
- 2The dominant failure mode is capacity that is freed and never banked. In the largest linked survey-to-payroll study available — one that sets what workers say they do beside what their employers actually paid them — 85% of chatbot users report reallocating their time savings to other tasks, and earnings and hours across 25,000 Danish workers show precise null effects: a result sharp enough to place the true effect at about zero, rather than a failure to detect one. It rules out changes larger than 2%. Time saved is not money saved unless someone converts it. High confidence
- 3Self-reported productivity is systematically, measurably wrong — in the optimistic direction. In METR's randomised trial, experienced developers were 19% slower with AI tools while estimating afterwards that they had been 20% faster — a 39-point gap between the felt effect and the measured one, inside the same people, on the same tasks. METR has since said the sign of that result should not be carried forward; the gap should. At the firm level the Atlanta Fed found the same shape in ~750 CFOs: reported AI-related productivity gains of 2.4 percentage points in 2025 against 1.0 implied — "implied" meaning what those same firms' own reported inputs and outputs work out to. Pilot business cases built on self-report inherit that gap at both levels. High confidence
- 4Buying beats building roughly two to one, and the gap is about maintenance, not talent. External partnerships reached deployment ~67% of the time against ~33% for internal builds in the NANDA sample, with employee usage rates nearly double for externally built tools. Moderate confidence
- 5A pilot's most expensive failures are silent ones that survive in production. Epic's sepsis model — live across hundreds of US hospitals — scored an AUC of 0.63 in external validation at Michigan, and in a 2023 two-hospital emergency-department cohort of 145,885 encounters returned a sensitivity of 14.7% and a median warning lead time of zero minutes. Three words are carrying that sentence. AUC scores how cleanly a model separates the patients who will deteriorate from those who will not, on a scale where a coin flip scores a half and a perfect separator scores one; sensitivity is the share of real cases the model actually flags; lead time is how far ahead of the event the warning arrives. External validation means somebody other than the vendor tested it, on a different hospital's patients. Nothing about that is visible on a dashboard. High confidence
- 6Regulators and courts now price the operating gap directly, which converts a soft governance failure into a hard cash one. A tribunal held Air Canada to its chatbot's answer and called the argument that the bot was a separate legal entity "a remarkable submission"; the SEC ordered Presto Automation to cease and desist after finding that its "human-free" voice AI needed an off-site human on 70% of orders at its best pilot sites and on 100% of orders at the substantial majority of locations. High confidence
- 7The programs that work are boring, narrow and owned — and they look nothing like the portfolio most enterprises are running. The Permanente Medical Group's ambient scribes — software that listens to a clinical visit and drafts the note the doctor would otherwise type up afterwards — reached 2.5 million uses and ~16,000 documentation hours in one year by attacking the single task clinicians most hated, with no claim of headcount reduction attached. Adoption concentrated where the pain was worst. High confidence
PART I — THE PILOTS
Fifteen sections, cases first and pattern after. Start from the top and you get a list of technology complaints; start from the cases and you get a list of operating mistakes.
1 · A $62 million cancer engine that never saw a patient — MD Anderson & IBM Watson, 2013–2016
The Oncology Expert Advisor was supposed to read a patient's chart, read the literature, and tell an oncologist what to do next. MD Anderson had been running it with IBM since 2014. By the time it stopped, the centre had, in the words of the University of Texas System audit that surfaced in February 2017, "sent US$62 million in the general direction of Big Blue and PwC" — without going through its normal IT procurement process.
The audit did not find what the headlines said it found. It stayed deliberately out of the science. Its findings were about governance: work performed under an amended scope that "extended beyond the OEA project and intent of funding as approved by the Board of Regents"; invoices "paid in full regardless of whether contracted services were delivered as agreed upon"; much of the money spent without competitive tender, with fees "consistently set just below the amount that would have required Board approval"; and $11.59 million of donor gifts spent before they had been received.
The system had been built against ClinicStation, the medical records system MD Anderson used before it moved to EpicThe dominant electronic health record system in US hospitals — the software holding the digital chart, the orders and the notes. Every clinical action passes through it, so whether a tool is built inside the record or beside it largely decides whether a clinician ever sees its output.. It was never rebuilt against Epic. Audit staff told the reviewers that OEA's drug protocol and clinical trial data "is now outdated and must be updated before OEA can be piloted again within MD Anderson." IBM ended support in September 2016. The system had never been piloted anywhere else.
So: a decision-support tool whose entire selling point was up-to-date medical knowledge, sitting on a record system the institution no longer used, holding drug protocols that had gone stale. Whether the underlying model was good is, at that point, an academic question. The value was never reachable. The organisation moved and the pilot did not, because nothing in the pilot's design made anyone responsible for the pilot moving with it.
The epilogue is corporate. IBM had assembled Watson Health through several billion dollars of acquisitions; in June 2022 Francisco Partners completed the purchase of the healthcare data and analytics assets and relaunched them as a standalone company, Merative, headquartered in Ann Arbor. The flagship demonstration of AI in medicine did not so much fail as get sold, which is a different and quieter kind of ending — and one that leaves no post-mortem behind.
MD Anderson was not the whole of it, and the two failures are different in kind. Watson for Oncology was sold to hospitals well beyond Houston, and in July 2018 internal IBM presentations obtained by STAT, decks prepared in June and July 2017 by the division's own deputy chief health officer and circulated to Watson Health management, recorded "multiple examples of unsafe and incorrect treatment recommendations." Customers described the output as "often inaccurate," raising "serious questions about the process for building content and the underlying technology." The documents name the mechanism plainly: the system had been drilled on a small number of synthetic cancer cases — hypothetical patients — rather than on real patient data, with recommendations derived from a handful of specialists per cancer type rather than from guidelines or evidence.
At MD Anderson the model's quality was never the binding constraint; the integration path was. Everywhere else, the constraint was that the evaluation set bore no resemblance to the population. One is a Gate-3 failureA failure of workflow integration. The model works, but its output never reaches the right person at the right moment in a form they will act on. The next section sets out all four gates a pilot must pass; this is the third. and the other a Gate-2 failureA failure of model capability — the system does not do the job accurately or reliably enough under real conditions. This is the only one of the four gates the AI industry sells tooling for. wearing a Gate-2 evaluation that could not have caught it — which is the vendor-demo mechanism of §10, four years early and inside the most famous AI programme of its decade.
- 01AmbitionA Watson-powered advisor reading charts and literature to recommend oncology treatment and match trials.
- 02SetupProcured outside the normal IT process; much of the spend untendered; fees kept just under Board-approval thresholds; scope amended beyond the funding the Board had approved.
- 03What happened$62m to IBM and PwC combined. Built on ClinicStation; the institution moved to Epic. IBM ended support in 2016; the project was terminated that September.
- 04Why value never landedNever in clinical use; never piloted outside MD Anderson; protocol and trial data stale before restart was possible.
- 05Transferable root causeThe pilot was funded as a science project and governed as a gift, so no one owned the integration path the value ran through.
2 · Four gates, one leak — the frame the rest of the cases fill in
Strip the vocabulary away and every enterprise AI pilot has to pass through four gates in order. Miss any one and the money does not arrive, regardless of how well the others went.
Gate 1 — is there value on the table? Is the problem worth solving, is the value ownable by this organisation, and would anyone pay for the solved version? A great many pilots die here and the death certificate says something else, because Gate 1 is assessed before any measurement exists.
Gate 2 — does the model work? Accuracy, latency, cost per call, reliability under the real distribution rather than the sampled one. This is the only gate the AI industry sells tooling for, and it is the gate enterprises are best at. It is also the gate that matters least often.
Gate 3 — does the work change? Does the output land inside a process, in front of a person with the authority to act on it, at the moment the decision is made, in a form that survives the person's judgment about whether to trust it? Deployment, on its own, is none of this — it only makes it possible.
Gate 4 — does anyone collect? Freed capacity has to be converted — into headcount not hired, hours redeployed to revenue work, a service level raised that a customer pays for, a loss avoided. Someone must own that conversion, and it is almost never the same person who owned the pilot.
"The tech isn't ready" survives as the standard explanation because it is the only one of the four a technologist can diagnose alone. The other three require somebody to look at the P&L and the org chart, and those are not in the pilot team's job description.
- 01Value on the tableIs the problem worth money, is the value ownable, and is the counterfactual (do nothing / buy it / fix the process) worse? Assessed before any measurement exists, which is why it is usually skipped.Rarely managed
- 02The model worksAccuracy, latency, unit cost, reliability under the production distribution. The one gate with a mature tooling market and a clear owner.Well managed
- 03The work changesOutput reaches a decision-maker at the decision moment, in a form they trust and act on. Requires process redesign, not deployment.Partly managed
- 04Someone collectsFreed hours, avoided losses or won revenue are converted into a P&L line by a named owner with the authority to do so. Nearly always unowned.Rarely managed
3 · The deflection trap — Commonwealth Bank and Klarna, 2024–2025
A tool can be logged in by everyone and used by nobody, or used hard and still leave the work where it was. The two programs here are the ones most 2026 pilot budgets resemble — a generative assistant put in front of customer contact and measured on how many conversations it keeps away from a human — and both published their numbers, made the headcount decision, and reversed it in public.
Commonwealth Bank of Australia, July–August 2025. CBA announced 45 redundancies in its call centre, citing technology including AI: "Our investment in technology, including AI, is making it easier and faster for customers to get help, especially in our call centres." The Finance Sector Union raised a dispute at the Fair Work Commission, and its members reported that call volumes had risen after the voice-bot went in, with management offering overtime and pulling team leaders onto the phones. On 21 August the bank reversed the redundancies and called the decision an "error", admitting it "did not adequately consider all relevant business considerations" and "should have been more thorough in our assessment of the roles required." Affected staff were offered their jobs back. This from a bank that had just posted a record A$10.25 billion cash profit — the money was never the constraint.
Nothing about that outcome is rare, and the arithmetic behind it runs backwards from intuition. A deflection botA contact-centre assistant whose measured job is to stop calls and chats from reaching a human. Its headline metric — the deflection rate — counts contacts intercepted, not problems solved, which is why it can rise while the work left for people gets harder and larger. takes the simple calls, which raises the average difficulty of everything still reaching a human. And when it handles a call badly, it generates a second one. Deflection rate goes up. Human workload does not go down proportionally, and may go up. The bank measured the first number and made a headcount decision on it.
Klarna, February 2024 to 2025. The most-cited enterprise GenAIShort for generative AI: systems that produce new text, images, code or speech rather than classifying or scoring something that already exists. The large language model behind a chat assistant is the common enterprise case. The distinction matters because generative systems fail differently — an older model gives a wrong score, a generative one writes a fluent wrong answer. success of its year. Klarna's own release of 27 February 2024 reported the OpenAI-powered assistant handling 2.3 million conversations in month one — two-thirds of all customer-service chats — "doing the equivalent work of 700 full-time agents," resolving errands in under two minutes against eleven previously, with a 25% drop in repeat inquiries and an estimated "$40 million USD in profit improvement to Klarna in 2024." Every number there is Klarna's own; none was independently audited.
By 2025 the company was recruiting humans again, with CEO Sebastian Siemiatkowski settling on a rule his 2024 self would not have written: "From a brand perspective, a company perspective… I just think it's so critical that you are clear to your customer that there will be always a human if you want." Klarna is now building an "Uber-type" flexible agent pool. This is not a failure — the automation of routine contacts held. What failed was the inference from "the bot handles two-thirds of chats" to "we need proportionally fewer people," which is the same inference CBA made.
One inference cuts across both programs. Deflection metrics and value metrics point in opposite directions more often than anyone admits, because deflecting the easy half of a workload raises the cost per remaining unit and multiplies the failure cost of the deflections that go wrong. An assistant that handles 70% of contacts may reduce total cost by far less than 70%, by nothing, or — where it generates repeat contact — by a negative number. CBA is the clean natural experiment: same firm, same quarter, the deflection number rose and the workload rose with it.
4 · The drive-thru lesson — McDonald’s and Taco Bell, 2021–2026
Voice ordering is the most public pilot program in America — it runs outdoors, in front of customers holding phones, so failures are filmed and retreats are announced. Two of the largest quick-service operators ran it at hundreds of sites; what each did next is the useful part.
Taco Bell, 2023–2025. Voice AI ordering rolled out to more than 500 drive-thrus on Yum Brands' Byte platform. Then the viral clips: an order for 18,000 cups of water among them. Chief digital and technology officer Dane Mathews told The Wall Street Journal in August 2025 that the company was in an "active conversation" about where to use it and where not, and the resolution is the interesting part: the company segmented instead of withdrawing. "For our teams, we'll help coach them: at your restaurant, at these times, we recommend you use voice AI or recommend that you actually really monitor voice AI and jump in as necessary." Busy restaurants with long lines may keep a human on the mic.
McDonald's, 2021–2026. Two years of automated order taking with IBM across more than 100 restaurants, ended by a system message from chief restaurant officer Mason Smoot in June 2024: "While there have been successes to date, we feel there is an opportunity to explore voice ordering solutions more broadly… the technology will be shut off in all restaurants currently testing it no later than July 26, 2024." McDonald's declined to say whether the test had succeeded or by what metric it was judged — a small silence that says a lot about how the pilot was scoped. The company nonetheless said the work "has given us the confidence that a voice-ordering solution for drive-thru will be part of our restaurants' future."
Two years later that confidence has a name. In June 2026 McDonald's began testing ArchIQ, with a voice assistant nicknamed "Archy," at five US locations on its Google Cloud partnership — five, against roughly 13,600 US restaurants. The scope collapse from "more than 100" to "five" is the most honest number in the entire story.
| Program | Dates | Benefit assumed at design | What was observed | Resolution |
|---|---|---|---|---|
| CBA voice-bot Retail banking, AU | Jul–Aug 2025 | Fewer inbound calls → 45 roles redundant | Call volumes rose; overtime offered; team leaders on phones | Redundancies reversed; bank calls it an "error" |
| Klarna AI assistant Fintech, SE | Feb 2024 → | 2.3m chats, "work of 700 agents", $40m profit uplift (company estimate) | Routine automation held; complex and emotive cases degraded | Human agents re-recruited; guaranteed human path |
| Taco Bell voice AI QSR, US | 2023–2025 | Faster service, better accuracy, automated upsell at 500+ sites | Order errors, interruptions, viral prank orders | Segmented by site and daypart; humans retained at busy stores |
| McDonald's AOT (IBM) Quick service, US | 2021–Jul 2024 | Operational savings and speed of service across 100+ sites | Not disclosed; company declined to state the success metric | Ended; relaunched Jun 2026 as ArchIQ at 5 of ~13,600 US sites |
5 · The chatbot that made a contract — Moffatt v. Air Canada, 2024 BCCRT 149
In November 2022 Jake Moffatt's grandmother died and he went to Air Canada's website to book a flight to Toronto. He asked the site's chatbot about bereavement fares. It told him he could book at full price and apply for the bereavement rate within 90 days. The airline's actual policy, on a page titled "Bereavement travel" elsewhere on the same site, said the opposite: no retroactive claims. He paid $1,630.36, applied, and was refused.
Air Canada's defence at the British Columbia Civil Resolution Tribunal is the reason this case matters. Tribunal member Christopher C. Rivers recorded it plainly: the airline argued it could not be held liable for information provided by "one of its agents, servants, or representatives — including a chatbot," and, as Rivers put it, "in effect, Air Canada suggests the chatbot is a separate legal entity that is responsible for its own actions. This is a remarkable submission."
The tribunal answered in a few plain sentences that are, for anyone deploying a customer-facing assistant, load-bearing. "While a chatbot has an interactive component, it is still just a part of Air Canada's website. It should be obvious to Air Canada that it is responsible for all the information on its website. It makes no difference whether the information comes from a static page or a chatbot." On the argument that Moffatt should have checked the other page: the airline "does not explain why the webpage titled 'Bereavement travel' was inherently more trustworthy than its chatbot… nor why customers should have to double-check information found in one part of its website on another part of its website." Damages of $650.88, plus $36.14 interest and $125 in fees. A rounding error, and entirely beside the point.
Everyone already knew a chatbot can hallucinateProduce a confident, fluent, well-formed answer that is simply invented. It is not a bug in the ordinary sense — the system is doing what it was built to do, which is generate plausible text, and plausibility and truth are not the same target. The dangerous property is that a hallucination looks exactly like a correct answer.. What the tribunal established is that the enterprise's own statement of what the assistant is — a helper, a convenience, an experiment, explicitly not authoritative — has no legal weight against a customer who reasonably relied on it. Every disclaimer in a deployment plan is an internal document. The counterparty gets to treat the output as the company speaking, because it is.
Price that. A conversational assistant deployed against a policy surfaceAny body of rules a company publishes and is expected to honour — fares, refunds, warranties, eligibility, pricing. When an assistant is pointed at one, it stops being a search box and starts being a mouth: whatever it says about the rule is what the company has said about the rule. is not a productivity tool with a small error rate; it is an unbounded, unmonitored authority to make representations on the company's behalf, at machine volume, without a review step. Most pilot business cases model the cost of a wrong answer as a customer-satisfaction ding. The correct model is the expected cost of a promise the company must honour, multiplied by conversation volume. Change that one line and a meaningful share of customer-facing assistant pilots stop clearing their hurdle rateThe minimum return a company requires before it will spend money on something. Every investment competes against it; a project that does not clear the hurdle is refused not because it loses money but because the money does better elsewhere. — which is the honest reason many of them were quietly narrowed after February 2024.
In January 2024 the parcel firm DPD disabled part of a chatbot it had run "successfully for a number of years" after a system update caused it to swear at a customer and compose a poem about how terrible the company was; the customer's account was viewed 800,000 times in a day. The failure was introduced by an update, not by the original design. The risk is not retired at launch, and a system that has behaved for years is not thereby safe.
The more serious version came in June 2023, when the National Eating Disorders Association took down its Tessa chatbot after users showed a newer version dispensing dieting and calorie-counting advice to a population for whom that advice is precisely the hazard. A helpline replacement that gives contraindicatedMedically the wrong thing to give this patient — advice or treatment that is fine in general but actively harmful for the condition in front of you. guidance has not underperformed its target; it has inverted its purpose. Any pilot whose failure mode is the opposite of its objective — rather than merely less of it — needs a different risk model than a productivity tool, and almost none get one.
6 · The pattern predates the technology — two deep-learning deployments, 2018–2024
Every case so far is generative AI, and every one is recent. The two in this section are neither — deep-learning classifiers designed and deployed before anyone had typed a prompt into ChatGPT — and they are here on purpose. The most comfortable rebuttal is that the failures above are teething problems of a young technology, cured by the next model generation. These two deployments are the control group for that claim. The technology was different, the vendors were different, the decade was different, and the leak was the same.
Google Health in Thailand, 2018–2019. Thailand's ministry of health had a target: screen 60% of diabetic patients for diabetic retinopathy, the leading cause of preventable blindness in that population. It had roughly 4.5 million patients and about 200 retinal specialists. The existing process — nurses photograph the eye, a specialist elsewhere reads the image — could take ten weeks. Google Health had a deep-learning systemA model trained by being shown a very large number of labelled examples — here, retinal photographs a specialist had already graded — until it learns the pattern for itself. Nobody writes the rules; the rules are inferred from the examples, which is why the examples the system was shown decide what it can handle. that identified the condition from a retinal scan at better than 90% accuracy, which the team characterised as human-specialist level, and could return a result in under ten minutes.
This is as clean a value hypothesis as enterprise AI ever gets. A real bottleneck, a quantified queue, a model that beats the queue on both accuracy and latency. The team deployed to eleven clinics and then did the thing almost nobody does — they watched.
The model had been trained on high-quality scans and, to protect its accuracy, was designed to refuse images below a quality thresholdThe cut-off an engineer sets on a model's own confidence score, deciding when it acts, warns, or declines. It is a dial, not a property of the model: turn it one way and you get more misses, the other way and you get more false alarms. Somebody chooses where it sits, and that choice is where a technical decision quietly becomes an operating one.. Nurses were photographing dozens of patients an hour in rooms with poor lighting. More than a fifth of images were rejected outright. A rejected image meant the patient was told to visit a specialist at another clinic on another day — the exact outcome the system existed to prevent — and nurses, who frequently believed the rejected scans showed no disease, burned time trying to retake or edit them. Because processing ran in the cloud, clinics with slow connections queued. One nurse: "Patients like the instant results, but the internet is slow and patients then complain. They've been waiting here since 6 a.m., and for the first two hours we could only screen 10 patients."
The team's own write-up, presented at CHI 2020, names the mechanism without euphemism: the study set out to characterise the "socio-environmental factors that impact model performance, nursing workflows, and patient experience," and found "tensions between the model's thresholds for data quality, and the quality of data that arise from an imperfect, resource-constrained environment." A threshold is a specification. An imperfect, resource-constrained environment is every place work actually happens.
The quality threshold is a decision. Engineers optimising for the metric they were graded on converted an accuracy risk into a throughput cost, and pushed that cost onto a nurse and a patient in a room they had never stood in. It is a perfectly defensible choice at Gate 2 and a catastrophic one at Gate 3. Nothing about it shows up in a validation report.
The Epic Sepsis Model, 2018–2024. Of every case in this gallery, this is the one that most completely defeats the "pilots fail to reach production" framing. It reached production. It reached production at hundreds of American hospitals, embedded in the dominant electronic health record, firing alerts into live clinical workflow for years. By the pilot-to-production metric it is a triumph.
In 2021, researchers at the University of Michigan externally validated it across 38,455 hospitalisations and reported an area under the curve of 0.63 — against the 0.76 to 0.83 range cited in the vendor's own documentation — with sensitivity of 33% and positive predictive valueOf all the times the system raises an alarm, the share that turn out to be real. It is the number the person receiving the alert actually experiences — a low value means most alarms are false, which is how a technically defensible model trains the staff around it to ignore it. of 12%. Two-thirds of sepsis cases missed; roughly eight false alarms for every true one.
The 2024 replication is worse, and more useful. Ostermayer and colleagues examined 145,885 emergency-department encounters across two county hospitals through 2023, alerting at Epic's recommended threshold of 6. Sensitivity: 14.7%. Positive predictive value: 7.6%. And the number that ends the argument — the median lead time on the alert was zero minutes (80% CI, −6h42m to +12h00m). The model, at the median, told clinicians a patient was becoming septic at the moment the patient was already recorded as septic.
Sensitivity, AUC and lead time are not measuring the same thing, and the distinction runs through the rest of this brief. Sensitivity and AUC measure discrimination — how well a model tells the sick apart from the well, judged on a pile of finished cases after the fact. Lead time measures decision value — whether the answer arrives while there is still something a person can do about it. A model can be respectable on the first and worthless on the second.
Start with the objection, because it is a good one: Epic fixed it. On 27 February 2026 a multicentre prospective validation of the updated model ran across four US health systems and 227,091 inpatient encounters and reported encounter-level AUROC of 0.82 to 0.92 — inside the range the first version was only ever claimed to hit — with median lead time before sepsis onset of 1.9 to 10.3 hours depending on the institution. That is a real improvement and it should be said plainly. Three details in the same study qualify it. The positive predictive value at a 60% sensitivity threshold still ran only 0.13 to 0.26, so between 21 and 35 alerts were needed per detected case; the spread across four sites was wide enough that the authors recommend "local validation prior to deployment"; and the study was, in their words, "not designed to evaluate its impact on clinical outcomes, such as mortality." A better discrimination statistic is not yet a better patient outcome.
What the correction does not do is rescue the deployment. Version one ran live in hundreds of hospitals for roughly six years. What ended it was not the vendor's monitoring and not a customer's: it was two academic external validations, published five years apart, by people with no commercial relationship to the alert. Every hospital paying for the model had the data to compute what Michigan computed. None did. That is the finding — the fix arrived through the literature because no one in the operating chain owned the measurement.
The second point is the obvious one. A model can be technically live and clinically inert, and the difference is invisible to everyone except a researcher who goes looking. Alert volume goes up, which reads as engagement. Nobody's dashboard says "median lead time: zero."
The third is less comfortable. Engineering did not fail here. Sensitivity of 14.7% with specificityThe mirror image of sensitivity: the share of the people who were not going to get sick that the model correctly leaves alone. High specificity means few false alarms as a fraction of the healthy — which can still be a great many alarms in absolute terms, because the healthy vastly outnumber the sick. of 95.3% is a model doing exactly what a low-signal prediction problemA problem where the data available simply does not contain much advance warning of the thing you want to predict. No modelling technique conjures signal that is not there, so the ceiling on performance is set by the world, not by the engineering — and the honest response is to ask whether a prediction is the right instrument at all. allows, deployed at a threshold chosen to keep alert volume tolerable. The organisation traded away almost all of the model's sensitivity to buy alert-fatigue relief, and then kept the alert on. That trade was made by clinical informatics staff balancing two operational pressures. It was rational locally and value-destroying globally, and no AI capability at any price would have changed it.
7 · The economics nobody priced — cost, latency, and the human in the loop
Amazon's Just Walk Out was the most technically impressive retail deployment of its decade: walk in, take things, walk out, get charged. In 2024 Amazon began removing it from Amazon Fresh grocery stores in the US, replacing it with Dash Carts — a smart trolley that asks the shopper to scan. Reporting at the time described roughly a thousand associates in India reviewing shopping sessions. Amazon disputes the characterisation: spokesperson Jessica Martin Strauss said "the characterization that Just Walk Out technology relies on human reviewers is inaccurate," that associates' primary role is annotating video to improve the model, and that they "may also validate a small minority of shopping visits where our computer vision technology cannot determine with complete confidence an individual's purchases."
Take Amazon's account at face value; the conclusion barely changes. Grocery baskets are large, long-dwell and visually cluttered — the hardest possible case for the technology and the lowest-value one, because a supermarket checkout is already cheap per basket. Just Walk Out was not a failure. It was pointed at the wrong basket size, and it remains deployed in stadiums, airports and convenience formats where the basket is three items and the alternative is a queue nobody has time for. That is a Gate-1 error corrected late, not a Gate-2 defect.
Where the human-in-the-loopA person kept inside the automated process to check, correct or complete what the machine produces. Sold as scaffolding that comes down once the model improves; in practice it is usually a permanent staffed tier, and its size is the honest measure of how much automation was actually achieved. economics were genuinely concealed, the SEC put it on the record. On 14 January 2025 the Commission instituted settled cease-and-desist proceedings against Presto Automation, a Nasdaq-listed restaurant technology company, over its Presto Voice drive-thru product. Two findings matter here.
First, from November 2021 to September 2022, every commercially deployed Presto Voice unit ran on a third party's speech technology while Presto described it in Commission filings as "our" and "Presto's" technology; the company had not begun building its own until early 2022. Second, and more consequential, when Presto did deploy its own AI it claimed the product "eliminat[es] human order taking." The order finds that the proprietary units "lacked the capability to take orders on their own and required substantial human involvement" — Presto "hired, trained, and supervised human order takers located abroad (primarily in the Philippines and India), who processed the vast majority of drive-thru orders." The original version was designed to depend on that support. The later version "required a human agent to enter the orders approximately 70% of the time."
That 70% is the figure everyone quotes, and quoting it alone is too kind. The order records what Presto itself had to concede on 14 December 2023: the 70% "referred to orders at the few locations where the most advanced version of Presto Voice was being piloted," and "human agent intervention was required on 100% of orders at the substantial majority of locations where the original version of Presto Voice units were installed." The same disclosure put the average non-intervention rateThe share of transactions the system completed with no human touching them — the mirror of the intervention rate. Vendors quote the flattering one, and neither number means anything until you know which transactions, at which sites, over which period, are in the bottom of the fraction. across all restaurants running Presto's own technology at 85% — so a 15% intervention rate — while the company had told investors it achieved "automated order completion" rates of 95% to 99%. Four numbers, four denominators, one product. Read Presto's Form 10-K for the fiscal year ended 30 June 2023 and the mechanism is visible in the qualifier: the filing claims "approximately 95% non-intervention rate at certain locations." Three words are carrying the entire representation.
Presto's dishonesty is the least portable part of this. An automation rate is meaningless without its denominator, and a vendor asked for "the intervention rate" will answer with the best-scoped one available. The buyer's job is to specify the population before the number is quoted: all units, all locations, last calendar month.
The Commission ordered Presto to cease and desistA regulator's order to stop the conduct and not repeat it. Where the proceedings are settled, the company accepts the order without contesting the findings — neither an admission nor a trial — and, as here, it can arrive with no fine attached at all. from violations of Securities Act §17(a)(2) and Exchange Act §13(a) and Rules 13a-11 and 13a-15(a). No civil penalty was imposed. Presto had been delisted from Nasdaq the previous September.
A third pattern kills pilots on cost alone, and Gartner put a number on it. Its July 2024 analysis — the source of the widely repeated forecast that at least 30% of GenAI projects would be abandoned after proof of concept by the end of 2025, "due to poor data quality, inadequate risk controls, escalating costs or unclear business value" — carried a cost range for business-model-innovation deployments of $5 million to $20 million. Rita Sallam's framing was blunt: "executives are impatient to see returns on GenAI investments, yet organizations are struggling to prove and realize value. As the scope of initiatives widen, the financial burden of developing and deploying GenAI models is increasingly felt."
Falling model prices do not fix this, and may make it worse. When inferenceThe act of running a trained model — one question in, one answer out, billed per use, as distinct from the one-off cost of training it. It is the line everyone watches, because it is the line that has been collapsing in price. gets cheaper, the binding constraint moves to the things that did not get cheaper — integration, evaluation, monitoring, escalation handling, the human review tier, the legal review of what the system is allowed to say. Those are labour and organisational costs, and they scale with the number of workflows touched rather than with tokens. A pilot whose economics only work if the model becomes free is a pilot that has misidentified its own cost structure.
8 · Scaled before it was right — Zillow Offers, 2018–2022
Zillow Offers is the most expensive case in this brief and the one most often mis-told. It is usually filed as "the algorithm was wrong." Zillow's own filings say something more precise and more useful.
On 2 November 2021 the company announced it would wind down the business. Rich Barton: "We've determined the unpredictability in forecasting home prices far exceeds what we anticipated and continuing to scale Zillow Offers would result in too much earnings and balance-sheet volatility." That quarter carried "a write-down of inventory of approximately $304 million within the Homes segment as a result of purchasing homes in Q3 at higher prices than the company's current estimates of future selling prices," with a further $240–265 million of losses expected in Q4. The Homes segment lost $421.6 million before tax in the quarter; consolidated net loss was $328.2 million. The wind-down would cut roughly 25% of Zillow's workforce.
The FY2021 Form 10-K totals it: inventory write-downs of $407.9 million for the year, $71.2 million of impairment and restructuring inside the Homes segment, $6.2 million winding down the financing facilities, $4.9 million of accelerated depreciation. The board's stated reasoning is a value judgment, not a technical one: the decision was made "in light of home pricing unpredictability, capacity constraints and other operational challenges… all of which led us to conclude that, despite its initial promise in earlier quarters, Zillow Offers was unlikely to be a sufficiently stable line of business to meet our goals and needs going forward."
And the risk factor, written by Zillow's own lawyers, is the cleanest statement of a Gate-2 limit anywhere in this brief: "We underwrite and price the homes we buy and sell through Zillow Offers using in-person evaluations and data science and proprietary algorithms… These assessments may be inaccurate… Our pricing model may not account for submarket nuances — for example, the location of a home on a hill or in a building — which could have a significant impact on price."
Two clarifications, because this case is easy to over-claim.
Zillow Offers was not a pilot by 2021; it was an operating business with a balance sheet. It earns its place here because of what the pilot phase did not establish before scaling — the width of the prediction intervalThe range a model's forecast is actually likely to fall in, rather than the single number it prints. A model that says "$400,000" is really saying "somewhere around $400,000, probably within this much" — and that this much, sometimes quoted as the standard error, is the number that decides whether you can trade on the forecast. A narrow interval is a business; a wide one is a bet. relative to the gross margin on a flip. An iBuyerA firm that buys homes directly from sellers for cash, on an algorithm's price estimate, then resells them. The seller trades price for speed and certainty; the iBuyer takes the house — and the risk of having valued it wrong — onto its own balance sheet.'s economics are a spread trade: buy at estimate minus fee, sell at market, keep the difference less holding and renovation cost. If the model's standard error on a home is a few percent and the spread is a few percent, the business is a coin flip with leverage, and no amount of model improvement inside the achievable range changes that. That is knowable at pilot scale, from the residual distributionThe spread of a model's misses — every past estimate set against what the house actually sold for, collected into a shape. It shows not just how wrong the model is on average but how wrong it can get on a bad one, which is the only version of the question a balance sheet cares about., before a single house is bought at volume.
The market moved violently in 2021 and Zillow was caught longHolding a large amount of an asset when its price falls. Being long is not itself a mistake — you cannot flip houses without owning houses — but the size of the holding decides how much a wrong price costs you.. It is fair to say bad luck contributed. It is not fair to say bad luck explains it, because the company's own explanation is that the variance exceeded expectation — which is a statement about the model's calibration, not about the draw. Calibration is a different property from accuracy, and here the distinction is the whole case. An accurate model gets the number close; a calibrated model knows how close, and can say when it is guessing. A model can be accurate on average and badly calibrated at the same time, which is the dangerous combination, because it hands over its shakiest estimates in exactly the same confident voice as its best ones. A well-calibrated pricing model that knows its own uncertainty produces a lower offer, buys fewer homes, and loses less. Opendoor ran a comparable book through the same market and did not exit. The differentiating variable was position sizingDeciding how much to stake on a view, given how sure you are of it. It is separate from being right: the same forecast, held in twice the size, loses twice the money. In an algorithmic business it is the executive decision that converts a modelling limitation into a financial one, or doesn't. against model uncertainty, which is a risk-management choice made by executives.
The question a pilot must answer is not "how accurate is the model" but "how does the model's error distribution compare to the margin it is being asked to protect." An error rate that is excellent for a recommendation engine is catastrophic for a balance sheet. Most pilot readouts report the first number and never compute the second.
9 · The governance kill — and why one of these is a success
One case in this section is filed everywhere as a famous AI failure and is, on the evidence, the opposite.
Amazon began building an automated résumé-ranking system in 2014, scoring candidates one to five stars. Trained on a decade of the company's own résumés in a male-dominated field, it learned to prefer men — penalising résumés containing the word "women's" and the names of certain all-women colleges. Amazon neutralised those specific terms and then killed the project anyway, because it "lost confidence that the program was indeed gender neutral in all other areas." Amazon's own position, in the reporting that broke the story, is that the tool "was never used by Amazon recruiters to evaluate candidates" — though the company did not deny that recruiters looked at its recommendations, which is a narrower denial than it first appears and is the reason this case is filed as a near-miss rather than a clean abstention.
An organisation detected a defect it could have papered over, tried the paper-over, judged the residual risk unquantifiable, and shut the program down before deployment. That is exactly what a functioning control environment does. It cost Amazon four years of engineering and saved it what the alternative cost others. The reason it is remembered as a failure is that our vocabulary has no word for a pilot that correctly kills itself — and that vocabulary gap is itself a cause of failure elsewhere, because a leader who reads Amazon's outcome as humiliation learns to keep quiet instead.
Now the ones that did deploy. In September 2023 the EEOCThe US Equal Employment Opportunity Commission — the federal agency that enforces the laws against discrimination in hiring, pay and promotion. It can investigate, sue on a worker's behalf, and settle on terms that bind an employer for years. settled with iTutorGroup for $365,000: the company had "programmed their tutor application software to automatically reject female applicants aged 55 or older and male applicants aged 60 or older," rejecting more than 200 qualified US applicants. No learned model produced that bias; it was a hard-coded rule. The AI framing is almost incidental; what the case establishes is that automated screening logic is legally identical to a hiring manager's decision, with the aggravating feature of being written down.
Which is the bridge to Mobley v. Workday. Derek Mobley applied to more than 100 employers using Workday's platform and was rejected every time. On 16 May 2025, Judge Rita Lin of the Northern District of California granted preliminary certification of a nationwide ADEA collectiveA group action under the Age Discrimination in Employment Act, the US law protecting workers aged 40 and over. A "collective" is the age-discrimination cousin of a class action: people join it rather than being swept in automatically. Preliminary certification decides only that the group may proceed together — it settles nothing about whether the claim is right. covering applicants aged 40 and over denied recommendations through Workday's platform since 24 September 2020. The court had earlier declined to dismiss on the theory that Workday was merely a software supplier. The doctrine at stake is agency: a vendor performing a core hiring function on an employer's behalf can be treated as the employer's agent and held directly liable.
Two years on, the case has widened. The preliminary collective was scoped in July 2025 to applicants screened using Workday's HiredScore features, over the company's objection that HiredScore was a separate product. In January 2026 the court allowed three further named plaintiffs and new claims under Title VIIThe section of the US Civil Rights Act of 1964 that bans employment discrimination on race, colour, religion, sex and national origin. It is the main federal hook for hiring-discrimination claims outside age. and California's Fair Employment and Housing Act. In the order of 6 March 2026 Judge Lin dismissed the FEHA counts with leave to amend for want of a pleaded California nexus — and denied dismissal of everything else, including the ADEA disparate-impactA discrimination claim that works without any proof of intent: it is enough that a neutral-looking rule or system falls harder on a protected group and cannot be justified by the job. This is the doctrine that makes a screening algorithm legally interesting: intent is not the question, outcome is. claim, while granting AARP leave to file an amicus briefLiterally "friend of the court" — an argument filed by an outside organisation with no direct stake in the case, offering the judge a wider view of what the ruling would mean. Courts grant leave for one when they think the decision reaches beyond the two parties.. On 22 June 2026, ruling on the re-pleaded Third Amended Complaint, the court granted in part and denied in part, rejecting Workday's argument that California's anti-discrimination law cannot reach applicants screened from California for jobs elsewhere. The docket runs on. No merits determination has been made against Workday and none is assumed here.
What has been established in three years of procedure is narrower than the headlines and more useful. A screening vendor has not persuaded a court that it is merely a supplier; a re-pleading has not made the age-discrimination theory go away; and the location of the algorithm, not the location of the applicant, is now doing work in the jurisdictional analysis. For a buyer: where the model runs is becoming a compliance fact about your hiring process.
"We bought it, so the vendor carries the risk" is not how this is resolving, and almost no pilot budget models what follows. Both parties are exposed, and the employer's exposure is not reduced by the vendor's indemnity when the claim is statutory discrimination brought by an applicant. A screening pilot's true cost includes a bias-testing regime, a records regime, and a legal review — before go-live, not after the first demand letter.
Then the FTC. In December 2023 the Commission banned Rite Aid from using facial recognition for surveillance for five years, finding that from 2012 to 2020 the retailer deployed the technology across hundreds of stores without reasonable procedures, that the system falsely flagged consumers — disproportionately people of colour — as shoplifters, and that employees publicly accused them of wrongdoing, "sometimes in front of friends or family." The order requires Rite Aid to discontinue any such automated system "if it cannot control potential risks to consumers." Eight years of deployment, terminated by an enforcement action, with the risk-control obligation imposed retroactively at the worst possible moment.
And the state AGs are moving faster than the federal agencies on accuracy claims specifically. In September 2024 Texas settled with Pieces Technologies, whose generative summarisation product was in use at four major Texas hospitals receiving patient data in real time. Pieces had advertised a "severe hallucination rate" of "<1 per 100,000." The investigation found the metrics "were likely inaccurate and may have deceived hospitals about the accuracy and safety of the company's products." No money changed hands. What changed was the obligation: for five years, Pieces must clearly disclose the meaning and calculation method of any accuracy metric it advertises, or have an independent auditor substantiate the claim.
That remedy attacks the exact mechanism by which pilots are sold. A hallucination rate without a stated denominator, evaluation set and severity definition is not a measurement; it is a marketing number. Texas has now made publishing one, in that state, in healthcare, an enforceable representation.
10 · The vendor-demo mirage — when the product is the pitch
Builder.ai raised on the promise that AI would assemble software the way a pizza is assembled from toppings. It collapsed into insolvency in May 2025. The Register's account is unsentimental: more than $500 million raised from blue-chip investors including Microsoft and Qatar's sovereign fund, a business model in which "the Builder.ai team actually built the apps," and a prior history — as Engineer.ai — of press scrutiny over exactly that gap. The company had installed a new CEO in February 2025, months before the end.
Put Builder.ai next to Presto and the SEC's first AI-washing enforcement actions of March 2024 — Delphia ($225,000) for claiming it put "collective data to work to make our artificial intelligence smarter so it can predict which companies and trends are about to make it big," and Global Predictions ($175,000) for calling itself the "first regulated AI financial advisor" — and a pattern with a shape emerges. In each case the "AI" was, in whole or in part, either absent or a person. Gary Gensler's framing: "when new technologies come along, they can create buzz from investors as well as false claims by those purporting to use those new technologies."
The Delphia order documents something more damning than a marketing exaggeration. The Commission found that from at least August 2019 to August 2023 the firm claimed to use AI and machine learning to analyse retail clients' spending and social-media data "when, in fact, no such data was being used in its investment process" — and that after an examination, Delphia agreed in 2021 to correct the statements, made some corrective efforts, and then continued making false statements through August 2023. Two years elapsed between the regulator identifying the gap and the conduct stopping. A firm overselling AI is ordinary; an incentive strong enough to survive a direct regulatory intervention is the story.
The FTC added the consumer-facing version in September 2024. Under "Operation AI Comply" it brought five actions, including one against DoNotPay — "the world's first robot lawyer," which promised consumers could "sue for assault without a lawyer" and "generate perfectly valid legal documents in no time." The complaint alleges the company "did not conduct testing to determine whether its AI chatbot's output was equal to the level of a human lawyer, and that the company itself did not hire or retain any attorneys." The proposed order carried $193,000 and a requirement to notify subscribers from 2021–2023 about the service's limitations. The absence of testing is the finding worth carrying into a procurement conversation: the claim was not exaggerated relative to a measurement, it was made in the absence of one.
One word is doing most of the work in the newest version of this pitch, and a great deal now rides on it. An agent, in the enterprise sense, is a system handed a goal rather than a question: it plans its own steps, reaches into other software, and acts — issues the refund, files the ticket, updates the record — instead of handing text back to a person who then does those things. That is a genuinely different and harder thing to build than a chatbot. The label, unfortunately, is much easier to apply than the capability is to deliver.
Gartner gave the phenomenon its enterprise name and, in June 2025, a magnitude. Its analysis warned that over 40% of agentic AI projects will be cancelled by the end of 2027, "due to escalating costs, unclear business value or inadequate risk controls," and described "agent washing" — "the rebranding of existing products, such as AI assistants, robotic process automation (RPA) and chatbots, without substantial agentic capabilities." Gartner's estimate: only about 130 of the thousands of agentic AI vendors are real. Anushree Verma: "Most agentic AI propositions lack significant value or return on investment… Many use cases positioned as agentic today don't require agentic implementations."
A vendor's compensation event is a signed proof of concept, not a delivered outcome twelve months later — which is most of the explanation, and dishonesty is very little of it. A demo is optimised against the buyer's imagination; a production system is optimised against the buyer's edge cases. The two artefacts are built by different people to different specs, and only one of them is what the buyer evaluated. Every structural feature of enterprise procurement — the pilot budget that is smaller than the approval threshold, the six-week evaluation window, the reference customer selected by the vendor — pushes in the same direction.
The practical defence is cheap and almost never used: require the vendor to disclose the human-intervention rate, the escalation rate, and the evaluation set composition as contractual representations, and make the pilot's success criterion a metric computed on the buyer's data by the buyer. Presto's 70% would have surfaced in week two.
The honest vendors describe this problem in their own filings. C3.ai — which structures customer engagements as paid "Initial Production Deployments" before any subscription — warns in its Form 10-K for the fiscal year ended 30 April 2026 that "after completing an initial production deployment or trial customers do not have an obligation to continue to license our products, and we may not be able to convert initial production deployment customers into customers purchasing ongoing subscriptions." The pilot-to-production chasm is not only a buyer's complaint; it is a disclosed revenue risk on the seller's side, which tells you the conversion rate is low enough to be material.
- What percentage of transactions in your reference deployment required human intervention last month, and how is "intervention" defined? (Presto: ~70%, undisclosed until an SEC investigation.)
- Show the evaluation set your accuracy figure was computed on, its size, its provenance, and the severity definition behind any error-rate claim. (Pieces: "<1 per 100,000" with no stated method; now a five-year disclosure obligation in Texas.)
- Which components of the system do you own, and which are licensed from a third party under a contract you could lose? (Presto: all deployed units ran on a supplier's technology while filings called it "our" technology.)
11 · Success theatre — the pilot kept alive because killing it costs more
In October 2023 New York City launched an AI chatbot to help small businesses find their way through municipal regulation. In March 2024 The Markup tested it and found it advising landlords that they need not accept tenants using housing vouchers — illegal source-of-income discrimination in New York City — and telling employers they could take workers' tips.
It stayed up. For nearly two more years.
The New York City Comptroller's audit of the MyCity system found that the Office of Technology and Innovation had "very little else to show after spending more than four years and $100 million on the program." Users still could not apply for benefits through a single form. The portal "largely redirects users to pre-existing City websites that have been rebranded and redesigned as MyCity applications." On the chatbot specifically: it "appears to be unable to provide accurate or consistent information." The audit's overall conclusion — "MyCity was poorly managed from both a project management and contract oversight perspective" — was accompanied by the detail that OTI had requested a further $81 million in the 2026 budget to maintain what existed and add functionality.
The chatbot was finally shut down in early 2026, and the reason it was shut down is the point. A new administration took office; Mayor Zohran Mamdani called the bot "functionally unusable", said it was costing "around half a million dollars," and cut it as part of closing a budget gap. Reporting put the original build cost near $600,000. The bot's page now says its beta test has ended.
Nearly two years elapsed between public documentation that the system gave illegal advice and its termination. In that window the failure was known, cheap to fix by removal, and structurally protected — because the thing being protected was not the chatbot. It was the announcement. A launched initiative is a political asset whose value is realised at launch and whose costs accrue afterwards to whoever inherits it. Withdrawal converts a past success into a present admission. The rational move for the sponsor is to leave it running and let the budget carry it, which is precisely what happened until an actor with no stake in the original announcement arrived.
The corporate version runs on a shorter clock but the same logic. On 9 July 2024 the HR software firm Lattice announced it would become the first company to give "digital workers" official employee records — AI agents onboarded, given goals, performance metrics, system access and an accountable manager. The backlash from HR practitioners was immediate. On 12 July, three days later, CEO Sarah Franklin withdrew the feature, telling Fortune: "This innovation sparked a lot of conversation and questions that have no clear answers yet… we will not further pursue digital workers in the product."
Lattice belongs here as the control case for New York. Same failure category — a launch that met reality badly — and a total elapsed time of 72 hours, because a product company that sells to HR professionals gets its feedback from the people whose approval it needs to survive, immediately and unambiguously. New York's feedback loop ran through a press investigation, a comptroller's audit, and an election. The difference in cure time is a difference in who is allowed to say no and how fast they can say it.
Sometimes the theatre is the deliverable itself. In 2025 Australia's Department of Employment and Workplace Relations received a report from Deloitte on welfare compliance under a contract valued at about A$440,000. Academics found fabricated references — non-existent academic papers and a made-up quotation from a Federal Court judgment. The revised version disclosed that a generative AI system, Azure OpenAI, had been used in its production. Deloitte refunded A$97,000 — less than a quarter of the contract value.
The refund is the smallest number in the story. The reputational and policy cost of a government welfare-compliance report containing invented case law is not A$97,000. More instructive for anyone running an internal program: the errors were caught by an outside academic reader, not by the firm's own review. The deliverable passed every internal quality gate the engagement had. That is the exact shape of the risk when generative tools are introduced into knowledge work without a corresponding change to verification: the failure mode is not visibly bad work, it is confidently formatted work that nobody re-checks because it looks like the work that used to be checked.
12 · Where the value actually leaks — the measured evidence
Everything so far has been reconstructed from cases. What follows is the small body of properly measured evidence on what happens to value between the tool and the P&L.
Finding one: freed time is reallocated, not banked. Anders Humlum and Emilie Vestergaard linked large-scale adoption surveys to Danish administrative payroll records — 25,000 workers across 7,000 workplaces in the latest round. In workplaces combining encouraged use, enterprise tools and training, 93% of workers reported using chatbots at work, 28% daily, and 19% reported saving more than an hour a day. Those are excellent adoption numbers by any enterprise standard. The earnings and hours result: "precise null effects… at both the worker and workplace levels, ruling out effects larger than 2% two years after the launch of ChatGPT."
The outcome in that study is not what anybody said happened; it is what the payroll system recorded — the survey supplies the adoption, the administrative data supplies the result. And a precise null is a positive finding rather than an empty one: the study did not fail to detect an effect, it measured the effect, found it sitting at about zero, and had the statistical room to rule out anything much larger. "We found nothing" and "we established there is nothing above this size" are very different sentences.
The paper states the mechanism: "most chatbot users (85%) report reallocating time savings from AI chatbots to other job tasks." New tasks appeared too — content generation, oversight of AI outputs, integration work — reported by about 8% of users without employer initiatives and roughly 17% where initiatives existed.
The time was really saved. The workers really felt it. And it went into other work — some of it new work created by the tool itself. No cost line fell because no mechanism existed to make one fall. Gate 4 was never built.
Finding two: self-reported productivity is wrong in a predictable direction. METR ran a randomised controlled trialThe design that carries the most weight in any evidence hierarchy: who gets the tool and who does not is decided by chance, so the two groups differ only in the tool. Everything else — skill, motivation, the difficulty of the work — averages out, which is what lets you say the tool caused the difference rather than merely accompanying it. with 16 experienced open-source developers across 246 real issues in repositories they knew well, averaging over 22,000 stars and a million lines of code, using Cursor Pro with Claude 3.5/3.7 Sonnet, frontier models at the time. Developers forecast a 24% speed-up. They were measured 19% slower. Asked afterwards, they estimated they had been sped up by 20%.
Two numbers, same people, same tasks: a 39-percentage-point gap between the experience and the measurement. METR is careful about generalisation, and so should anyone citing it be — 16 developers, mature codebases they already knew intimately, a setting where the AI's context disadvantage is largest. It does not show AI slows most developers. What it shows is narrower and far more damaging to standard practice: practitioner self-report is not a valid instrument for measuring AI productivity effects, even among skilled practitioners reporting on their own recent work. The overwhelming majority of enterprise AI business cases are built on exactly that instrument.
METR has since revised its own position, and the revision should be carried by anyone who quotes the 19%. On 24 February 2026 the group reported that a larger follow-on experiment begun in August 2025 "gives us an unreliable signal of the current productivity effect of AI tools" — because 30% to 50% of developers were declining to submit tasks they did not want to do without AI, which biases the estimated speed-up downward. Its own reading now: "it is likely that developers are more sped up from AI tools now — in early 2026 — compared to our estimates from early 2025," while cautioning that "our data is only very weak evidence for the size of this increase."
The sign of the 2025 effect is now the wrong thing to argue about, and this brief does not rest on it. What the trial established is the gap between what skilled people believed about their own recent work and what a clock recorded — and nothing in the 2026 update repairs that instrument. A firm that could not measure itself in 2025 cannot measure itself in 2026 either; it can only be wrong in a more flattering direction.
Finding three: the gains are real, and they are concentrated where nobody points the pilot. Brynjolfsson, Li and Raymond studied the staggered rolloutGiving the tool to different groups at different times rather than to everyone at once. The people who have not got it yet act as a comparison group for the people who have — a natural experiment you get for free from an ordinary deployment schedule, and the closest most firms will come to a controlled trial without running one. of a generative assistant across 5,179 customer-support agents: productivity up 14% on average in issues resolved per hour, "including a 34% improvement for novice and low-skilled workers but with minimal impact on experienced and highly skilled workers," plus improved customer sentiment and higher retention. The mechanism is knowledge transfer — the model disseminates the practices of the best agents to the newest.
Put finding three beside finding two. The value concentrates in the least experienced workers on the most routine tasks; the enthusiasm, the budget and the pilot sponsorship concentrate in senior people doing complex work, where the measured effect is smallest or negative. MIT NANDA observed the budget half of this directly: sales and marketing captured the largest share of GenAI allocation "because it's easier to attribute metrics," while "back-office automation often yields better ROI." Note honestly that NANDA's own text is internally inconsistent on the size of that share — its takeaway line says 50% and its body says "approximately 70 percent" — so the direction is what should be carried, not the number.
Finding four: badly-used AI has a measurable negative externalityA cost your activity imposes on somebody else, which never appears on your own ledger. Pollution is the classic case. Here it is quieter: the time you save by generating something is spent again, with interest, by the colleague who has to work out whether it is any good.. BetterUp Labs and Stanford's Social Media Lab surveyed 1,150 US full-time desk workers in September 2025 and found 40% had received "workslop" — AI-generated content that looks like work and lacks the substance to advance it — in the prior month, costing roughly two hours per incident to sort out, which they price at about $186 per employee per month, or around $9 million a year at 10,000 employees. The trust effects were worse than the time cost: recipients rated senders as less capable and less reliable.
Workslop is what individual-level productivity gains look like when they are not governed. Person A's time saving becomes Person B's time cost, plus a trust debit. Netted across the firm the tool can be productivity-negative while every individual user honestly reports a gain — and every one of those honest reports lands in the pilot's success survey.
Finding five: the same gap shows up in the accounts, and it is not a story about engineers. METR measured sixteen developers. In March 2026 the Federal Reserve Bank of Atlanta, with the Richmond Fed and Duke, ran the firm-level version — a survey of nearly 750 corporate financial executives — and found the same shape in CFO answers. Relative to firms not investing in AI, investing firms reported AI-related labour-productivity growth 2.4 percentage points higher in 2025 and expected 3.3 points higher in 2026. The productivity gains implied by those same firms' own reported inputs and outputs were 1.0 and 1.8 points. The authors name it: "a productivity paradox, in which perceived productivity gains are larger than measured productivity gains."
Be careful what this is evidence for, because it cuts against a lazy reading of this brief as well as for it. The gains here are positive, they are largest in high-skill services and finance, and the paper expects them to strengthen. Anyone arguing that enterprise AI does nothing has to explain that away, and this brief does not try to. What the paper also finds is that the effect on employment is close to nothing — firm-size-and-sector-weighted aggregate employment is expected to fall by less than 0.4% due to AI in 2026, with large firms shedding and small firms adding — while the composition shifts, routine clerical work falling by more than two percentage points of employment share over three years. Real productivity, real task reallocation, almost no headcount. That is the Danish result and the Census result arriving a third time, from a third method, in the mouths of the people who own the budget.
| Study | Design & sample | Headline result | What it establishes | What it cannot support |
|---|---|---|---|---|
| Humlum & Vestergaard NBER WP 33777, May 2025 (rev. Mar 2026) | Surveys linked to Danish administrative payroll; 25,000 workers, 7,000 workplaces; difference-in-differences | Precise null on earnings and hours; rules out effects >2%. 85% of users reallocate time savings to other tasks | That firm-level financial impact does not follow automatically from adoption or from felt time savings | That AI has no productivity effect; the paper documents substantial task reorganisation |
| METR 10 Jul 2025 | RCT; 16 experienced OSS developers, 246 real issues in their own mature repos | 19% slower with AI; forecast +24%, self-estimated +20% after | That practitioner self-report is an invalid measurement instrument for AI effects | Generalisation to most developers or to unfamiliar codebases; sample is 16. METR's own 24 Feb 2026 update reports that its follow-on experiment gives "an unreliable signal" and that developers are likely more sped up in early 2026 — so the sign of the 2025 effect carries no weight here, only the perception gap |
| Brynjolfsson, Li & Raymond NBER WP 31161 | Staggered rollout; 5,179 customer-support agents; issues resolved per hour | +14% average; +34% novice/low-skill; minimal for experienced | That real gains exist and concentrate in routine work done by less-experienced staff | That the firm captured the gain financially — throughput is measured, P&L is not |
| Atlanta Fed / Richmond Fed / Duke Working Paper 2026-4, Mar 2026 | Survey of ~750 corporate financial executives; extensive- and intensive-margin regressions | Reported labour-productivity growth +2.4pp (2025) and +3.3pp expected (2026) for AI investors, against +1.0pp and +1.8pp implied; aggregate employment expected to fall <0.4% in 2026 | That the perception-versus-measurement gap holds at the level of the budget holder, and that real gains coexist with almost no headcount effect | A causal estimate; responses are executive self-report and expectations, and "implied" is a construction, not an audited outturn |
| BetterUp Labs / Stanford SML Sep 2025 | Survey; 1,150 US full-time desk workers | 40% received "workslop" in prior month; ~2 hrs per incident; ~$186/employee/month | That ungoverned AI output transfers cost to colleagues and debits trust | A firm-level net productivity figure; the cost estimate is derived from self-reported time |
13 · The incentive map — everybody behaved rationally and the value still died
Almost nothing in this gallery required anyone to act in bad faith. Lay out what each participant around a pilot is actually rewarded for and most of the failures stop being surprising.
One under-quoted NANDA finding is the end-user row made visible. While only 40% of companies had bought an official LLMLarge language model — the kind of system behind a chat assistant, trained on enormous quantities of text to predict what words come next, and general enough to draft, summarise, translate or answer without being built for any one of those jobs. Its generality is what makes it easy to adopt informally and hard to govern. subscription, workers at over 90% of surveyed companies reported regular use of personal AI tools for work. Employees crossed the divide privately while the corporate program sat in pilot. The people are not resisting AI. They are resisting your AI, and they have a substitute.
14 · The taxonomy — nine failure modes, built from the cases up
Data quality, model accuracy, integration, skills: that is the list you get when a taxonomy of this subject is imposed from the top. The taxonomy that emerges from the cases above is a list of operating decisions, and hardly any of its modes would be fixed by a better model.
Sixteen cards, nine modes. Most of these programmes fail at two or three modes at once, and the modes they share turn out to be the ones decided before a model was chosen.
| # | Failure mode | Causal mechanism | Defining case | Detection question |
|---|---|---|---|---|
| 1 | No value on offer | The problem was never worth money, or the saving is too small to survive its own overhead | Lattice "digital workers"; McDonald's AOT (metric never stated) | Can you name the P&L line, its owner, and the size of the move — before the model exists? |
| 2 | Value not ownable | The gain is real but accrues to customers, competitors or the vendor, not to the firm | Representative archetype, not a specific company — commodity chat deflection; generic productivity assistants | If every competitor deploys the same tool from the same vendor, what do you still keep? |
| 3 | Context mismatch | The model was optimised for conditions that do not obtain where the work happens | Google Health Thailand (21% images rejected); MD Anderson (built on the prior EHR) | Has anyone on the team watched the work being done, in the room, for a full shift? |
| 4 | Silent accuracy decay | Live performance diverges from validation and nothing in production surfaces it | Epic Sepsis Model (AUC 0.63 external; 14.7% sensitivity in 2023 EDs) | What is your production metric, measured on your data, and when was it last recomputed? |
| 5 | Trust & liability collapse | Output binds the firm or loses user confidence faster than accuracy improves | Moffatt v. Air Canada; NYC MyCity chatbot; Taco Bell | What is the expected cost of one wrong answer × conversation volume — not the error rate? |
| 6 | Broken unit economics | Cost, latency or an irreducible human tier makes the per-transaction case negative | Presto Voice (~70% human intervention); Just Walk Out in grocery baskets | What is the fully loaded cost per transaction including escalation, review and rework? |
| 7 | Capacity never banked | Time is genuinely freed and reallocated to other tasks; no cost line moves | Denmark: 85% reallocate savings, null effect on earnings and hours | Who has committed, in writing and by date, to convert the freed hours into what? |
| 8 | Governance kill | A control, legal or regulatory constraint terminates a working system after sunk cost | FTC/Rite Aid; EEOC/iTutorGroup; Mobley v. Workday | Was legal, risk and compliance in the room at problem selection, or at go-live? |
| 9 | Success theatre | Termination costs the sponsor more than the program costs the firm, so it persists | NYC MyCity ($100m+, 4 years, chatbot killed only on a change of administration) | What written criterion would kill this program, who applies it, and on what date? |
15 · The aggregate evidence, and how much of it to believe — the numbers everyone quotes
The "95% of AI pilots fail" statistic has done more to shape enterprise conversation in the past year than any case in this gallery. So handle it properly: neither repeat it nor dismiss it.
It comes from "The GenAI Divide: State of AI in Business 2025", produced by Project NANDA at the MIT Media Lab, research period January–June 2025. The executive summary: "Despite $30–40 billion in enterprise investment into GenAI, this report uncovers a surprising result in that 95% of organizations are getting zero return… Just 5% of integrated AI pilots are extracting millions in value, while the vast majority remain stuck with no measurable P&L impact."
Four things about it, in order of importance.
What it actually measured. Not pilots in general. The funnel is explicit and it splits by tool class: for general-purpose LLMs, 80% investigated → 60% piloted → 50% successfully implemented; for embedded or task-specific enterprise GenAI, 40% → 20% → 5%. The 95% is the complement of that last figure. The research note defining success is doing an enormous amount of work: "We define successfully implemented for task-specific GenAI tools as ones users or executives have remarked as causing a marked and sustained productivity and/or P&L impact."
That is a high bar and a soft instrument at once — a demanding criterion assessed by asking people. A tool delivering a genuine 3% cost reduction in one department, unremarked by an executive, scores as a failure. So does a tool that a manager enthusiastically credits with a "marked" impact that never appears in the accounts.
What the authors themselves say about it. The report's own research limitations are candid and almost never quoted: "These figures are directionally accurate based on individual interviews rather than official company reporting. Sample sizes vary by category, and success definitions may differ across organizations." The methodology is 300 publicly disclosed initiatives reviewed, 52 organisations interviewed, and 153 senior-leader survey responses "collected across four major industry conferences" — a sampling frameThe pool a survey draws its respondents from, which silently decides what the answers can mean. Ask AI questions of people who paid to attend AI conferences and you have not surveyed enterprises; you have surveyed enterprises already doing this, which is a different and much smaller country. that over-represents firms actively pursuing AI.
Its internal inconsistency. In the investment section the takeaway line states "50% of GenAI budgets go to sales and marketing" while the body of the same section says "Sales and marketing functions captured approximately 70 percent of AI budget allocation." Both cannot be right. The allocation itself came from asking executives to divide a hypothetical $100. This is not a fatal flaw in a directional finding; it is a decisive reason not to quote either number to two significant figures.
Why it survives the criticism anyway. Because it is not the only measurement pointing the same way, and the others use different methods. That convergence — not NANDA's precision — is what makes the conclusion durable.
| Source & date | Figure | Population & method | Exact definition of "failure" | Tier |
|---|---|---|---|---|
| MIT NANDA Jul 2025 (research Jan–Jun 2025) | 95% zero return; 5% of task-specific tools reach production | 300 public initiatives, 52 interviews, 153 conference surveys | No "marked and sustained" productivity or P&L impact as remarked by users or executives | 3 |
| Gartner 29 Jul 2024 | At least 30% abandoned after PoC by end-2025 | Analyst forecast (not a measurement) | Abandonment after proof of concept, from poor data quality, inadequate risk controls, escalating costs or unclear business value | 2 |
| Gartner 25 Jun 2025 | Over 40% of agentic projects cancelled by end-2027 | Analyst forecast; separate Jan 2025 poll of 3,412 webinar attendees on investment posture | Project cancellation, same three causes | 2 |
| S&P Global Market Intelligence Mar 2025 | 42% abandoned most AI initiatives (17% in 2024); 46% of PoCs scrapped | >1,000 enterprises, North America and Europe | Abandonment of the majority of initiatives; PoCs discarded before production | 2 |
| RAND 13 Aug 2024 | "More than 80% of AI projects fail" | 65 structured interviews with data scientists and engineers | Not measured by RAND — attributed to "some estimates"; RAND's own contribution is the five root causes | 2 |
| McKinsey 5 Nov 2025 | 39% attribute any enterprise EBIT impact; ~6% are high performers | 1,993 respondents, 105 countries | Inability to attribute enterprise-level EBIT impact to AI; most of the 39% report under 5% | 2 |
| Deloitte Tech Trends 2026 (2025 survey) | 11% of organisations run agentic AI in production; 38% piloting | Deloitte 2025 Emerging Technology Trends survey | Not "failure" — a production-adoption count, frequently misquoted as an 89% failure rate | 2 |
| US Census Bureau / Federal Reserve May 2026 / Apr 2026 | 19.8% of US firms use AI in a business function; 78% of the labour force works at an adopting firm | BTOS (large probability sample); Fed comparison of three surveys | Not a failure measure — the denominator problem itself | 1 |
One more piece of arithmetic, and it belongs to the Federal Reserve rather than to any consultancy. In April 2026 a Fed FEDS Note compared three high-quality surveys and found that Census Bureau data put firm-level AI adoption at about 18% at year-end 2025, individual work-related generative AI use at about 41%, and the Survey of Business Uncertainty at 78% of the labour force working at firms that have adopted AI. All three are correct. They count firms, people and employment-weighted exposure respectively. The note's own list of causes for the spread is instructive: "differences in sampling distributions and units of analysis… question framing, the materiality of reported usage, information asymmetries between different target respondents, and social desirability bias."
Census's May 2026 release puts the national rate at 19.8% of firms as of 3 May 2026, having hovered between 17% and 20% since December 2025 — and firms expected to be using AI in six months have hovered between 20% and 23% across the same period. The forward expectation has been running about three points above the current rate for half a year, and the current rate has barely moved. That is what a plateau looks like in survey data.
Beneath that headline sits the best single piece of evidence in this entire brief, and almost nobody has read it. In 2026 Census researchers published "The Microstructure of AI Diffusion", using a dedicated AI supplement to BTOS — a nationally representative probability sampleA survey where every firm in the country has a known, non-zero chance of being picked, so the answers can be scaled up to describe the whole economy. It is the gold standard, and it is why a government statistical series and a vendor survey are not two versions of the same thing: one is measuring the population, the other is measuring whoever answered., not a conference survey — to look at three layers at once: firm-level use, deployment by business function, and worker-task use. Over the November 2025–January 2026 reference period, 18% of firms used AI in a business function, 32% employment-weighted, with 50–60% use (60–70% employment-weighted) among very large firms in Information, Professional Services and Finance.
Four findings from it displace a good deal of what circulates as fact.
Adoption is narrow even where it exists. "Among adopting firms, the scope of use remains limited: 57% of users integrate AI in three or fewer business functions, most commonly Sales and Marketing (52%), Strategy and Business Development (45%), and IT (41%)." That independently corroborates NANDA's sales-and-marketing skew — from a probability sample rather than a hypothetical $100 allocation exercise — and it settles the direction even though NANDA's own number could not be pinned down.
Augmentation, not substitution, is what is actually happening. "Most users (66%) rely on AI solely to augment tasks, while AI-related employment decreases are rare, occurring in only 2% of firms."
Diffusion runs in both directions and neither implies the other. "Worker task use sometimes occurs without formal firm-level adoption, and firm-level adoption sometimes occurs without worker task use." The first half is shadow AI measured properly; the second half is the deployment-without-behaviour-change failure measured properly.
And the finding that matters most. Regression shows a positive correlation between commercial performance and the breadth of AI integration, and it holds up. But "a distinct divergence emerges… with respect to labor outcomes. Functional breadth and operational investment are positively associated with employment decreases, whereas worker-task integration shows no significant link to headcount reduction once functional integration and operational investment are taken into account."
Giving people tools does not reduce headcount. Restructuring functions around AI, and investing operationally to do it, does. It is the American, cross-sectional, probability-sampled version of the Danish result — arrived at by a different method, in a different labour market, with different institutions — and it says the same thing. The operating change, not the tool, is the mechanism.
PART II — THE DISCIPLINE
Fourteen sections, conclusion-forward. Every remedy below is matched to a numbered failure mode from Figure 13, and every one of them is an operating change that can be made before a model is chosen.
16 · Write the P&L line first — problem selection and the value hypothesis · fixes modes 1, 2
The single cheapest intervention available to an enterprise AI program costs nothing and takes an afternoon: before any tool is evaluated, write one sentence naming the account that will move, the person who owns that account, the size of the move, and the date it will be visible. If the sentence cannot be written, the pilot is not ready. The inability to write it is the diagnosis.
Test it against the cases. McDonald's could not, or would not, state the metric by which its two-year IBM test was judged — and the test ended without an expansion decision anyone could evaluate. CBA wrote a version of the sentence ("investment in technology, including AI, is making it easier and faster for customers to get help") that named no account, no baseline and no measurement, and then made a 45-person headcount decision on it, which it reversed six weeks later. Lattice's digital-worker feature had a narrative and no account at all.
A value hypothesis that passes has four components, and the fourth is the one that gets skipped.
The account. Not "efficiency." A general ledger line: contact-centre labour cost, claims leakage, days sales outstanding, first-pass yield, clinician documentation overtime. If nobody can name the account, no one will ever be able to prove the pilot worked, which means it will be judged on anecdote — and anecdote, per METR, is systematically biased upward.
The size, relative to the noise. A projected 2% improvement in a line that swings 8% quarter to quarter is unprovable and therefore unbankable, however real it is. This is the Zillow lesson generalised: compare the expected effect to the variance of the thing you are trying to move, not to zero.
The counterfactual. What happens if you do nothing, fix the process without AI, or buy the outcome? A surprising share of pilots are competing against a process fix that costs less and works better. Amazon's Dash Cart is exactly that judgment made in public: the outcome was worth having and the cheaper mechanism won.
The claim on the freed resource. Who has agreed, before the pilot starts, to give up what — a role not backfilled, a contractor budget cut, a queue that will absorb more volume without more people. This is Gate 4, moved to the front. Without it the honest projected value of a time-saving pilot is zero, and the Danish evidence says so: 85% of users reallocate savings to other tasks, and earnings and hours do not move.
- 01The accountName the general-ledger line and its owner. "Efficiency" is not a line. If none exists, stop.
- 02The size vs the noiseExpected move compared to that line's historical quarterly variance. Smaller than the noise means unprovable.
- 03The counterfactualDo nothing / fix the process / buy the outcome. AI must beat the cheapest of these, not beat zero.
- 04The claim on the resourceNamed owner commits in advance to what is given up when capacity is freed, and by when.
17 · Portfolio and kill rates — ruthless prioritisation · fixes modes 1, 9
Most enterprises set out to reduce the number of failed pilots. That is the wrong objective, and pursuing it produces exactly the pathology in §11: fewer terminations, longer zombie programs, more sunk cost.
A healthy portfolio has a high kill rate and a short time to kill. What NANDA found about speed is the sharpest operating datum in the report and has nothing to do with technology: mid-market top performers reported average timelines of 90 days from pilot to full implementation; enterprises took nine months or longer, and reported the lowest pilot-to-scale conversion rates despite leading in pilot count and staffing. More pilots, more people, more time, worse outcomes.
The mechanism is arithmetic. A pilot's expected value is roughly (probability of value) × (size of value) − (cost of running it) − (cost of the delay in learning). Enterprises optimise the third term, which is small, and ignore the fourth, which compounds. A nine-month pilot cycle gives you at most one learning iteration a year per team; a 90-day cycle gives four. Over two years that is an eightfold difference in accumulated evidence about your own operating context, which is the scarce input.
Three portfolio rules follow. Each is unpopular.
Write the kill criterion into the charter, with a date and a named person who applies it. Not "we will review progress." A number and a deadline: "if human intervention exceeds 25% of transactions at day 60, this stops." This works for a reason that has nothing to do with analytical rigour: it moves the termination decision away from the sponsor, who is punished for making it, and onto a pre-commitment made when nobody's reputation was yet attached.
Cap the number of concurrent pilots at the number of business owners you actually have. Not the number of use cases, not the budget divided by unit cost. Gate 4 requires an owner with authority over the resource; if you have six such people, you can run six pilots. NANDA's enterprise paradox — pilot volume up, conversion down — is what happens when this cap is ignored.
Report the kill rate to the board as a health metric, not a failure metric. A program that has terminated nothing in eighteen months is not succeeding — it has stopped measuring.
18 · Sponsorship and decision rights — org design · fixes modes 1, 7, 9
RAND's contribution to this literature is often mis-cited for the 80% figure it borrowed from elsewhere. Its actual finding, from 65 structured interviews with data scientists and engineers of five-plus years' experience, is that the leading root cause of AI project failure is that "industry stakeholders often misunderstand — or miscommunicate — what problem needs to be solved using AI. Too often, trained AI models are deployed that have been optimized for the wrong metrics or do not fit into the overall business workflow and context."
Its first recommendation follows: "Ensure that technical staff understand the project purpose and domain context… Misunderstandings and miscommunications about the intent and purpose of the project are the most common reasons for AI project failure." Its second is the one no enterprise wants to hear: "Choose enduring problems. AI projects require time and patience… leaders should be prepared to commit each product team to solving a specific problem for at least a year. If an AI project is not worth such a long-term commitment, it most likely is not worth committing to at all."
That seems to contradict §17's argument for 90-day cycles. It does not. RAND's year is a commitment to the problem. The 90 days is a cycle on the solution. A team that owns "reduce claims-handling cycle time" for a year and runs four experiments against it accumulates domain knowledge; a team that owns "deploy the vendor's claims assistant" for nine months accumulates a vendor relationship. Almost every enterprise AI operating model in the field commits to solutions and rotates through problems, which is precisely backwards.
On structure, NANDA is direct: "The dominant barrier to crossing the GenAI Divide is not integration or budget, it is organizational design. Our data shows that companies succeed when they decentralize implementation authority but retain accountability." Decentralised authority with retained accountability is a specific arrangement, not a slogan: the business unit chooses and owns the use case and the value commitment; a central function owns the platform, the evaluation harness, the model contracts and the veto. Central teams that own use-case selection produce a queue. Business units that own the platform produce nine incompatible ones.
One further point on sponsorship, drawn from the incentive map. The sponsor should not be the person who benefits from the announcement. Where possible, make the accountable executive the one who must give up the resource — the operations leader whose headcount line falls, not the technology leader whose budget rises. This single reassignment collapses most of modes 1, 7 and 9 at once, because it puts the person with the strongest reason to be sceptical in charge of the go/no-go.
19 · Buy the workflow, build the edge — sourcing · fixes modes 3, 6, 9
The build-versus-buy answer in the data is lopsided. In NANDA's sample, "external partnerships with learning-capable, customized tools reached deployment ~67% of the time, compared to ~33% for internally built tools," with pilots built through strategic partnerships "2x as likely to reach full deployment" and employee usage rates "nearly double for externally built tools." The report is careful — self-reported outcomes, 52 organisations, correlation not causation — and Deloitte's independent 2025 survey reports the same 2× relationship for agentic pilots.
The internal build's cost curve is back-loaded and invisible at approval time. What gets approved is the build. What kills it is year two: the evaluation harnessThe standing test rig that scores a model against a fixed set of cases with known right answers, so you can tell whether this week's version is better or worse than last week's. Without one you are not managing a system, you are reacting to complaints — and it is invariably the line cut first from a build budget, because at launch there is nothing yet to compare against. nobody budgeted, the model version upgrade, the departure of the one engineer who understood the retrieval layerThe machinery that finds the right internal documents and feeds them to the model before it answers, so the answer is grounded in your policies and records rather than in whatever the model absorbed during training. Most of the difficulty in a corporate assistant lives here, not in the model., the integration that breaks when the source system is upgraded. MD Anderson is this failure in its purest form — the asset did not decay, its environment moved, and no one owned keeping up.
The rule that follows is a decomposition:
Buy the workflow. If the process is one thousands of firms run similarly — contract review, call summarisation and routing, document intake, code assistance, expense classification — buy it, and buy it from a vendor whose product is that workflow rather than a platform you must assemble it on. NANDA's own list of categories that succeeded is exactly this: "voice AI for call summarisation and routing, document automation for contracts and forms, code generation for repetitive engineering tasks." Its list of failures is equally instructive: "complex internal logic, opaque decision support, or optimization based on proprietary heuristics."
Build the edge. Build only where the asset is genuinely yours and durable — your proprietary data, your pricing logic, your risk model, the thing a competitor buying the same vendor product cannot replicate. This is failure mode 2 answered in the sourcing decision: if every competitor can buy your advantage, you have bought a cost, not an advantage, and you should buy it as cheaply as possible.
Never build the plumbing. Retrieval, evaluation harnesses, observabilityInstrumentation that lets you see what a live system is actually doing — what it was asked, what it answered, where it slowed down, when its behaviour drifted. The Epic sepsis case is what an unobserved deployment looks like from the inside: running fine, by all appearances., guardrail layersThe filters wrapped around a model that block certain inputs and outputs — the topics it may not discuss, the claims it may not make, the data it may not repeat. They sit outside the model rather than inside it, which is why they can be changed without retraining anything., model routingSending each request to whichever model suits it — a cheap fast one for the easy majority, an expensive one for the hard remainder — instead of paying top rate for everything. Plumbing, and the reason a stated price per query tells you very little about a real bill.. These are commodity, they are where internal builds sink their year-two budget, and the buy-versus-build calculation on them has been settled for two years.
One caution on the 67/33 figure, since it will be quoted. It is self-reported, from 52 organisations, in a report whose author acknowledges "the correlation between external partnerships and success does not necessarily prove causation" — and there is an obvious selection effectWhen the way cases end up in each group, rather than the treatment itself, produces the difference you are measuring. Here: firms build in-house precisely when nothing can be bought, so the "build" column is stocked with the hardest problems before anyone writes a line of code., since organisations attempt internal builds precisely where no product exists, which is also where problems are hardest. The direction is well-supported; the ratio is not a constant of nature.
- Voice AI for call summarisation and routingBounded task, immediate visible output, low configuration burden, a human still owns the decision
- Document automation — contracts and formsHigh-volume structured extraction where the error is caught downstream by an existing check
- Code generation for repetitive engineeringThe user is also the reviewer; verification is instant and free
- Ambient clinical documentationAttacks the task practitioners most resent; adoption is voluntary and self-sustaining
- Complex internal logic and orchestrationDeep enterprise specificity; the configuration burden exceeds the value before go-live
- Opaque decision supportUsers cannot verify the output cheaply, so they do not act on it — deployment without behaviour change
- Optimisation on proprietary heuristicsRequires codifying knowledge the organisation has never written down
- Broad-scope, complex-execution agents"Fails" in NANDA's own scope/complexity matrix; Gartner's agentic cancellation forecast points the same way
20 · Production engineering — evals, monitoring, reliability, cost · fixes modes 3, 4, 6
The Epic Sepsis Model makes a boring proposition unarguable: a deployed model without a live, on-your-own-data performance measurement is an unmonitored liability that looks like an asset. It ran for years. Its degradation relative to the vendor's cited range was found by academics, not by the hospitals paying for it.
Four engineering commitments follow, each mapped to a case above.
Compute the metric on your data, at your threshold, on a schedule. A vendor's accuracy figure is a claim about their evaluation set. Michigan's AUC of 0.63 against a cited 0.76–0.83 is the size of the gap that is possible. Texas made this an enforceable obligation for one vendor in one sector; buyers should impose it contractually everywhere.
Measure decision value, not just discrimination. The sepsis model's 14.7% sensitivity is bad. Its zero-minute median lead time is disqualifying, and it is a different kind of measurement — it asks whether the output arrives in time to change anything. Every predictive deployment needs the equivalent: for a churn model, how long before the churn; for a fraud model, before or after the funds move; for a maintenance model, before or after the failure window closes.
Instrument the human tier as a first-class metric. Presto's ~70% human-intervention rate was, in the SEC's account, invisible to investors until an investigation. Whatever your equivalent — escalation rate, override rate, share of outputs edited before use, images rejected — that number is the honest measure of automation achieved, and it should sit on the same dashboard as accuracy, computed weekly. Google Health's 21% rejection rate is the same statistic wearing different clothes; had it been a tracked KPI rather than a finding, the deployment design would have changed.
Model the fully loaded cost per transaction, including rework. Inference is the cheapest line and falling. Escalation handling, review labour, incident response and the legal review of what the system may say are the expensive lines and they are not falling. Gartner's $5–20 million range for business-model-innovation deployments is mostly not compute. A pilot business case that shows only model cost is not a business case.
| Metric | Definition | Cadence | Failure it would have caught |
|---|---|---|---|
| Own-data performance | The vendor's headline metric, recomputed on your population at your operating threshold | Monthly, and on every model version change | Epic Sepsis Model — AUC 0.63 on external data against a cited 0.76–0.83 |
| Decision lead time | Time between the output and the moment the decision must be made | Monthly, reported as a median with interval | Epic Sepsis Model — median lead time of zero minutes across 145,885 encounters |
| Human intervention rate | Share of transactions requiring a person to enter, correct, validate or escalate | Weekly, on the same dashboard as accuracy | Presto Voice — ~70% of orders on the proprietary system; Google Health Thailand — 21% of images rejected |
| Fully loaded cost per transaction | Inference + escalation labour + review + rework + incident and legal overhead | Quarterly, against the pre-agreed baseline | Just Walk Out in large-basket grocery; Gartner's $5–20m deployment cost range |
21 · Attribution a board can audit — value measurement · fixes mode 7
Modes 3, 4 and 6 are engineering problems with engineering answers. Mode 7 — capacity freed and never banked — is a measurement and authority problem, and it defeats programs where everything else went right.
Multiplying users by hours saved by loaded cost, where hours saved comes from a survey, produces a number that is both large and false. METR's developers, reporting on their own recent work in a controlled setting, were off by 39 percentage points in the flattering direction. There is no reason to expect a claims processor's estimate of time saved to be better calibrated than a senior engineer's.
Three methods do work, in ascending order of rigour and descending order of convenience.
The holdout. Withhold the tool from a randomly selected, comparable group for a defined period and compare the outcome metric — issues resolved per hour, cases closed, cycle time — not the self-report. This is what Brynjolfsson and colleagues effectively exploited with a staggered rollout across 5,179 agents, and it is why their 14% average and 34% novice effect are believable in a way that vendor case studies are not. Most enterprises can do this and decline to, because it delays universal rollout by a quarter and someone has to explain to the withheld group why.
The pre-registered baseline. Before deployment, agree with the finance function on the account, the measurement window, the seasonality adjustment and the definition of the counterfactual. Write it down. Pre-registration is not a statistical nicety. After deployment every party's incentive is to redefine success, and a document written before anyone's reputation was attached is the only defence.
The capacity conversion ledger. The one nobody keeps. For every unit of capacity a pilot frees, record what happened to it: converted to headcount not backfilled, converted to volume absorbed without hiring, converted to a service level a customer pays for, or reallocated to other work. The fourth category is the honest destination of most of it, and naming it is what converts an inflated business case into a real one. The Danish study's 85% is the base rateWhat normally happens, before you account for anything special about your own case. It is the number a forecast should start from and argue away from, rather than the number people reach for only after their own estimate has already been proved wrong.; a program that assumes better than that needs to explain the mechanism.
Figure 9 showed the arithmetic, and it inverts standard practice. Raising conversion from 15% to 50% at fixed 40% adoption adds about $1.58 million on those parameters; doubling adoption from 40% to 80% at fixed 15% conversion adds about $675,000. Enterprises spend nearly all their change-management energy on adoption and almost none on conversion, and conversion is worth more than twice as much per unit of movement. At zero conversion the program is negative at every adoption level including 80%, which is the mathematical statement of why so many well-adopted tools show up nowhere in the accounts.
22 · Adoption in fact — change management · fixes modes 3, 5, 7
The best-documented adoption success in enterprise AI is dictation.
The Permanente Medical Group deployed ambient AI scribes — systems that listen to a clinical encounter and draft the note — across Northern California. In the first year, physicians used the technology more than 2.5 million times, and analysis found it saved nearly 16,000 hours of documentation time. Of 102 adult and family medicine physicians surveyed, two-thirds used it five or more days a week and 63% used it in every in-person visit. Crucially: "We found the highest adoption rates in departments that typically suffer from the highest levels of documentation and burnout," and the users who benefited most were the highest-volume users, whose time savings "substantially surpassed the time savings among their peers who used the technology infrequently or not at all."
Four properties made that work, and each is an operating choice available to any organisation.
It attacked the task the users most hated, so adoption did not need to be mandated. It produced output the user verifies in seconds — the physician reads the note they were going to write anyway — so trust was established transactionally rather than institutionally. It was voluntary, so non-adoption was informative rather than hidden. And no headcount reduction was attached to it, which removed the single largest reason for users to sandbagQuietly underperform on purpose — use the tool sparingly, report modest gains, keep the old process running alongside it. Rarely a conspiracy and almost never visible in a metric, because a person protecting their job looks exactly like a person who finds the tool unhelpful. a deployment.
The last of those should be uncomfortable. If a tool's stated purpose is to reduce headcount, the people who must adopt it are being asked to build the case for their own redundancy, and they will not do it well. CBA announced 45 redundancies and then discovered the workload had not fallen. Klarna's most consequential retreat was not technical but positional — from replacement to augmentation, with a guaranteed human path. The organisations getting adoption in fact are, so far, overwhelmingly the ones that separated the tool from the headcount question in time.
Adoption alone is not the goal. BetterUp's workslop finding is the counterweight: 40% of desk workers received AI-generated content that looked like work and was not, costing about two hours each time and debiting the sender's credibility. High adoption of an ungoverned generative tool can be net-negative at the firm level while every individual reports a gain. The governance that prevents this is unglamorous and cheap — a norm that the person who generates output owns its verification, and an expectation that AI-assisted work is disclosed to the colleague who must build on it.
NANDA found that while only 40% of companies had bought an official LLM subscription, workers at over 90% of surveyed companies used personal AI tools for work. That is shadow AIEmployees using AI tools they bought or signed up for themselves, on company work, outside any policy or contract. It is what adoption looks like when nobody has to be persuaded — and it runs on personal accounts, which is why it is invisible to the programme and visible only in surveys., and the security lens gets it backwards. It is free, high-quality market research about which tasks your people find AI genuinely useful for — a revealed-preferenceWhat people's behaviour shows they want, as opposed to what they say they want in a survey. Behaviour costs something and answers do not, which is why the tool an employee pays for out of their own pocket is better evidence than the one they rated highly in a pilot questionnaire. dataset most organisations are trying to suppress rather than read. The tasks where shadow use is heaviest are the tasks where an official deployment will get adoption in fact.
23 · Governance as accelerator — fixes mode 8, and makes 5 survivable
The standard complaint is that risk and legal slow AI down. The case record says something more precise: late governance slows AI down, and it does so by destroying work that has already been paid for. Rite Aid deployed for eight years and then had the capability removed by consent orderA settlement a regulator writes and a court enforces. The company admits no wrongdoing and still accepts binding obligations — here, a ban with a term of years — that it cannot later argue its way out of. Cheaper than losing a trial, and considerably more restrictive than winning one.. iTutorGroup shipped screening logic and paid $365,000 plus five years of EEOC monitoring. Workday is three years into litigation over a product function.
Three questions move from go-live to problem selection, where answering them is nearly free.
What can this system say or decide on our behalf, and what is the ceiling on the resulting obligation? This is the Air Canada question. It has a numeric answer — expected cost of an erroneous representation multiplied by volume — and it belongs in the business case, not the risk register.
What claims will we make about this system's performance, and can we substantiate them by the method we will publish? This is the Pieces and Presto question, and it now has regulators attached in at least two jurisdictions. Internally, the same discipline kills mode 9: a program that must state its metric methodology in writing cannot survive on a vibe.
Who is a protected party in this decision, and what is our evidentiary record if challenged? This is Mobley. If the system touches hiring, credit, insurance, housing or healthcare access, the record you keep from day one is the defence, and it cannot be constructed retroactively.
Two dated obligations now make the timing question concrete rather than theoretical. New York City's Local Law 144 has, since enforcement began on 5 July 2023, prohibited employers and employment agencies from using an automated employment decision tool "unless the tool has been subject to a bias audit within one year of the use of the tool, information about the bias audit is publicly available, and certain notices have been provided to employees or job candidates." A bias audit is not something you can produce retrospectively for a system already in use; the record either exists or it does not.
And the EU AI Act became applicable on 2 August 2026, three weeks before this brief's evidence cutoff, having entered into force on 1 August 2024. Prohibited practices and AI-literacy obligations applied from 2 February 2025; general-purpose model obligations from 2 August 2025. Crucially for planning, the timetable then moved — and it moved as enacted law, not as a proposal. The "AI Omnibus" simplification package, put forward by the Commission in November 2025, was published in the Official Journal on 24 July 2026 and entered into force on 27 July 2026, one month before this brief's cutoff. It defers the Annex III high-risk obligations — the sensitive-area list covering employment, credit, education and essential services — to 2 December 2027, and Annex I product-embedded systems to 2 August 2028. The prohibitions, the AI-literacy duty, the general-purpose model obligations and the Article 50 transparency and content-labelling duties are unaffected and already live.
Read that as an operator rather than a lawyer. An organisation planning against the original August 2026 date now has roughly sixteen further months on its highest-risk use cases, and it is a settled sixteen months rather than a hoped-for one — which is exactly the window in which the evidentiary record a bias claim will later be judged against can still be built cheaply. An organisation reading only the deferral will miss that the duties that bite first, on disclosure and prohibited practice, never moved at all. The deferral bought time on the paperwork, not on the conduct.
A decision boundary cleared in advance — "the assistant may quote published policy verbatim and may not construct an answer about entitlement" — lets a team ship in weeks without a bespoke legal review per feature. That is the acceleration, and it runs backwards from expectation. The slow path is the one where the boundary is negotiated after the demo has been shown to the board, when the sunk cost makes every restriction a fight.
There is a live counter-argument, and it is a fair one. Heavy pre-deployment governance can be an excuse for institutional paralysis, and some organisations use "responsible AI" processes to avoid deciding anything. The test that distinguishes the two: a governance function that produces reusable, written decision boundaries is an accelerator; one that produces case-by-case reviews with no accumulating precedent is a tax. Ask how many of last quarter's reviews created a rule that removed the need for a future review. If the answer is none, the complaint about slowness is correct.
24 · Sector overlays — the same failure, wearing local clothes
The nine modes are general. What varies by sector is which mode binds first and how expensive it is when it does.
| Sector | Binding mode | Local aggravator | Case in this brief | Where value has landed |
|---|---|---|---|---|
| Healthcare | 4 — silent accuracy decay; 3 — context mismatch | Clinicians cannot verify a prediction cheaply, and alert fatigue forces thresholds that destroy sensitivity | Epic Sepsis Model; MD Anderson OEA; Google Health Thailand; Pieces Technologies | Documentation, not diagnosis — 2.5m ambient scribe uses at TPMG in one year |
| Financial services | 7 — capacity never banked; 5 — trust and liability | Regulated advice and dispute exposure; deflection metrics that diverge from workload | CBA voice-bot; Klarna; SEC AI-washing actions | Routine contact automation with a guaranteed human path; fraud and document work |
| Retail & QSR | 6 — unit economics; 5 — trust | Franchise heterogeneity, accents and dialects, viral failure at consumer scale | McDonald's AOT; Taco Bell; Just Walk Out; Presto | Small-basket, high-friction formats; back-of-house forecasting and scheduling |
| Public sector | 9 — success theatre; 8 — governance kill | Launch is the political deliverable; termination is an admission; audit cycles run in years. US federal agencies must now name a Chief AI Officer and apply minimum risk-management practices to "high-impact AI" under OMB M-25-21 | NYC MyCity ($100m+, chatbot live ~2 years after documented illegal advice); Deloitte Australia report | Constrained, verifiable transactions — the MyCity childcare portal did digitise a real application |
| Manufacturing & industrials | 3 — context mismatch; 1 — no value on offer | Sensor and process data collected for control, not for learning; site-by-site variation defeats one model | Not directly evidenced in this brief — see "not verified" | NANDA records maintenance pilots with no major supply-chain shift; treat claims here sceptically |
| Technology | 7 — capacity never banked | Gains are real and land in individual throughput, where the firm has no mechanism to collect them | METR RCT; workslop; Salesforce support reduction from ~9,000 to ~5,000 | Support automation with a measured baseline; code assistance where the user is the reviewer |
25 · The maturity model — self-assess honestly · covers all nine modes
Five levels, six dimensions. Score yourself at the lowest level you can honestly claim across the six, not the highest: the binding constraint is what determines the outcome. An organisation at Level 4 on engineering and Level 1 on value attribution is a Level 1 organisation with expensive infrastructure — which describes a large share of the programs behind the statistics in §15.
| Dimension | L1 · Theatre | L2 · Experimenting | L3 · Shipping | L4 · Capturing | L5 · Compounding |
|---|---|---|---|---|---|
| Problem selection | Use cases arrive from vendors and conferences | Internal list, ranked by enthusiasm | Written value hypothesis with a named account | Hypothesis includes the counterfactual and the size-vs-noise test | Problems owned for a year; solutions cycled quarterly against them |
| Sponsorship | Sponsor benefits from the announcement | Technology leader accountable | Business leader accountable for delivery | Accountable executive is the one who gives up the resource | Decentralised authority, retained central accountability, veto exercised |
| Production engineering | Vendor's accuracy figure is the metric | Own-data evaluation at go-live only | Scheduled own-data evaluation on your threshold | Decision lead time and human-intervention rate tracked weekly | Fully loaded cost per transaction, with automatic rollback triggers |
| Value attribution | Users × hours saved × loaded cost, from a survey | Pre/post comparison, no control | Pre-registered baseline agreed with finance | Holdout group; measured outcome, not self-report | Capacity conversion ledger reconciled to the general ledger |
| Adoption | Seats provisioned counted as adoption | Weekly-active users tracked | Adoption voluntary; non-use treated as a signal | Shadow-AI usage read as demand data and acted on | Verification norms in force; workslop measured and falling |
| Governance | Reviewed at go-live, or after an incident | Checklist applied per project | Risk and legal present at problem selection | Written decision boundaries reusable across projects | Kill criteria pre-committed with dates and named owners; kill rate reported to the board |
26 · The first 90 days — inheriting a pile of stalled pilots
Assume the realistic situation: you have arrived, there are somewhere between nine and forty active AI initiatives, nobody can tell you what any of them are worth, and the board wants a number by the next meeting. The instinct is to build a strategy. Do the inventory instead — the strategy falls out of it, and the inventory is defensible in a way the strategy is not.
- 01Days 1–15 · InventoryOne page per initiative: the account it moves, its owner, spend to date, run-rate, the last date its metric was computed on your data, and the human-intervention rate. Initiatives whose page cannot be completed are your answer to step 2.
- 02Days 15–30 · KillTerminate everything with no named account or no resource claim. Announce the kill rate as a governance result, not a failure. Redirect the run-rate, not the headcount.
- 03Days 30–45 · InstrumentOn the survivors, stand up the four metrics of Figure 20 and a pre-registered baseline agreed in writing with finance. Nothing new is approved until these exist.
- 04Days 45–75 · Prove onePick the single highest-volume bounded task with cheap verification and heavy existing shadow use. Run it with a holdout. The objective is one auditable number, not a portfolio.
- 05Days 75–90 · ProposeTake the board the kill rate, the one measured result with its method, and a capped portfolio sized to the number of business owners who will sign a resource claim. Not a roadmap.
27 · Alternatives considered and rejected — the explanations that do not survive the cases
Four rival explanations for the value gap deserve a fair hearing. Each has real evidence behind it. None survives as the primary cause.
"The models aren't good enough yet." The strongest version: most of these failures date from a period of rapidly improving capability, and a sufficiently capable system would have handled Thailand's poor-quality images, Taco Bell's noisy drive-thru, and the sepsis signal. There is something to this — capability genuinely constrained several cases. But it cannot explain MD Anderson, where the binding constraint was an EHR migration; or CBA, where the bot worked and the workload rose; or Denmark, where the tools worked, the time was saved, and no earnings moved; or NYC, where the failure was documented for two years and the system stayed up. And the capability improvement is not in doubt: Stanford's 2026 AI Index records performance on SWE-bench VerifiedA standard test in which a model is given a real, previously reported bug in a real open-source codebase and has to produce a fix that passes the project's own tests. It is a hard benchmark and a narrow one — a graded exam, not a day at work — which is why saturating it says a great deal about capability and very little about value. rising "from 60% to near 100% in a single year," agent task success on OSWorldA benchmark that scores an AI system on operating a real computer the way a person does, clicking through actual applications, files and browsers to finish a task end to end. It tests the "agent" claim rather than the writing one. jumping "from 12% to ~66%," US private AI investment of $285.9 billion in 2025, and organisational adoption at 88%. Capability, money and adoption all moved hard. The value-realisation numbers did not move with them. If capability were the binding constraint, that combination should be impossible. Rejected as primary cause — high confidence
"The data isn't ready." Gartner names poor data quality first among abandonment causes and Deloitte reports nearly half of organisations citing data searchability and reusability as obstacles. Data readiness is a real and expensive constraint. But it functions more often as the acceptable public explanation than the operative one, because it is the only failure a technology leader can announce without implicating a business decision. Note that in the cases where the data was genuinely the problem — Amazon's résumé model learning from a decade of male-skewed hiring — the failure was correctly diagnosed and the project correctly killed. Real but secondary — moderate confidence
"Enterprises are too slow and risk-averse." NANDA demolishes this directly: enterprises "lead in pilot count and allocate more staff to AI-related initiatives," 90% "have seriously explored buying an AI solution," and they report the lowest pilot-to-scale conversion. The problem is that pilot volume without owned value hypotheses produces a queue, not a portfolio. Rejected — high confidence
"It's too early; the J-curve takes a decade." This is the strongest counter-argument in the set and it should not be dismissed. The J-curve is a shape, and worth having clearly in mind. A genuinely general technology first costs more than it returns, because the spending goes into rewiring the organisation around it — new processes, new skills, new plant — and the return only arrives once the rewiring is finished. Measured productivity therefore sags before it climbs, and the path drawn out looks like the letter. Electrification took roughly forty years to show in US productivity statistics; the productivity payoff from computing arrived long after the investment. Humlum and Vestergaard's own framing is compatible with it — "technological change reshapes work well before it surfaces in earnings or hours," and they document substantial task reorganisation with no earnings effect, which is precisely what an early J-curve looks like. If this is right, most of what this brief calls failure is transition cost.
Three things constrain it. First, it is unfalsifiable on any near horizon, which makes it useless as a management input — a sponsor cannot fund on it and a board cannot audit it. Second, the historical analogues involved complementary capital and organisational redesign taking decades, which is an argument for the operating discipline in this Part, not against it. Third, and decisively for the practical question: the J-curve explains why aggregate statistics lag. It does not explain why one organisation's program lands and its competitor's does not, and that variance — visible in every study here — is where operators actually work. Partly true, and not actionable — moderate confidence
28 · Second-order effects — what follows if this is right
If the operating diagnosis in this brief is right, the consequences run further than the case record shows.
Cheaper models make the operating gap wider, not narrower. When inference cost falls, the binding constraint moves to integration, evaluation, monitoring, escalation and legal review — costs that scale with the number of workflows touched and with headcount, not with tokens. Falling model prices therefore increase the number of technically viable pilots faster than they increase the number of organisations capable of capturing value from one. Expect the ratio of pilots to value-realising deployments to get worse as capability gets cheaper, which will be widely misread as evidence that the technology disappoints.
The advantage accrues to firms with strong management accounting, not strong data science. The scarce capability in this brief is the ability to define an account, agree a baseline, run a holdout, and make an executive give up a resource on a date. That capability lives in finance and operations. The corollary is that the firms best positioned are not the most technically sophisticated but the most operationally disciplined — which is a very different list, and one that public-market narratives about AI adoption largely ignore.
Vendor liability is migrating toward the vendor, which will reprice the software. This prediction has survived a year it could easily have failed. Workday has now been through three rounds of pleading and the "we are merely a software supplier" defence has not carried; the ADEA disparate-impact theory survived dismissal again in March 2026, and in June 2026 the court declined to accept that California's law stops at the applicant's state line when the screening happens in California. None of that is a merits finding, and it may yet become one in Workday's favour — but the cost of defending the theory is already being paid, and that is what prices software. Mobley's agency theory, if it holds, makes an AI vendor performing a core business function liable alongside its customer. Vendors will respond by narrowing warranties, raising prices in regulated functions, or withdrawing from them. Buyers should expect the "AI does the deciding" product category to become materially more expensive in hiring, credit, insurance and healthcare — and should read a vendor's willingness to accept contractual accuracy representations as the single best signal of whether their claims are real.
The measurement discipline that fixes AI value will expose everything else. A capacity conversion ledger applied honestly to an AI program is applicable to any productivity investment, and most enterprises have never had one. The first organisation to run it rigorously will discover that a decade of process-improvement programs also freed capacity that was never banked. That discovery is politically explosive and is one reason the ledger does not get built.
Delegation is rising faster than the controls around it. Anthropic's Economic Index reports that "directive" conversations — where a user hands over a complete task rather than iterating — rose from 27% to 39% of usage, and that lower-adoption regions skew toward delegation while high-adoption regions skew toward augmentation. The same report notes 40% of US employees using AI at work, up from 20% two years earlier. If delegation is growing fastest where usage is least mature, the aggregate volume of unverified AI output entering enterprises is growing faster than the verification capacity around it. That is the workslop mechanism with a trend line attached, and it argues for the verification norms in §22 becoming urgent rather than optional.
The macro data will keep saying "not yet", and that will be misread as "not ever." The Federal Reserve's own mid-2026 assessment, a FEDS Note of 17 July 2026, finds that "labor market impacts remain concentrated and have not yet broadened in the aggregate" and reads the evidence as "a buildout phase rather than the onset of broad-based displacement." That is the correct reading of the aggregate series and it is the wrong input for an operator, for the same reason the J-curve is: a national statistic averages the firms that built the conversion mechanism with the many that did not. Expect the gap between the two groups to widen while the average stays flat, and expect the flat average to be quoted at you in budget season.
A backlash is coming and it will overshoot. The gap between what was promised in 2024–25 and what the accounts show in 2026–27 is large enough to produce a governance reaction — budget consolidation, centralised approval, a freeze on new initiatives. The firms that suffer most from the freeze will be the ones that never built attribution, because they will have no defensible evidence to bring to it. Attribution is not a reporting nicety; it is the survival mechanism for a program in a downcycle.
29 · Falsifiers and dated predictions — what would change this view
- A large, well-identified study (holdout or staggered-rollout design, not self-report) finds firm-level financial effects that scale with model capability rather than with operating practice — i.e. the same organisational design produces materially better value purely because the model improved. That would move the cause back to technology.
- Aggregate value realisation rises sharply — say McKinsey's "any enterprise EBITEarnings before interest and taxes — operating profit, the money the business itself makes before the financing and the tax authorities take their turns. It is the number a chief executive is judged on, which is why "did it move EBIT" is the version of the value question that survives contact with a board. impact" figure moving from 39% to above 60%, with the sub-5% share falling — without any observable change in attribution practice or decision rights. That would suggest the constraint was capability all along.
- Agentic deployments demonstrably close Gate 4 automatically by executing end-to-end processes with no freed-capacity conversion step, so that cost lines fall without an executive decision. Gartner's cancellation forecast and Deloitte's 11%-in-production figure both argue against this today; a reversal by 2027 would be decisive.
- Replication of the Danish null result fails in a comparable economy with comparable data — for example, linked administrative data elsewhere showing earnings or hours effects well above the 2% bound — which would undercut the capacity-conversion mechanism at the centre of this brief. Note that the Census Bureau's 2026 microstructure study already provides one independent US corroboration, so this falsifier now requires overturning two datasets built on different methods.
| Prediction | Horizon | Settled by | Confidence |
|---|---|---|---|
| US firm-level AI adoption (Census BTOS, firm-weighted) remains below 30% — the plateau at 17–20% through mid-2026 does not break upward | By Dec 2027 | Census Bureau BTOS releases | High |
| Gartner's forecast that over 40% of agentic AI projects are cancelled by end-2027 is met or exceeded, with "unclear business value" the most-cited cause rather than model capability | By Dec 2027 | Gartner follow-up research; Deloitte production-adoption tracking | Moderate |
| At least one further US state AG or federal agency action turns on an AI vendor's accuracy or automation-rate claim rather than on downstream consumer harm, extending the Pieces/Presto pattern | By Dec 2027 | State AG and FTC/SEC enforcement announcements | Moderate |
| McKinsey's share of organisations reporting any enterprise EBIT impact stays below 55%, and the share reporting above 5% of EBIT stays in single digits | Next two annual surveys, to end-2027 | McKinsey State of AI survey | Moderate |
| At least one additional large enterprise publicly reverses an AI-justified headcount reduction after workload fails to fall, following the CBA pattern | By Dec 2027 | Company statements, union disputes, labour tribunal filings | Moderate |
Evidence register — dated and tiered; every item opened and read for this brief
Tier 1 = primary disclosures, filings, court and regulator records, first-party post-mortems and official statistics. Tier 2 = reputable analyst, academic and established trade reporting. Tier 3 = single-source, self-reported or estimated, used only where flagged. Event dates and publication dates are distinguished where they differ.
| Source | Date | Tier | What it carries here |
|---|---|---|---|
| Zillow Group Q3 2021 results release (Exhibit 99.1 to Form 8-K) | 2 Nov 2021 | 1 | Wind-down announcement; Barton quote; $304m Q3 write-down; $421.6m Homes pre-tax loss; $240–265m expected Q4 loss; ~25% workforce reduction |
| Zillow Group Form 10-K, FY2021 — Item 1A risk factors and MD&A | Feb 2022 | 1 | $407.9m FY21 inventory write-downs; $71.2m impairment and restructuring; the "submarket nuances" pricing-model risk factor; the board's stated wind-down rationale |
| SEC Order, In re Presto Automation Inc., Rel. 33-11352 | 14 Jan 2025 | 1 | ~70% human-intervention rate; reliance on off-site agents in the Philippines and India; undisclosed third-party technology; cease-and-desist, no civil penalty |
| SEC press release 2024-36, first AI-washing actions | 18 Mar 2024 | 1 | Delphia $225,000 and Global Predictions $175,000 penalties; Gensler quote |
| SEC Order, In re Delphia (USA) Inc., Rel. IA-6573 | 18 Mar 2024 | 1 | Aug 2019–Aug 2023 misstatements; "no such data was being used"; continuation of false statements after a 2021 corrective undertaking |
| Moffatt v. Air Canada, 2024 BCCRT 149 (full reasons) | Event Nov 2022; decision 14 Feb 2024 | 1 | "A remarkable submission"; responsibility for all website information; $650.88 damages plus $36.14 interest and $125 fees; failure to produce the tariff |
| FTC press release, Rite Aid facial-recognition order | Conduct 2012–2020; order 19 Dec 2023 | 1 | Five-year ban; false flags disproportionately affecting people of colour; obligation to discontinue uncontrollable automated systems |
| FTC, Operation AI Comply | 25 Sep 2024 | 1 | DoNotPay "robot lawyer" complaint — no testing against a human-lawyer standard, no attorneys retained; five actions on deceptive AI claims |
| EEOC press release, iTutorGroup consent decree | 11 Sep 2023 | 1 | $365,000; software programmed to auto-reject women 55+ and men 60+; 200+ applicants; five years of monitoring |
| Texas Attorney General, Pieces Technologies settlement | Sep 2024 | 1 | "<1 per 100,000" severe-hallucination claim found likely inaccurate; four Texas hospitals; five-year metric-methodology disclosure obligation |
| NYC Comptroller, audit of OTI's MyCity system | 2025 | 1 | $100m+ over four years; chatbot "unable to provide accurate or consistent information"; further $81m requested for FY2026; poor project and contract management |
| US Census Bureau, Business Trends and Outlook Survey release | Collection to 3 May 2026 | 1 | National AI use 19.8%; 17–20% range Dec 2025–May 2026; 37% at 250+ employees; Information 39.7%, Finance 33.9% |
| US Census Bureau CES Working Paper 26-25, "The Microstructure of AI Diffusion" | Reference period Nov 2025–Jan 2026 | 1 | 18% of firms (32% employment-weighted); 57% of adopters use AI in ≤3 functions; Sales and Marketing 52%; 66% augment only; AI-related employment decreases in just 2% of firms; worker-task use shows no significant link to headcount reduction once functional integration and operational investment are controlled for |
| European Commission, AI Act regulatory framework and application timeline | In force 1 Aug 2024; applicable 2 Aug 2026 | 1 | Prohibitions and AI literacy from 2 Feb 2025; GPAI obligations from 2 Aug 2025; Annex III high-risk extended to 2 Dec 2027 and Annex I to 2 Aug 2028 under the AI Omnibus agreed Nov 2025 |
| NYC Department of Consumer and Worker Protection, Local Law 144 on automated employment decision tools | Enforcement from 5 Jul 2023 | 1 | Bias audit within one year of use, published summary, and candidate notice required before an AEDT may be used |
| C3.ai Form 10-K, FY ended 30 Apr 2026 — risk factors | Filed 24 Jun 2026 | 1 | Paid "Initial Production Deployment" engagement model; disclosed risk that trial customers may not convert to ongoing subscriptions — the pilot-to-production chasm as a seller-side revenue risk |
| Federal Reserve FEDS Note, "Monitoring AI Adoption in the US Economy" | 3 Apr 2026 | 1 | ~18% of firms adopted at year-end 2025; ~41% individual work-related GenAI use; 78% of the labour force at adopting firms; sources of divergence between surveys |
| Klarna press release, AI assistant month-one results | 27 Feb 2024 | 3 | 2.3m conversations; two-thirds of chats; "equivalent work of 700 full-time agents"; 2 vs 11 minutes; 25% fewer repeat inquiries; estimated $40m profit improvement — all self-reported and unaudited |
| Francisco Partners, completion of the IBM Watson Health assets acquisition | Jun 2022 | 1 | Watson Health's data and analytics assets divested and relaunched as Merative |
| JPMorganChase, Chairman and CEO letter to shareholders, 2025 Annual Report | Apr 2026 | 1 | "AI will affect virtually every function, application and process"; adoption pace framed as faster than electricity or the internet — used as an expectations datum, not a value datum |
| OMB Memorandum M-25-21, "Accelerating Federal Use of AI" | Apr 2025 | 1 | Chief AI Officer requirement; minimum risk-management practices for "high-impact AI"; used as the public-sector governance anchor in §24 |
| Stanford HAI, 2026 AI Index Report | 2026 | 2 | SWE-bench Verified from 60% to near 100% in a year; OSWorld agent success 12% → ~66% with ~1 in 3 still failing; $285.9bn US private AI investment in 2025; 88% organisational adoption — the capability counter-evidence in §27 |
| Anthropic Economic Index, September 2025 report | Sep 2025 | 3 | "Directive" (full-delegation) conversations rising 27% → 39%; 40% of US employees report using AI at work, up from 20% in 2023; automation-versus-augmentation skew by region — first-party platform telemetry, not a representative survey |
| RAND, "The Root Causes of Failure for Artificial Intelligence Projects," RR-A2680-1 | 13 Aug 2024 | 2 | 65 interviews; five root causes led by misunderstood problem framing; five recommendations including "choose enduring problems"; the >80% figure attributed to external estimates |
| METR, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity" | 10 Jul 2025 | 2 | 16 developers, 246 issues; 19% slower; +24% forecast, +20% self-estimate after; authors' own generalisation caveats |
| Humlum & Vestergaard, NBER Working Paper 33777 | May 2025, rev. Mar 2026 | 2 | 25,000 workers / 7,000 workplaces; null effects ruling out >2%; 85% reallocate time savings; 93%/28%/19% adoption and saving figures in best-supported workplaces |
| Brynjolfsson, Li & Raymond, "Generative AI at Work," NBER WP 31161 | 2023 (QJE 2025) | 2 | 5,179 support agents; +14% issues per hour; +34% novice; minimal for experienced; improved sentiment and retention |
| Wong et al., "External Validation of a Widely Implemented Proprietary Sepsis Prediction Model," JAMA Internal Medicine | 21 Jun 2021 (cohort Dec 2018–Oct 2019) | 2 | 38,455 hospitalisations; AUC 0.63; sensitivity 33%; PPV 12% |
| Ostermayer et al., external validation of the Epic sepsis model in two county emergency departments | Cohort Jan–Dec 2023 | 2 | 145,885 encounters; sensitivity 14.7%; PPV 7.6%; median lead time 0 minutes (80% CI −6h42m to +12h) |
| Kaiser Permanente Division of Research, ambient AI scribes at The Permanente Medical Group | 2025 (first year of deployment) | 1 | >2.5m uses; ~16,000 documentation hours; adoption highest in highest-burnout departments; 102 physician survey responses |
| BetterUp Labs and Stanford Social Media Lab, "workslop" research | Sep 2025 | 2 | 1,150 US desk workers; 40% received workslop in the prior month; ~2 hours per incident; ~$186 per employee per month |
| MIT NANDA, "The GenAI Divide: State of AI in Business 2025" | Research Jan–Jun 2025 | 3 | The 95% figure and its exact success definition; the 80/60/50 and 40/20/5 funnels; 67% vs 33% buy-vs-build; 90 days vs nine months; shadow-AI 40% vs 90%; the report's own stated limitations and its internal 50%/70% inconsistency |
| Deloitte Insights, "Agentic AI strategy," Tech Trends 2026 | 2026 (2025 Emerging Technology Trends survey) | 2 | 11% running agentic AI in production; 38% piloting; 14% ready to deploy; 48%/47% data searchability and reusability obstacles; partnerships twice as likely to reach deployment |
| Gartner forecast on GenAI abandonment, as republished by THE Journal | Release 29 Jul 2024 | 2 | ≥30% abandoned after PoC by end-2025; the four stated causes; Sallam quotes; $5m–$20m deployment cost range |
| Gartner forecast on agentic AI cancellations, as republished by BigDATAwire | 25 Jun 2025 | 2 | >40% cancelled by end-2027; "agent washing"; only ~130 of thousands of agentic vendors judged real; Jan 2025 poll of 3,412 attendees |
| CIO Dive on S&P Global Market Intelligence's enterprise AI survey | 14 Mar 2025 | 2 | >1,000 enterprises; 42% abandoning most initiatives (17% in 2024); 46% of PoCs scrapped; cost, privacy and security as top obstacles |
| Silicon Canals report of McKinsey's State of AI survey | Survey mid-2025; McKinsey published 5 Nov 2025 | 2 | 1,993 respondents, 105 countries; 88% using AI; 7% fully scaled; 39% attributing any enterprise EBIT impact, most under 5%; ~6% high performers |
| The Register on the University of Texas System audit of the OEA project | Audit Nov 2016; report 20 Feb 2017 | 2 | $62m total; $39.2m to IBM; fees set just below Board-approval thresholds; $11.59m gift deficit; never in clinical use; ClinicStation/Epic integration gap |
| ABC News on Commonwealth Bank's reversal of AI-attributed redundancies | Cuts Jul 2025; reversal 21 Aug 2025 | 2 | 45 roles; the "error" admission; rising call volumes; FSU dispute at the Fair Work Commission; A$10.25bn FY25 cash profit |
| CX Dive on Klarna's return to human customer service | 2025 | 2 | Siemiatkowski on always guaranteeing a human path; the "Uber-type" flexible agent model |
| CIO Dive on McDonald's ending the IBM automated order-taking test | Jun 2024 | 2 | Partnership since 2021; test purpose stated as operational savings and speed; McDonald's declined to state whether it succeeded or by what metric |
| Restaurant Business on the AOT wind-down memo | 17 Jun 2024 | 2 | Mason Smoot's system message; more than 100 restaurants; shut off no later than 26 Jul 2024 |
| Restaurant Business on the ArchIQ / "Archy" test | 2 Jun 2026 | 2 | Five US locations; Google partnership; ~13,600 US restaurants; dialect and format heterogeneity as the stated obstacle |
| TechCrunch on Taco Bell's voice-AI reassessment | 30 Aug 2025 | 2 | 500+ drive-thrus; 18,000-water-cup prank order; Dane Mathews' "active conversation"; the coaching and segmentation response |
| Retail Dive on Just Walk Out's removal from Amazon Fresh | 2024 | 2 | Replacement by Dash Carts; Amazon's rebuttal of the human-reviewer characterisation, quoted in full in §7 |
| MIT Technology Review on Amazon's scrapped recruiting model | 10 Oct 2018 (development from 2014) | 2 | Five-star ranking; penalties on "women's" and all-women colleges; project killed after loss of confidence in neutrality |
| MIT Technology Review on the Google Health diabetic-retinopathy deployment in Thailand | 27 Apr 2020 (field study 2018–19, Beede et al., CHI 2020) | 2 | 11 clinics; >90% lab accuracy; more than a fifth of images rejected; upload delays; the nurse's "10 patients in two hours" quote; 4.5m patients to ~200 specialists |
| Holland & Knight on the preliminary collective certification in Mobley v. Workday | Order 16 May 2025 | 2 | Nationwide ADEA collective; applicants 40+ since 24 Sep 2020; rejection of the "just a software provider" defence; Mobley's 100+ applications |
| The Markup on NYC's MyCity chatbot giving illegal advice | 29 Mar 2024 (launch Oct 2023) | 2 | Advice that landlords may refuse housing-voucher tenants and employers may take tips |
| The Markup on the chatbot's termination | 30 Jan 2026 (shut down Feb 2026) | 2 | "Functionally unusable"; ~$500,000 annual cost; ~$600,000 reported build cost; termination as a budget measure by a new administration |
| Fortune on Lattice withdrawing its "digital workers" feature | Announced 9 Jul 2024; withdrawn 12 Jul 2024 | 2 | Three-day reversal; Franklin's "questions that have no clear answers yet" statement |
| CFO Dive on Deloitte's partial refund to the Australian government | Oct 2025 | 2 | A$97,000 refunded on a ~A$440,000 contract; fabricated references; Azure OpenAI use disclosed in the revised report |
| The Register on Builder.ai's insolvency | 21 May 2025 | 2 | >$500m raised; apps built by the team rather than by the advertised automation; prior scrutiny as Engineer.ai |
| CNBC on Salesforce's customer-support headcount reduction | 2 Sep 2025 | 2 | Support workforce from ~9,000 to ~5,000; Benioff's "I need less heads" framing |
| BBC News on the DPD chatbot failure | Jan 2024 | 2 | Failure introduced by a system update to a chatbot run "successfully for a number of years"; 800,000 views in 24 hours; component disabled |
| NPR on NEDA taking down the Tessa chatbot | 8 Jun 2023 | 2 | Chatbot dispensing dieting and calorie-counting advice to an eating-disorder population; withdrawal |
| Presto Automation Form 10-K, FY ended 30 Jun 2023 | Filed 11 Oct 2023 | 1 | Company's own claim of "approximately 95% non-intervention rate at certain locations"; Hi Auto subcontract disclosed — the filing against which the SEC order should be read |
| European Commission, "AI Omnibus enters into force" | In force 27 Jul 2026 (OJ 24 Jul 2026) | 1 | Annex III standalone high-risk obligations deferred to 2 Dec 2027 and Annex I embedded systems to 2 Aug 2028; prohibitions, AI-literacy, GPAI and Article 50 transparency duties unchanged |
| Mobley v. Workday, order on the Second Amended Complaint (N.D. Cal., Dkt 267) | 6 Mar 2026 | 1 | FEHA counts dismissed with leave to amend for want of a California nexus; ADEA disparate-impact claim survives; AARP granted leave to file amicus |
| Civil Rights Litigation Clearinghouse docket, Mobley v. Workday, 3:23-cv-00770 | Docket to 2026 | 1 | Case chronology: Jul 2024 agency ruling; 16 May 2025 preliminary certification; Jul 2025 HiredScore scope order; Jan and Mar 2026 amendments and orders |
| Federal Reserve Bank of Atlanta Working Paper 2026-4, "Artificial Intelligence, Productivity, and the Workforce" | Mar 2026 | 1 | ~750 corporate financial executives; reported labour-productivity growth +2.4pp (2025) and +3.3pp expected (2026) against +1.0pp and +1.8pp implied; aggregate employment expected to fall <0.4% in 2026; routine clerical share down >2pp over three years |
| Federal Reserve FEDS Note, "The AI Buildout and the Economy" | 17 Jul 2026 | 1 | "Labor market impacts remain concentrated and have not yet broadened in the aggregate"; the evidence read as "a buildout phase rather than the onset of broad-based displacement" |
| "Multicenter Prospective Validation of an Updated Proprietary Sepsis Prediction Model," JAMA Network Open | 27 Feb 2026 | 2 | Epic Sepsis Model v2 across 4 US health systems and 227,091 encounters; AUROC 0.82–0.92; PPV 0.13–0.26 at 60% sensitivity; median lead time 1.9–10.3 h; authors recommend local validation before deployment |
| METR, "We are Changing our Developer Productivity Experiment Design" | 24 Feb 2026 | 2 | The follow-on experiment "gives us an unreliable signal"; 30–50% of developers declined to submit tasks without AI; authors' own view that developers are likely more sped up in early 2026 |
| Duane Morris analysis of the 22 Jun 2026 order on the Third Amended Complaint in Mobley | 24 Jun 2026 | 2 | Granted in part and denied in part; extraterritoriality argument rejected — wrongful conduct within California is not "extraterritorial" regardless of applicants' locations |
| STAT on internal IBM documents concerning Watson for Oncology | 25 Jul 2018 (decks Jun–Jul 2017) | 2 | "Multiple examples of unsafe and incorrect treatment recommendations"; customers describing output as "often inaccurate"; training on a small number of synthetic rather than real patient cases |
| Beede et al., "A Human-Centered Evaluation of a Deep Learning System Deployed in Clinics for the Detection of Diabetic Retinopathy," CHI 2020 | 2020 (field study 2018–19) | 2 | Eleven clinics in Thailand; the study's own framing of "tensions between the model's thresholds for data quality, and the quality of data that arise from an imperfect, resource-constrained environment" |
Not verified / not load-bearing — stated plainly so no conclusion rests on it
Sources that could not be retrieved this session, and how that was handled. Gartner's own website, McKinsey's, BCG's, the ACM Digital Library, JAMA Network and the Boston Consulting Group's January 2026 AI Radar release all refused automated retrieval. Where a Gartner figure is used, it is cited to a trade publication that reproduced the release verbatim (THE Journal for the July 2024 forecast, BigDATAwire for the June 2025 agentic forecast), and labelled as such. McKinsey's survey figures are cited to a trade report of the survey, not to McKinsey, and the confidence attached to them is correspondingly lower. Beede et al.'s CHI 2020 paper is unreachable through the ACM Digital Library, but the authors' own publication page is not, and it is cited directly in §6 alongside MIT Technology Review's contemporaneous account, which quotes the study and its participants. Three further primaries refused automated retrieval and are named rather than quietly replaced: Reuters' original report of the Amazon recruiting model (401), which MIT Technology Review's account stands in for; CanLII and the BC tribunal's own site for Moffatt (403), for which a full PDF of the reasons is cited instead; and the Texas Attorney General's release (intermittently 402/200), which was reachable and is cited. BCG's 2026 figures were not obtainable and no BCG number appears in this brief.
Numbers circulating widely that this brief declines to use. Several figures in current circulation could not be traced to a retrievable primary and are therefore excluded: a "74% of enterprises have rolled back a customer-facing AI agent" statistic attributed to a vendor report; an "89% agent failure rate" attributed to Deloitte, which appears to be a misreading of Deloitte's 11%-in-production figure (the complement of "in production" is not "failed"); and various 2026 "AI ROI" percentages attributed to composite surveys whose underlying instruments are not published. Gartner has since published a figure of at least 50% of GenAI projects abandoned after proof of concept by the end of 2025 — an outturn rather than the 2024 forecast, and one that would, if citable, confirm this brief's own reading. It is not used, for a narrow reason worth stating precisely: every Gartner surface carrying it refuses automated retrieval, so the figure could not be read in a Gartner document this session. Only the July 2024 ≥30% forecast, verified through a verbatim trade republication, appears here. Readers should treat this brief's abandonment evidence as conservative on that account.
Case details deliberately not asserted. Widely repeated specifics of the MD Anderson contract — an original six-month, $2.4 million scope, twelve amendments, $39.2 million paid to IBM and roughly $21–23 million to PwC — appear in secondary accounts that could not be opened this session (The Cancer Letter is paywalled; the JNCI article and Medscape refused automated retrieval). The audit itself — since removed from utsystem.edu — has now been retrieved from the Internet Archive's capture and is cited directly in §1. It confirms the contract history: a six-month, $2.4 million original IBM agreement extended twelve times to $39.2 million in total fees, plus approximately $21.2 million of OEA-related PwC fees. Every audit quotation in §1 was checked against the retrieved text verbatim; the $62 million combined total is asserted as before. Epic's subsequent overhaul of the model is verified and is now treated in §6 on its own evidence: the updated version was prospectively validated across four US health systems and 227,091 encounters and published in February 2026. The argument in §6 concerns version one as externally validated in 2019 and 2023 and as deployed for roughly six years, and it is not weakened by the improvement — but the improvement is real and is stated rather than omitted. What remains unverified is any claim about how many hospitals have moved to the updated model or retrained it locally; no number is asserted. Amazon's characterisation of the human role in Just Walk Out is disputed by Amazon and both positions are given in §7; no claim about the true intervention rate is made. The Klarna "$40 million profit improvement" and "700 full-time agents" figures are Klarna's own estimates, never audited or restated in a financial filing, and no conclusion here rests on them. McDonald's never disclosed the success metric for the IBM test, so no inference is drawn about whether it met one.
Named cases versus archetypes. Every case treated at length in Part I is a named, publicly documented programme with a dated source: MD Anderson, Watson for Oncology beyond it, Google Health Thailand, the Epic Sepsis Model, Air Canada, DPD, NEDA, Commonwealth Bank, Klarna, Taco Bell, McDonald's, Just Walk Out, Presto, Zillow, Amazon's recruiting model, iTutorGroup, Workday, Rite Aid, Pieces, Builder.ai, Delphia, Global Predictions, DoNotPay, NYC MyCity, Lattice, Deloitte Australia, Salesforce and The Permanente Medical Group. Figure 11a contains no archetype. Exactly one entry in this brief is a composite rather than a company: failure mode 2 in Figure 13, "value not ownable," whose defining case is labelled in the figure as a representative archetype, not a specific company — commodity chat deflection and generic productivity assistants. It is drawn from the documented pattern across the cases above and no company is implied by it. Every other row in that figure names a real programme.
Under-evidenced areas. Manufacturing and industrials are materially under-represented in the public case record used here; the sector row in Figure 22 is flagged rather than filled, and the corresponding claims are held at low confidence. Public-sector evidence outside New York City and the Australian Deloitte engagement was not developed — Australia's Robodebt Royal Commission report could not be retrieved this session and is not cited. No European or Asian enterprise case is developed at length; the Danish labour-market study is the only non-Anglophone primary. Readers should treat the geographic generality of the taxonomy as an inference, not a demonstration.
Where this brief is most likely to be wrong. The capacity-conversion mechanism rests substantially on one very good study of one small, high-trust, high-union-density labour market. Denmark's institutions make headcount conversion harder than in the United States, which could mean the 85% reallocation figure is an upper bound rather than a base rate. The Census Bureau's 2026 microstructure paper is an independent US corroboration by a different method — cross-sectional regression on a probability sample rather than difference-in-differences on linked payroll — and it points the same way, which raises confidence materially. It does not eliminate the risk: both are observational, neither identifies a causal effect of an operating change on headcount, and a well-identified US study finding otherwise would require the weight placed on Gate 4 to fall.
This is an internal educational and analytical reference on why enterprise AI pilots fail to deliver value. It is not investment advice, not legal advice, and not an endorsement or criticism of any vendor, employer or public body beyond what the cited public record supports. Litigation described here that has not reached judgment — including Mobley v. Workday — is at a procedural stage only, and no finding of liability is implied. Evidence cutoff: 26 August 2026.