The Primer Desk.↗ ShareDiscuss on X
2026-08-27·34,499 words·67 sources·~138 min read

The argument — stated once, then earned

In September 2016, MD Anderson switched off an IBM Watson system it had spent $62 million teaching to recommend cancer treatments. The Oncology Expert Advisor had never advised on the treatment of a single patient outside a test. It did not fail for lack of capability. It failed because they had not built the path from the system to the actual practice of medicine.

The consensus explanation for the enterprise AI disappointment — the models aren't good enough, the data isn't ready — describes a minority of the wreckage. The majority of pilots that produce nothing are pilots where the technology worked. The demo was crisp. The accuracy target was hit. Some even shipped, ran in production for a year, and still created no tangible value.

Getting the right answer clears one gate out of four. The system still has to be pointed at a problem that is financially worth solving. It has to land inside a workflow, in front of people willing to change how they work. And then comes the gate almost nobody plans for: the capacity it frees or the revenue it wins is only real when it shows up on the bottom line — and somebody has to put it there. That is the distance between technical success and value delivery. Value leaks at every one of these gates, and each leak traces back to an operating choice made, or not made, at pilot design time — usually months before a model is trained.

Everything that follows turns on what counts as a pilot and what counts as value — two terms management uses freely and rarely defines. A pilot here is any funded, bounded attempt to put AI in front of real work with the intention of expanding it if it works — a proof of conceptA short, small-scale build whose only job is to show that something can work: a sample of data, a handful of users, a few weeks. It establishes feasibility, not value, and it is usually run outside the process that would govern a real system., a limited production rollout, a departmental deployment. Value is something a controller can find: a cost line that fell, a revenue line that rose, a mission outcome measurably better, cash. Hours "saved" don't qualify, and neither do satisfaction scores or adoption numbers. Saved hours are an input to value — and, as the evidence below shows, they are where the trail goes cold in most enterprises.

By the numbers
$62mSpent at MD Anderson — never advised a patient outside a test
95%Of organizations getting zero return on GenAI (MIT NANDA)
19%Slower with AI — while feeling 20% faster (METR trial)
85%Of users' AI time savings reallocated, never banked
70%Of orders still needed a human at the "human-free" drive-thru AI (SEC)
$407.9mZillow's one-year write-down — scaled before the model's error was priced
Evidence: §1 (UT System audit), §15 (MIT NANDA), §12 (METR; NBER WP 33777), §7 (SEC v. Presto), §8 (Zillow FY21 10-K) — each cited in place.

Executive takeaways — the seven claims this brief proves

  1. 1The famous "95% of AI pilots fail" number is roughly right about the outcome and badly wrong about the cause — and it is not a measurement of pilots. MIT NANDA's own text — NANDA being a research project at the MIT Media Lab, and the study an interview-and-survey exercise rather than an audit of company accounts — defines success as tools "users or executives have remarked as causing a marked and sustained productivity and/or P&L impact," from 52 interviews and 153 conference surveys, and warns its figures are "directionally accurate based on individual interviews rather than official company reporting." Moderate confidence
  2. 2The dominant failure mode is capacity that is freed and never banked. In the largest linked survey-to-payroll study available — one that sets what workers say they do beside what their employers actually paid them — 85% of chatbot users report reallocating their time savings to other tasks, and earnings and hours across 25,000 Danish workers show precise null effects: a result sharp enough to place the true effect at about zero, rather than a failure to detect one. It rules out changes larger than 2%. Time saved is not money saved unless someone converts it. High confidence
  3. 3Self-reported productivity is systematically, measurably wrong — in the optimistic direction. In METR's randomised trial, experienced developers were 19% slower with AI tools while estimating afterwards that they had been 20% faster — a 39-point gap between the felt effect and the measured one, inside the same people, on the same tasks. METR has since said the sign of that result should not be carried forward; the gap should. At the firm level the Atlanta Fed found the same shape in ~750 CFOs: reported AI-related productivity gains of 2.4 percentage points in 2025 against 1.0 implied — "implied" meaning what those same firms' own reported inputs and outputs work out to. Pilot business cases built on self-report inherit that gap at both levels. High confidence
  4. 4Buying beats building roughly two to one, and the gap is about maintenance, not talent. External partnerships reached deployment ~67% of the time against ~33% for internal builds in the NANDA sample, with employee usage rates nearly double for externally built tools. Moderate confidence
  5. 5A pilot's most expensive failures are silent ones that survive in production. Epic's sepsis model — live across hundreds of US hospitals — scored an AUC of 0.63 in external validation at Michigan, and in a 2023 two-hospital emergency-department cohort of 145,885 encounters returned a sensitivity of 14.7% and a median warning lead time of zero minutes. Three words are carrying that sentence. AUC scores how cleanly a model separates the patients who will deteriorate from those who will not, on a scale where a coin flip scores a half and a perfect separator scores one; sensitivity is the share of real cases the model actually flags; lead time is how far ahead of the event the warning arrives. External validation means somebody other than the vendor tested it, on a different hospital's patients. Nothing about that is visible on a dashboard. High confidence
  6. 6Regulators and courts now price the operating gap directly, which converts a soft governance failure into a hard cash one. A tribunal held Air Canada to its chatbot's answer and called the argument that the bot was a separate legal entity "a remarkable submission"; the SEC ordered Presto Automation to cease and desist after finding that its "human-free" voice AI needed an off-site human on 70% of orders at its best pilot sites and on 100% of orders at the substantial majority of locations. High confidence
  7. 7The programs that work are boring, narrow and owned — and they look nothing like the portfolio most enterprises are running. The Permanente Medical Group's ambient scribes — software that listens to a clinical visit and drafts the note the doctor would otherwise type up afterwards — reached 2.5 million uses and ~16,000 documentation hours in one year by attacking the single task clinicians most hated, with no claim of headcount reduction attached. Adoption concentrated where the pain was worst. High confidence

PART I — THE PILOTS

Fifteen sections, cases first and pattern after. Start from the top and you get a list of technology complaints; start from the cases and you get a list of operating mistakes.

1 · A $62 million cancer engine that never saw a patient — MD Anderson & IBM Watson, 2013–2016

The Oncology Expert Advisor was supposed to read a patient's chart, read the literature, and tell an oncologist what to do next. MD Anderson had been running it with IBM since 2014. By the time it stopped, the centre had, in the words of the University of Texas System audit that surfaced in February 2017, "sent US$62 million in the general direction of Big Blue and PwC" — without going through its normal IT procurement process.

The audit did not find what the headlines said it found. It stayed deliberately out of the science. Its findings were about governance: work performed under an amended scope that "extended beyond the OEA project and intent of funding as approved by the Board of Regents"; invoices "paid in full regardless of whether contracted services were delivered as agreed upon"; much of the money spent without competitive tender, with fees "consistently set just below the amount that would have required Board approval"; and $11.59 million of donor gifts spent before they had been received.

The system had been built against ClinicStation, the medical records system MD Anderson used before it moved to EpicThe dominant electronic health record system in US hospitals — the software holding the digital chart, the orders and the notes. Every clinical action passes through it, so whether a tool is built inside the record or beside it largely decides whether a clinician ever sees its output.. It was never rebuilt against Epic. Audit staff told the reviewers that OEA's drug protocol and clinical trial data "is now outdated and must be updated before OEA can be piloted again within MD Anderson." IBM ended support in September 2016. The system had never been piloted anywhere else.

So: a decision-support tool whose entire selling point was up-to-date medical knowledge, sitting on a record system the institution no longer used, holding drug protocols that had gone stale. Whether the underlying model was good is, at that point, an academic question. The value was never reachable. The organisation moved and the pilot did not, because nothing in the pilot's design made anyone responsible for the pilot moving with it.

The epilogue is corporate. IBM had assembled Watson Health through several billion dollars of acquisitions; in June 2022 Francisco Partners completed the purchase of the healthcare data and analytics assets and relaunched them as a standalone company, Merative, headquartered in Ann Arbor. The flagship demonstration of AI in medicine did not so much fail as get sold, which is a different and quieter kind of ending — and one that leaves no post-mortem behind.

MD Anderson was not the whole of it, and the two failures are different in kind. Watson for Oncology was sold to hospitals well beyond Houston, and in July 2018 internal IBM presentations obtained by STAT, decks prepared in June and July 2017 by the division's own deputy chief health officer and circulated to Watson Health management, recorded "multiple examples of unsafe and incorrect treatment recommendations." Customers described the output as "often inaccurate," raising "serious questions about the process for building content and the underlying technology." The documents name the mechanism plainly: the system had been drilled on a small number of synthetic cancer cases — hypothetical patients — rather than on real patient data, with recommendations derived from a handful of specialists per cancer type rather than from guidelines or evidence.

At MD Anderson the model's quality was never the binding constraint; the integration path was. Everywhere else, the constraint was that the evaluation set bore no resemblance to the population. One is a Gate-3 failureA failure of workflow integration. The model works, but its output never reaches the right person at the right moment in a form they will act on. The next section sets out all four gates a pilot must pass; this is the third. and the other a Gate-2 failureA failure of model capability — the system does not do the job accurately or reliably enough under real conditions. This is the only one of the four gates the AI industry sells tooling for. wearing a Gate-2 evaluation that could not have caught it — which is the vendor-demo mechanism of §10, four years early and inside the most famous AI programme of its decade.

Figure 1Ambition to autopsy: the OEA chain, and where it actually broke
  1. 01
    Ambition
    A Watson-powered advisor reading charts and literature to recommend oncology treatment and match trials.
  2. 02
    Setup
    Procured outside the normal IT process; much of the spend untendered; fees kept just under Board-approval thresholds; scope amended beyond the funding the Board had approved.
  3. 03
    What happened
    $62m to IBM and PwC combined. Built on ClinicStation; the institution moved to Epic. IBM ended support in 2016; the project was terminated that September.
  4. 04
    Why value never landed
    Never in clinical use; never piloted outside MD Anderson; protocol and trial data stale before restart was possible.
  5. 05
    Transferable root cause
    The pilot was funded as a science project and governed as a gift, so no one owned the integration path the value ran through.
Evidence: University of Texas System Audit Office special review of OEA procurement (released 11 Nov 2016; since removed from utsystem.edu, retrieved from the Internet Archive's capture and quoted directly), corroborated by The Register's contemporaneous account, 20 Feb 2017.
Method: Chain reconstructed from the audit's findings as quoted directly by The Register. The $62m is the audit's combined total across IBM and PwC; the split between the two vendors is reported inconsistently elsewhere and is not asserted here. The audit explicitly stayed "out of the scientific merits of the project," so no claim is made about model quality.
Synthesis: AIR classification — a Gate-3 failure (workflow integration), not a Gate-2 failure (model capability). The determining fact is an EHR migration, which is an IT-portfolio event, not an AI one.
The lessonA pilot procured outside the process that keeps its underlying systems running has a shelf life measured in the next platform migration.

2 · Four gates, one leak — the frame the rest of the cases fill in

Strip the vocabulary away and every enterprise AI pilot has to pass through four gates in order. Miss any one and the money does not arrive, regardless of how well the others went.

Gate 1 — is there value on the table? Is the problem worth solving, is the value ownable by this organisation, and would anyone pay for the solved version? A great many pilots die here and the death certificate says something else, because Gate 1 is assessed before any measurement exists.

Gate 2 — does the model work? Accuracy, latency, cost per call, reliability under the real distribution rather than the sampled one. This is the only gate the AI industry sells tooling for, and it is the gate enterprises are best at. It is also the gate that matters least often.

Gate 3 — does the work change? Does the output land inside a process, in front of a person with the authority to act on it, at the moment the decision is made, in a form that survives the person's judgment about whether to trust it? Deployment, on its own, is none of this — it only makes it possible.

Gate 4 — does anyone collect? Freed capacity has to be converted — into headcount not hired, hours redeployed to revenue work, a service level raised that a customer pays for, a loss avoided. Someone must own that conversion, and it is almost never the same person who owned the pilot.

"The tech isn't ready" survives as the standard explanation because it is the only one of the four a technologist can diagnose alone. The other three require somebody to look at the P&L and the org chart, and those are not in the pilot team's job description.

Figure 2The four gates — and which ones enterprises actually manage
  1. 01
    Value on the table
    Is the problem worth money, is the value ownable, and is the counterfactual (do nothing / buy it / fix the process) worse? Assessed before any measurement exists, which is why it is usually skipped.
    Rarely managed
  2. 02
    The model works
    Accuracy, latency, unit cost, reliability under the production distribution. The one gate with a mature tooling market and a clear owner.
    Well managed
  3. 03
    The work changes
    Output reaches a decision-maker at the decision moment, in a form they trust and act on. Requires process redesign, not deployment.
    Partly managed
  4. 04
    Someone collects
    Freed hours, avoided losses or won revenue are converted into a P&L line by a named owner with the authority to do so. Nearly always unowned.
    Rarely managed
Evidence: Derived from the case record assembled in this brief (MD Anderson OEA; Epic Sepsis Model; Commonwealth Bank; Humlum & Vestergaard, NBER WP 33777; MIT NANDA 2025), each cited in place.
Method: Qualitative classification, not a measured distribution. The "managed" verdicts reflect whether a named owner and a measurement exist at each gate in the cases reviewed, not a survey.
Synthesis: AIR framework. Enterprises are strongest exactly where the failure rate is lowest, which is why more model investment does not move the outcome.
Figure 3Two scoreboards that look alike and share almost no entries
Technical successWhat the pilot review measures
Accuracy / F1 against a held-out set · latency at p95 · uptime · cost per 1,000 calls · number of users provisioned · "in production" · stakeholder satisfaction survey · demo completed for the board.
Value deliveryWhat a controller can find
Cost line that fell versus a pre-agreed baseline · revenue attributable against a holdout · headcount not hired · loss ratio moved · cycle time that a customer pays for · capacity converted, by a named owner, with a date.
Evidence: Left column reflects the success criteria described in MIT NANDA (2025) and Gartner's July 2024 abandonment analysis; right column reflects the EBIT-attribution test used in McKinsey's Nov 2025 survey and the earnings/hours test used in NBER WP 33777.
Method: Editorial contrast, not a paired dataset. No claim that the two columns were measured on the same population.
Synthesis: Almost every entry on the left can be fully satisfied while every entry on the right stays flat. That is the definition of the value gap.

3 · The deflection trap — Commonwealth Bank and Klarna, 2024–2025

A tool can be logged in by everyone and used by nobody, or used hard and still leave the work where it was. The two programs here are the ones most 2026 pilot budgets resemble — a generative assistant put in front of customer contact and measured on how many conversations it keeps away from a human — and both published their numbers, made the headcount decision, and reversed it in public.

Commonwealth Bank of Australia, July–August 2025. CBA announced 45 redundancies in its call centre, citing technology including AI: "Our investment in technology, including AI, is making it easier and faster for customers to get help, especially in our call centres." The Finance Sector Union raised a dispute at the Fair Work Commission, and its members reported that call volumes had risen after the voice-bot went in, with management offering overtime and pulling team leaders onto the phones. On 21 August the bank reversed the redundancies and called the decision an "error", admitting it "did not adequately consider all relevant business considerations" and "should have been more thorough in our assessment of the roles required." Affected staff were offered their jobs back. This from a bank that had just posted a record A$10.25 billion cash profit — the money was never the constraint.

Nothing about that outcome is rare, and the arithmetic behind it runs backwards from intuition. A deflection botA contact-centre assistant whose measured job is to stop calls and chats from reaching a human. Its headline metric — the deflection rate — counts contacts intercepted, not problems solved, which is why it can rise while the work left for people gets harder and larger. takes the simple calls, which raises the average difficulty of everything still reaching a human. And when it handles a call badly, it generates a second one. Deflection rate goes up. Human workload does not go down proportionally, and may go up. The bank measured the first number and made a headcount decision on it.

Klarna, February 2024 to 2025. The most-cited enterprise GenAIShort for generative AI: systems that produce new text, images, code or speech rather than classifying or scoring something that already exists. The large language model behind a chat assistant is the common enterprise case. The distinction matters because generative systems fail differently — an older model gives a wrong score, a generative one writes a fluent wrong answer. success of its year. Klarna's own release of 27 February 2024 reported the OpenAI-powered assistant handling 2.3 million conversations in month one — two-thirds of all customer-service chats — "doing the equivalent work of 700 full-time agents," resolving errands in under two minutes against eleven previously, with a 25% drop in repeat inquiries and an estimated "$40 million USD in profit improvement to Klarna in 2024." Every number there is Klarna's own; none was independently audited.

By 2025 the company was recruiting humans again, with CEO Sebastian Siemiatkowski settling on a rule his 2024 self would not have written: "From a brand perspective, a company perspective… I just think it's so critical that you are clear to your customer that there will be always a human if you want." Klarna is now building an "Uber-type" flexible agent pool. This is not a failure — the automation of routine contacts held. What failed was the inference from "the bot handles two-thirds of chats" to "we need proportionally fewer people," which is the same inference CBA made.

One inference cuts across both programs. Deflection metrics and value metrics point in opposite directions more often than anyone admits, because deflecting the easy half of a workload raises the cost per remaining unit and multiplies the failure cost of the deflections that go wrong. An assistant that handles 70% of contacts may reduce total cost by far less than 70%, by nothing, or — where it generates repeat contact — by a negative number. CBA is the clean natural experiment: same firm, same quarter, the deflection number rose and the workload rose with it.

The lessonA deflection rate is not a cost line. The easy contacts leave first, the hard ones stay, and the failed deflections come back as second calls — which is why headcount decisions made on the deflection number keep getting reversed in public.

4 · The drive-thru lesson — McDonald’s and Taco Bell, 2021–2026

Voice ordering is the most public pilot program in America — it runs outdoors, in front of customers holding phones, so failures are filmed and retreats are announced. Two of the largest quick-service operators ran it at hundreds of sites; what each did next is the useful part.

Taco Bell, 2023–2025. Voice AI ordering rolled out to more than 500 drive-thrus on Yum Brands' Byte platform. Then the viral clips: an order for 18,000 cups of water among them. Chief digital and technology officer Dane Mathews told The Wall Street Journal in August 2025 that the company was in an "active conversation" about where to use it and where not, and the resolution is the interesting part: the company segmented instead of withdrawing. "For our teams, we'll help coach them: at your restaurant, at these times, we recommend you use voice AI or recommend that you actually really monitor voice AI and jump in as necessary." Busy restaurants with long lines may keep a human on the mic.

McDonald's, 2021–2026. Two years of automated order taking with IBM across more than 100 restaurants, ended by a system message from chief restaurant officer Mason Smoot in June 2024: "While there have been successes to date, we feel there is an opportunity to explore voice ordering solutions more broadly… the technology will be shut off in all restaurants currently testing it no later than July 26, 2024." McDonald's declined to say whether the test had succeeded or by what metric it was judged — a small silence that says a lot about how the pilot was scoped. The company nonetheless said the work "has given us the confidence that a voice-ordering solution for drive-thru will be part of our restaurants' future."

Two years later that confidence has a name. In June 2026 McDonald's began testing ArchIQ, with a voice assistant nicknamed "Archy," at five US locations on its Google Cloud partnership — five, against roughly 13,600 US restaurants. The scope collapse from "more than 100" to "five" is the most honest number in the entire story.

Figure 4The gap between the benefit assumed and the benefit observed
ProgramDatesBenefit assumed at designWhat was observedResolution
CBA voice-bot
Retail banking, AU
Jul–Aug 2025Fewer inbound calls → 45 roles redundantCall volumes rose; overtime offered; team leaders on phonesRedundancies reversed; bank calls it an "error"
Klarna AI assistant
Fintech, SE
Feb 2024 →2.3m chats, "work of 700 agents", $40m profit uplift (company estimate)Routine automation held; complex and emotive cases degradedHuman agents re-recruited; guaranteed human path
Taco Bell voice AI
QSR, US
2023–2025Faster service, better accuracy, automated upsell at 500+ sitesOrder errors, interruptions, viral prank ordersSegmented by site and daypart; humans retained at busy stores
McDonald's AOT (IBM)
Quick service, US
2021–Jul 2024Operational savings and speed of service across 100+ sitesNot disclosed; company declined to state the success metricEnded; relaunched Jun 2026 as ArchIQ at 5 of ~13,600 US sites
Evidence: ABC News (21 Aug 2025) on CBA and the FSU dispute; Klarna press release (27 Feb 2024) and CX Dive (2025) on the reversal; TechCrunch on Mathews' WSJ interview (30 Aug 2025); Restaurant Dive (17 Jun 2024) and Restaurant Business (2 Jun 2026) on McDonald's.
Method: "Benefit assumed" is the sponsor's own stated rationale at launch, not a third-party estimate. Klarna's figures are self-reported and unaudited; McDonald's success metric was never disclosed, so that row's assumption is taken from the company's stated test purpose.
Synthesis: In all four, the technology performed close to specification on the narrow task. The value assumption that failed was about the rest of the system — call mix, escalation, queue length, franchise variance.
The lessonWatch the relaunch scale, not the launch scale. McDonald’s went from more than 100 sites to five; the scope of the second attempt is the honest measure of what the first one proved.

5 · The chatbot that made a contract — Moffatt v. Air Canada, 2024 BCCRT 149

In November 2022 Jake Moffatt's grandmother died and he went to Air Canada's website to book a flight to Toronto. He asked the site's chatbot about bereavement fares. It told him he could book at full price and apply for the bereavement rate within 90 days. The airline's actual policy, on a page titled "Bereavement travel" elsewhere on the same site, said the opposite: no retroactive claims. He paid $1,630.36, applied, and was refused.

Air Canada's defence at the British Columbia Civil Resolution Tribunal is the reason this case matters. Tribunal member Christopher C. Rivers recorded it plainly: the airline argued it could not be held liable for information provided by "one of its agents, servants, or representatives — including a chatbot," and, as Rivers put it, "in effect, Air Canada suggests the chatbot is a separate legal entity that is responsible for its own actions. This is a remarkable submission."

The tribunal answered in a few plain sentences that are, for anyone deploying a customer-facing assistant, load-bearing. "While a chatbot has an interactive component, it is still just a part of Air Canada's website. It should be obvious to Air Canada that it is responsible for all the information on its website. It makes no difference whether the information comes from a static page or a chatbot." On the argument that Moffatt should have checked the other page: the airline "does not explain why the webpage titled 'Bereavement travel' was inherently more trustworthy than its chatbot… nor why customers should have to double-check information found in one part of its website on another part of its website." Damages of $650.88, plus $36.14 interest and $125 in fees. A rounding error, and entirely beside the point.

Everyone already knew a chatbot can hallucinateProduce a confident, fluent, well-formed answer that is simply invented. It is not a bug in the ordinary sense — the system is doing what it was built to do, which is generate plausible text, and plausibility and truth are not the same target. The dangerous property is that a hallucination looks exactly like a correct answer.. What the tribunal established is that the enterprise's own statement of what the assistant is — a helper, a convenience, an experiment, explicitly not authoritative — has no legal weight against a customer who reasonably relied on it. Every disclaimer in a deployment plan is an internal document. The counterparty gets to treat the output as the company speaking, because it is.

Price that. A conversational assistant deployed against a policy surfaceAny body of rules a company publishes and is expected to honour — fares, refunds, warranties, eligibility, pricing. When an assistant is pointed at one, it stops being a search box and starts being a mouth: whatever it says about the rule is what the company has said about the rule. is not a productivity tool with a small error rate; it is an unbounded, unmonitored authority to make representations on the company's behalf, at machine volume, without a review step. Most pilot business cases model the cost of a wrong answer as a customer-satisfaction ding. The correct model is the expected cost of a promise the company must honour, multiplied by conversation volume. Change that one line and a meaningful share of customer-facing assistant pilots stop clearing their hurdle rateThe minimum return a company requires before it will spend money on something. Every investment competes against it; a project that does not clear the hurdle is refused not because it loses money but because the money does better elsewhere. — which is the honest reason many of them were quietly narrowed after February 2024.

In January 2024 the parcel firm DPD disabled part of a chatbot it had run "successfully for a number of years" after a system update caused it to swear at a customer and compose a poem about how terrible the company was; the customer's account was viewed 800,000 times in a day. The failure was introduced by an update, not by the original design. The risk is not retired at launch, and a system that has behaved for years is not thereby safe.

The more serious version came in June 2023, when the National Eating Disorders Association took down its Tessa chatbot after users showed a newer version dispensing dieting and calorie-counting advice to a population for whom that advice is precisely the hazard. A helpline replacement that gives contraindicatedMedically the wrong thing to give this patient — advice or treatment that is fine in general but actively harmful for the condition in front of you. guidance has not underperformed its target; it has inverted its purpose. Any pilot whose failure mode is the opposite of its objective — rather than merely less of it — needs a different risk model than a productivity tool, and almost none get one.

Representational exposureThe liability an organisation takes on when an automated system makes statements a customer, regulator or counterparty can reasonably rely on. It scales with conversation volume, not with error rate, and it is not reduced by an internal disclaimer.

6 · The pattern predates the technology — two deep-learning deployments, 2018–2024

Every case so far is generative AI, and every one is recent. The two in this section are neither — deep-learning classifiers designed and deployed before anyone had typed a prompt into ChatGPT — and they are here on purpose. The most comfortable rebuttal is that the failures above are teething problems of a young technology, cured by the next model generation. These two deployments are the control group for that claim. The technology was different, the vendors were different, the decade was different, and the leak was the same.

Google Health in Thailand, 2018–2019. Thailand's ministry of health had a target: screen 60% of diabetic patients for diabetic retinopathy, the leading cause of preventable blindness in that population. It had roughly 4.5 million patients and about 200 retinal specialists. The existing process — nurses photograph the eye, a specialist elsewhere reads the image — could take ten weeks. Google Health had a deep-learning systemA model trained by being shown a very large number of labelled examples — here, retinal photographs a specialist had already graded — until it learns the pattern for itself. Nobody writes the rules; the rules are inferred from the examples, which is why the examples the system was shown decide what it can handle. that identified the condition from a retinal scan at better than 90% accuracy, which the team characterised as human-specialist level, and could return a result in under ten minutes.

This is as clean a value hypothesis as enterprise AI ever gets. A real bottleneck, a quantified queue, a model that beats the queue on both accuracy and latency. The team deployed to eleven clinics and then did the thing almost nobody does — they watched.

The model had been trained on high-quality scans and, to protect its accuracy, was designed to refuse images below a quality thresholdThe cut-off an engineer sets on a model's own confidence score, deciding when it acts, warns, or declines. It is a dial, not a property of the model: turn it one way and you get more misses, the other way and you get more false alarms. Somebody chooses where it sits, and that choice is where a technical decision quietly becomes an operating one.. Nurses were photographing dozens of patients an hour in rooms with poor lighting. More than a fifth of images were rejected outright. A rejected image meant the patient was told to visit a specialist at another clinic on another day — the exact outcome the system existed to prevent — and nurses, who frequently believed the rejected scans showed no disease, burned time trying to retake or edit them. Because processing ran in the cloud, clinics with slow connections queued. One nurse: "Patients like the instant results, but the internet is slow and patients then complain. They've been waiting here since 6 a.m., and for the first two hours we could only screen 10 patients."

The team's own write-up, presented at CHI 2020, names the mechanism without euphemism: the study set out to characterise the "socio-environmental factors that impact model performance, nursing workflows, and patient experience," and found "tensions between the model's thresholds for data quality, and the quality of data that arise from an imperfect, resource-constrained environment." A threshold is a specification. An imperfect, resource-constrained environment is every place work actually happens.

The quality threshold is a decision. Engineers optimising for the metric they were graded on converted an accuracy risk into a throughput cost, and pushed that cost onto a nurse and a patient in a room they had never stood in. It is a perfectly defensible choice at Gate 2 and a catastrophic one at Gate 3. Nothing about it shows up in a validation report.

A rejection threshold transfers cost from the model's scoreboard to somebody else's afternoon.
On the Thailand deployment

The Epic Sepsis Model, 2018–2024. Of every case in this gallery, this is the one that most completely defeats the "pilots fail to reach production" framing. It reached production. It reached production at hundreds of American hospitals, embedded in the dominant electronic health record, firing alerts into live clinical workflow for years. By the pilot-to-production metric it is a triumph.

In 2021, researchers at the University of Michigan externally validated it across 38,455 hospitalisations and reported an area under the curve of 0.63 — against the 0.76 to 0.83 range cited in the vendor's own documentation — with sensitivity of 33% and positive predictive valueOf all the times the system raises an alarm, the share that turn out to be real. It is the number the person receiving the alert actually experiences — a low value means most alarms are false, which is how a technically defensible model trains the staff around it to ignore it. of 12%. Two-thirds of sepsis cases missed; roughly eight false alarms for every true one.

The 2024 replication is worse, and more useful. Ostermayer and colleagues examined 145,885 emergency-department encounters across two county hospitals through 2023, alerting at Epic's recommended threshold of 6. Sensitivity: 14.7%. Positive predictive value: 7.6%. And the number that ends the argument — the median lead time on the alert was zero minutes (80% CI, −6h42m to +12h00m). The model, at the median, told clinicians a patient was becoming septic at the moment the patient was already recorded as septic.

Sensitivity, AUC and lead time are not measuring the same thing, and the distinction runs through the rest of this brief. Sensitivity and AUC measure discrimination — how well a model tells the sick apart from the well, judged on a pile of finished cases after the fact. Lead time measures decision value — whether the answer arrives while there is still something a person can do about it. A model can be respectable on the first and worthless on the second.

Figure 5The warning that arrives at the event
0 min
Median lead time of the sepsis alert
Across 145,885 emergency-department encounters at two county hospitals in 2023, at the vendor-recommended alert threshold. 80% CI: −6h42m to +12h00m.
Evidence: Ostermayer et al., "External validation of the Epic sepsis predictive model in 2 county emergency departments," JAMIA Open / PMC11560849, cohort period 1 Jan–31 Dec 2023.
Method: Sepsis defined by Sepsis-3 (≥2-point SOFA rise with blood culture and antibiotic ordered); 6-hour window; alert threshold 6 as recommended by the vendor. Median, not mean — the confidence interval is wide in both directions and the point estimate should not be read as "always zero."
Synthesis: A prediction that arrives at the moment of the event has no decision value, whatever its discrimination statistics say. This is a Gate-3 failure that Gate-2 metrics conceal.

Start with the objection, because it is a good one: Epic fixed it. On 27 February 2026 a multicentre prospective validation of the updated model ran across four US health systems and 227,091 inpatient encounters and reported encounter-level AUROC of 0.82 to 0.92 — inside the range the first version was only ever claimed to hit — with median lead time before sepsis onset of 1.9 to 10.3 hours depending on the institution. That is a real improvement and it should be said plainly. Three details in the same study qualify it. The positive predictive value at a 60% sensitivity threshold still ran only 0.13 to 0.26, so between 21 and 35 alerts were needed per detected case; the spread across four sites was wide enough that the authors recommend "local validation prior to deployment"; and the study was, in their words, "not designed to evaluate its impact on clinical outcomes, such as mortality." A better discrimination statistic is not yet a better patient outcome.

What the correction does not do is rescue the deployment. Version one ran live in hundreds of hospitals for roughly six years. What ended it was not the vendor's monitoring and not a customer's: it was two academic external validations, published five years apart, by people with no commercial relationship to the alert. Every hospital paying for the model had the data to compute what Michigan computed. None did. That is the finding — the fix arrived through the literature because no one in the operating chain owned the measurement.

The second point is the obvious one. A model can be technically live and clinically inert, and the difference is invisible to everyone except a researcher who goes looking. Alert volume goes up, which reads as engagement. Nobody's dashboard says "median lead time: zero."

The third is less comfortable. Engineering did not fail here. Sensitivity of 14.7% with specificityThe mirror image of sensitivity: the share of the people who were not going to get sick that the model correctly leaves alone. High specificity means few false alarms as a fraction of the healthy — which can still be a great many alarms in absolute terms, because the healthy vastly outnumber the sick. of 95.3% is a model doing exactly what a low-signal prediction problemA problem where the data available simply does not contain much advance warning of the thing you want to predict. No modelling technique conjures signal that is not there, so the ceiling on performance is set by the world, not by the engineering — and the honest response is to ask whether a prediction is the right instrument at all. allows, deployed at a threshold chosen to keep alert volume tolerable. The organisation traded away almost all of the model's sensitivity to buy alert-fatigue relief, and then kept the alert on. That trade was made by clinical informatics staff balancing two operational pressures. It was rational locally and value-destroying globally, and no AI capability at any price would have changed it.

The lesson"In production" is a deployment fact, not a value fact — and a model tuned for tolerable alert volume can be perfectly live and clinically worthless at the same time.

7 · The economics nobody priced — cost, latency, and the human in the loop

Amazon's Just Walk Out was the most technically impressive retail deployment of its decade: walk in, take things, walk out, get charged. In 2024 Amazon began removing it from Amazon Fresh grocery stores in the US, replacing it with Dash Carts — a smart trolley that asks the shopper to scan. Reporting at the time described roughly a thousand associates in India reviewing shopping sessions. Amazon disputes the characterisation: spokesperson Jessica Martin Strauss said "the characterization that Just Walk Out technology relies on human reviewers is inaccurate," that associates' primary role is annotating video to improve the model, and that they "may also validate a small minority of shopping visits where our computer vision technology cannot determine with complete confidence an individual's purchases."

Take Amazon's account at face value; the conclusion barely changes. Grocery baskets are large, long-dwell and visually cluttered — the hardest possible case for the technology and the lowest-value one, because a supermarket checkout is already cheap per basket. Just Walk Out was not a failure. It was pointed at the wrong basket size, and it remains deployed in stadiums, airports and convenience formats where the basket is three items and the alternative is a queue nobody has time for. That is a Gate-1 error corrected late, not a Gate-2 defect.

Where the human-in-the-loopA person kept inside the automated process to check, correct or complete what the machine produces. Sold as scaffolding that comes down once the model improves; in practice it is usually a permanent staffed tier, and its size is the honest measure of how much automation was actually achieved. economics were genuinely concealed, the SEC put it on the record. On 14 January 2025 the Commission instituted settled cease-and-desist proceedings against Presto Automation, a Nasdaq-listed restaurant technology company, over its Presto Voice drive-thru product. Two findings matter here.

First, from November 2021 to September 2022, every commercially deployed Presto Voice unit ran on a third party's speech technology while Presto described it in Commission filings as "our" and "Presto's" technology; the company had not begun building its own until early 2022. Second, and more consequential, when Presto did deploy its own AI it claimed the product "eliminat[es] human order taking." The order finds that the proprietary units "lacked the capability to take orders on their own and required substantial human involvement" — Presto "hired, trained, and supervised human order takers located abroad (primarily in the Philippines and India), who processed the vast majority of drive-thru orders." The original version was designed to depend on that support. The later version "required a human agent to enter the orders approximately 70% of the time."

That 70% is the figure everyone quotes, and quoting it alone is too kind. The order records what Presto itself had to concede on 14 December 2023: the 70% "referred to orders at the few locations where the most advanced version of Presto Voice was being piloted," and "human agent intervention was required on 100% of orders at the substantial majority of locations where the original version of Presto Voice units were installed." The same disclosure put the average non-intervention rateThe share of transactions the system completed with no human touching them — the mirror of the intervention rate. Vendors quote the flattering one, and neither number means anything until you know which transactions, at which sites, over which period, are in the bottom of the fraction. across all restaurants running Presto's own technology at 85% — so a 15% intervention rate — while the company had told investors it achieved "automated order completion" rates of 95% to 99%. Four numbers, four denominators, one product. Read Presto's Form 10-K for the fiscal year ended 30 June 2023 and the mechanism is visible in the qualifier: the filing claims "approximately 95% non-intervention rate at certain locations." Three words are carrying the entire representation.

Presto's dishonesty is the least portable part of this. An automation rate is meaningless without its denominator, and a vendor asked for "the intervention rate" will answer with the best-scoped one available. The buyer's job is to specify the population before the number is quoted: all units, all locations, last calendar month.

Figure 6One product, four automation rates — what "eliminates human order taking" meant, by denominator
Original version — substantial majority of locations100%Most advanced version — few pilot sites70%Average across all Presto-technology restaurants15%Presto's public claim — “at certain locations”5%Share of drive-thru orders requiring an off-site human agent. Red = the population the buyer actually had.
Evidence: SEC Order Instituting Cease-and-Desist Proceedings, In the Matter of Presto Automation Inc., Rel. 33-11352, 14 Jan 2025, ¶¶36–38; Presto Automation Form 10-K for the fiscal year ended 30 June 2023 (filed 11 Oct 2023).
Method: All four are Presto’s or the Commission’s own figures, on four different populations, as of the dates in the Order. 100% and 70% are the Commission’s findings for the original and most-advanced versions respectively; 15% is the complement of the 85% average non-intervention rate Presto disclosed on 14 Dec 2023 across all restaurants running its proprietary technology; 5% is the complement of the “approximately 95% non-intervention rate at certain locations” claimed in the FY2023 Form 10-K. These are not four measurements of one quantity and are not averaged.
Synthesis: The gap between the top bar and the bottom one is not a measurement error; it is a choice of denominator, and it was the whole business case. An automation rate quoted without its population is a marketing number — which is the same defect Texas made enforceable against Pieces, one section on.

The Commission ordered Presto to cease and desistA regulator's order to stop the conduct and not repeat it. Where the proceedings are settled, the company accepts the order without contesting the findings — neither an admission nor a trial — and, as here, it can arrive with no fine attached at all. from violations of Securities Act §17(a)(2) and Exchange Act §13(a) and Rules 13a-11 and 13a-15(a). No civil penalty was imposed. Presto had been delisted from Nasdaq the previous September.

A third pattern kills pilots on cost alone, and Gartner put a number on it. Its July 2024 analysis — the source of the widely repeated forecast that at least 30% of GenAI projects would be abandoned after proof of concept by the end of 2025, "due to poor data quality, inadequate risk controls, escalating costs or unclear business value" — carried a cost range for business-model-innovation deployments of $5 million to $20 million. Rita Sallam's framing was blunt: "executives are impatient to see returns on GenAI investments, yet organizations are struggling to prove and realize value. As the scope of initiatives widen, the financial burden of developing and deploying GenAI models is increasingly felt."

Falling model prices do not fix this, and may make it worse. When inferenceThe act of running a trained model — one question in, one answer out, billed per use, as distinct from the one-off cost of training it. It is the line everyone watches, because it is the line that has been collapsing in price. gets cheaper, the binding constraint moves to the things that did not get cheaper — integration, evaluation, monitoring, escalation handling, the human review tier, the legal review of what the system is allowed to say. Those are labour and organisational costs, and they scale with the number of workflows touched rather than with tokens. A pilot whose economics only work if the model becomes free is a pilot that has misidentified its own cost structure.

The lessonPrice the human-in-the-loop tier at pilot design as a permanent operating line, not a temporary scaffold — because in every case where it was treated as temporary, it wasn't.

8 · Scaled before it was right — Zillow Offers, 2018–2022

Zillow Offers is the most expensive case in this brief and the one most often mis-told. It is usually filed as "the algorithm was wrong." Zillow's own filings say something more precise and more useful.

On 2 November 2021 the company announced it would wind down the business. Rich Barton: "We've determined the unpredictability in forecasting home prices far exceeds what we anticipated and continuing to scale Zillow Offers would result in too much earnings and balance-sheet volatility." That quarter carried "a write-down of inventory of approximately $304 million within the Homes segment as a result of purchasing homes in Q3 at higher prices than the company's current estimates of future selling prices," with a further $240–265 million of losses expected in Q4. The Homes segment lost $421.6 million before tax in the quarter; consolidated net loss was $328.2 million. The wind-down would cut roughly 25% of Zillow's workforce.

The FY2021 Form 10-K totals it: inventory write-downs of $407.9 million for the year, $71.2 million of impairment and restructuring inside the Homes segment, $6.2 million winding down the financing facilities, $4.9 million of accelerated depreciation. The board's stated reasoning is a value judgment, not a technical one: the decision was made "in light of home pricing unpredictability, capacity constraints and other operational challenges… all of which led us to conclude that, despite its initial promise in earlier quarters, Zillow Offers was unlikely to be a sufficiently stable line of business to meet our goals and needs going forward."

And the risk factor, written by Zillow's own lawyers, is the cleanest statement of a Gate-2 limit anywhere in this brief: "We underwrite and price the homes we buy and sell through Zillow Offers using in-person evaluations and data science and proprietary algorithms… These assessments may be inaccurate… Our pricing model may not account for submarket nuances — for example, the location of a home on a hill or in a building — which could have a significant impact on price."

Figure 7What a pricing model's error bars cost when they exceed the margin
$407.9mFY21 inventory write-downs
$71.2mHomes impairment & restructuring
$421.6mQ3'21 Homes pre-tax loss
~25%Of Zillow's workforce cut
Evidence: Zillow Group Form 10-K for FY2021 (filed Feb 2022), Item 1A risk factors and MD&A; Zillow Group Q3 2021 results release, Exhibit 99.1 to Form 8-K, 2 Nov 2021.
Method: Figures as disclosed; the $407.9m annual write-down is not additive with the $304m Q3 figure (the Q3 amount is a component of the year). The $71.2m is Homes-segment impairment and restructuring only, and excludes the $6.2m financing wind-down charge and $4.9m accelerated depreciation disclosed separately.
Synthesis: The loss is not the model's error; it is the model's error multiplied by an inventory position the company took on the strength of the model. The scaling decision, not the estimator, is what converted a forecasting limitation into half a billion dollars.

Two clarifications, because this case is easy to over-claim.

Zillow Offers was not a pilot by 2021; it was an operating business with a balance sheet. It earns its place here because of what the pilot phase did not establish before scaling — the width of the prediction intervalThe range a model's forecast is actually likely to fall in, rather than the single number it prints. A model that says "$400,000" is really saying "somewhere around $400,000, probably within this much" — and that this much, sometimes quoted as the standard error, is the number that decides whether you can trade on the forecast. A narrow interval is a business; a wide one is a bet. relative to the gross margin on a flip. An iBuyerA firm that buys homes directly from sellers for cash, on an algorithm's price estimate, then resells them. The seller trades price for speed and certainty; the iBuyer takes the house — and the risk of having valued it wrong — onto its own balance sheet.'s economics are a spread trade: buy at estimate minus fee, sell at market, keep the difference less holding and renovation cost. If the model's standard error on a home is a few percent and the spread is a few percent, the business is a coin flip with leverage, and no amount of model improvement inside the achievable range changes that. That is knowable at pilot scale, from the residual distributionThe spread of a model's misses — every past estimate set against what the house actually sold for, collected into a shape. It shows not just how wrong the model is on average but how wrong it can get on a bad one, which is the only version of the question a balance sheet cares about., before a single house is bought at volume.

The market moved violently in 2021 and Zillow was caught longHolding a large amount of an asset when its price falls. Being long is not itself a mistake — you cannot flip houses without owning houses — but the size of the holding decides how much a wrong price costs you.. It is fair to say bad luck contributed. It is not fair to say bad luck explains it, because the company's own explanation is that the variance exceeded expectation — which is a statement about the model's calibration, not about the draw. Calibration is a different property from accuracy, and here the distinction is the whole case. An accurate model gets the number close; a calibrated model knows how close, and can say when it is guessing. A model can be accurate on average and badly calibrated at the same time, which is the dangerous combination, because it hands over its shakiest estimates in exactly the same confident voice as its best ones. A well-calibrated pricing model that knows its own uncertainty produces a lower offer, buys fewer homes, and loses less. Opendoor ran a comparable book through the same market and did not exit. The differentiating variable was position sizingDeciding how much to stake on a view, given how sure you are of it. It is separate from being right: the same forecast, held in twice the size, loses twice the money. In an algorithmic business it is the executive decision that converts a modelling limitation into a financial one, or doesn't. against model uncertainty, which is a risk-management choice made by executives.

The question a pilot must answer is not "how accurate is the model" but "how does the model's error distribution compare to the margin it is being asked to protect." An error rate that is excellent for a recommendation engine is catastrophic for a balance sheet. Most pilot readouts report the first number and never compute the second.

9 · The governance kill — and why one of these is a success

One case in this section is filed everywhere as a famous AI failure and is, on the evidence, the opposite.

Amazon began building an automated résumé-ranking system in 2014, scoring candidates one to five stars. Trained on a decade of the company's own résumés in a male-dominated field, it learned to prefer men — penalising résumés containing the word "women's" and the names of certain all-women colleges. Amazon neutralised those specific terms and then killed the project anyway, because it "lost confidence that the program was indeed gender neutral in all other areas." Amazon's own position, in the reporting that broke the story, is that the tool "was never used by Amazon recruiters to evaluate candidates" — though the company did not deny that recruiters looked at its recommendations, which is a narrower denial than it first appears and is the reason this case is filed as a near-miss rather than a clean abstention.

An organisation detected a defect it could have papered over, tried the paper-over, judged the residual risk unquantifiable, and shut the program down before deployment. That is exactly what a functioning control environment does. It cost Amazon four years of engineering and saved it what the alternative cost others. The reason it is remembered as a failure is that our vocabulary has no word for a pilot that correctly kills itself — and that vocabulary gap is itself a cause of failure elsewhere, because a leader who reads Amazon's outcome as humiliation learns to keep quiet instead.

Now the ones that did deploy. In September 2023 the EEOCThe US Equal Employment Opportunity Commission — the federal agency that enforces the laws against discrimination in hiring, pay and promotion. It can investigate, sue on a worker's behalf, and settle on terms that bind an employer for years. settled with iTutorGroup for $365,000: the company had "programmed their tutor application software to automatically reject female applicants aged 55 or older and male applicants aged 60 or older," rejecting more than 200 qualified US applicants. No learned model produced that bias; it was a hard-coded rule. The AI framing is almost incidental; what the case establishes is that automated screening logic is legally identical to a hiring manager's decision, with the aggravating feature of being written down.

Which is the bridge to Mobley v. Workday. Derek Mobley applied to more than 100 employers using Workday's platform and was rejected every time. On 16 May 2025, Judge Rita Lin of the Northern District of California granted preliminary certification of a nationwide ADEA collectiveA group action under the Age Discrimination in Employment Act, the US law protecting workers aged 40 and over. A "collective" is the age-discrimination cousin of a class action: people join it rather than being swept in automatically. Preliminary certification decides only that the group may proceed together — it settles nothing about whether the claim is right. covering applicants aged 40 and over denied recommendations through Workday's platform since 24 September 2020. The court had earlier declined to dismiss on the theory that Workday was merely a software supplier. The doctrine at stake is agency: a vendor performing a core hiring function on an employer's behalf can be treated as the employer's agent and held directly liable.

Two years on, the case has widened. The preliminary collective was scoped in July 2025 to applicants screened using Workday's HiredScore features, over the company's objection that HiredScore was a separate product. In January 2026 the court allowed three further named plaintiffs and new claims under Title VIIThe section of the US Civil Rights Act of 1964 that bans employment discrimination on race, colour, religion, sex and national origin. It is the main federal hook for hiring-discrimination claims outside age. and California's Fair Employment and Housing Act. In the order of 6 March 2026 Judge Lin dismissed the FEHA counts with leave to amend for want of a pleaded California nexus — and denied dismissal of everything else, including the ADEA disparate-impactA discrimination claim that works without any proof of intent: it is enough that a neutral-looking rule or system falls harder on a protected group and cannot be justified by the job. This is the doctrine that makes a screening algorithm legally interesting: intent is not the question, outcome is. claim, while granting AARP leave to file an amicus briefLiterally "friend of the court" — an argument filed by an outside organisation with no direct stake in the case, offering the judge a wider view of what the ruling would mean. Courts grant leave for one when they think the decision reaches beyond the two parties.. On 22 June 2026, ruling on the re-pleaded Third Amended Complaint, the court granted in part and denied in part, rejecting Workday's argument that California's anti-discrimination law cannot reach applicants screened from California for jobs elsewhere. The docket runs on. No merits determination has been made against Workday and none is assumed here.

What has been established in three years of procedure is narrower than the headlines and more useful. A screening vendor has not persuaded a court that it is merely a supplier; a re-pleading has not made the age-discrimination theory go away; and the location of the algorithm, not the location of the applicant, is now doing work in the jurisdictional analysis. For a buyer: where the model runs is becoming a compliance fact about your hiring process.

"We bought it, so the vendor carries the risk" is not how this is resolving, and almost no pilot budget models what follows. Both parties are exposed, and the employer's exposure is not reduced by the vendor's indemnity when the claim is statutory discrimination brought by an applicant. A screening pilot's true cost includes a bias-testing regime, a records regime, and a legal review — before go-live, not after the first demand letter.

Then the FTC. In December 2023 the Commission banned Rite Aid from using facial recognition for surveillance for five years, finding that from 2012 to 2020 the retailer deployed the technology across hundreds of stores without reasonable procedures, that the system falsely flagged consumers — disproportionately people of colour — as shoplifters, and that employees publicly accused them of wrongdoing, "sometimes in front of friends or family." The order requires Rite Aid to discontinue any such automated system "if it cannot control potential risks to consumers." Eight years of deployment, terminated by an enforcement action, with the risk-control obligation imposed retroactively at the worst possible moment.

And the state AGs are moving faster than the federal agencies on accuracy claims specifically. In September 2024 Texas settled with Pieces Technologies, whose generative summarisation product was in use at four major Texas hospitals receiving patient data in real time. Pieces had advertised a "severe hallucination rate" of "<1 per 100,000." The investigation found the metrics "were likely inaccurate and may have deceived hospitals about the accuracy and safety of the company's products." No money changed hands. What changed was the obligation: for five years, Pieces must clearly disclose the meaning and calculation method of any accuracy metric it advertises, or have an independent auditor substantiate the claim.

That remedy attacks the exact mechanism by which pilots are sold. A hallucination rate without a stated denominator, evaluation set and severity definition is not a measurement; it is a marketing number. Texas has now made publishing one, in that state, in healthcare, an enforceable representation.

Figure 8The regulatory arc — from "AI is unregulated" to "the claim is the liability"
Sep 2023EEOC v. iTutorGroup — $365,000 consent decree. Screening software hard-coded to reject women 55+ and men 60+; 200+ applicants affected. Automated screening logic held to the same standard as a human decision.
Dec 2023FTC v. Rite Aid — five-year ban on facial recognition for surveillance after eight years of deployment across hundreds of stores; obligation to discontinue any automated system whose consumer risk cannot be controlled.
Feb 2024Moffatt v. Air Canada (2024 BCCRT 149) — the chatbot's answer binds the company. "It should be obvious to Air Canada that it is responsible for all the information on its website."
Mar 2024SEC's first AI-washing actions — Delphia ($225,000) and Global Predictions ($175,000) penalised for claiming AI capabilities they did not have.
Sep 2024Texas AG / Pieces Technologies — first state settlement over healthcare generative-AI accuracy claims; five years of mandatory metric-methodology disclosure or third-party audit.
Jan 2025SEC v. Presto Automation — cease-and-desist. "Eliminates human order taking" met a product needing a human agent ~70% of the time.
May 2025Mobley v. Workday — nationwide ADEA collective preliminarily certified; the AI vendor is potentially liable as the employer's agent.
Mar–Jun 2026Mobley, two further orders — ADEA disparate-impact claim survives dismissal (6 Mar); on the Third Amended Complaint the court rejects the argument that California law cannot reach applicants screened from California for jobs elsewhere (22 Jun). Procedural stage only.
Jul 2026EU AI Omnibus in force — Annex III high-risk duties deferred to 2 Dec 2027, Annex I to 2 Aug 2028; prohibitions and transparency duties unchanged and live.
Evidence: EEOC press release 11 Sep 2023; FTC press release 19 Dec 2023; CRT decision 2024 BCCRT 149; SEC press release 2024-36 (18 Mar 2024); Texas OAG release (Sep 2024); SEC Rel. 33-11352 (14 Jan 2025); Holland & Knight analysis of the 16 May 2025 order in Mobley; N.D. Cal. order of 6 Mar 2026 (Dkt 267) and Duane Morris analysis of the 22 Jun 2026 order; European Commission, "AI Omnibus enters into force" (27 Jul 2026).
Method: Dates are the action date, not the underlying conduct date; conduct in Rite Aid ran 2012–2020 and in Air Canada from Nov 2022. Mobley remains at a procedural stage — no merits determination has been made against Workday, and the 2026 orders resolve pleading questions only. The July 2026 entry is a legislative act, not an enforcement action, and is included because it fixes the dates against which the other entries must be planned.
Synthesis: The arc runs from punishing outcomes to punishing claims. Pieces and Presto were both sanctioned for what they said their systems did, not for harm proved downstream — which prices the vendor demo directly.

10 · The vendor-demo mirage — when the product is the pitch

Builder.ai raised on the promise that AI would assemble software the way a pizza is assembled from toppings. It collapsed into insolvency in May 2025. The Register's account is unsentimental: more than $500 million raised from blue-chip investors including Microsoft and Qatar's sovereign fund, a business model in which "the Builder.ai team actually built the apps," and a prior history — as Engineer.ai — of press scrutiny over exactly that gap. The company had installed a new CEO in February 2025, months before the end.

Put Builder.ai next to Presto and the SEC's first AI-washing enforcement actions of March 2024 — Delphia ($225,000) for claiming it put "collective data to work to make our artificial intelligence smarter so it can predict which companies and trends are about to make it big," and Global Predictions ($175,000) for calling itself the "first regulated AI financial advisor" — and a pattern with a shape emerges. In each case the "AI" was, in whole or in part, either absent or a person. Gary Gensler's framing: "when new technologies come along, they can create buzz from investors as well as false claims by those purporting to use those new technologies."

The Delphia order documents something more damning than a marketing exaggeration. The Commission found that from at least August 2019 to August 2023 the firm claimed to use AI and machine learning to analyse retail clients' spending and social-media data "when, in fact, no such data was being used in its investment process" — and that after an examination, Delphia agreed in 2021 to correct the statements, made some corrective efforts, and then continued making false statements through August 2023. Two years elapsed between the regulator identifying the gap and the conduct stopping. A firm overselling AI is ordinary; an incentive strong enough to survive a direct regulatory intervention is the story.

The FTC added the consumer-facing version in September 2024. Under "Operation AI Comply" it brought five actions, including one against DoNotPay — "the world's first robot lawyer," which promised consumers could "sue for assault without a lawyer" and "generate perfectly valid legal documents in no time." The complaint alleges the company "did not conduct testing to determine whether its AI chatbot's output was equal to the level of a human lawyer, and that the company itself did not hire or retain any attorneys." The proposed order carried $193,000 and a requirement to notify subscribers from 2021–2023 about the service's limitations. The absence of testing is the finding worth carrying into a procurement conversation: the claim was not exaggerated relative to a measurement, it was made in the absence of one.

One word is doing most of the work in the newest version of this pitch, and a great deal now rides on it. An agent, in the enterprise sense, is a system handed a goal rather than a question: it plans its own steps, reaches into other software, and acts — issues the refund, files the ticket, updates the record — instead of handing text back to a person who then does those things. That is a genuinely different and harder thing to build than a chatbot. The label, unfortunately, is much easier to apply than the capability is to deliver.

Gartner gave the phenomenon its enterprise name and, in June 2025, a magnitude. Its analysis warned that over 40% of agentic AI projects will be cancelled by the end of 2027, "due to escalating costs, unclear business value or inadequate risk controls," and described "agent washing" — "the rebranding of existing products, such as AI assistants, robotic process automation (RPA) and chatbots, without substantial agentic capabilities." Gartner's estimate: only about 130 of the thousands of agentic AI vendors are real. Anushree Verma: "Most agentic AI propositions lack significant value or return on investment… Many use cases positioned as agentic today don't require agentic implementations."

A vendor's compensation event is a signed proof of concept, not a delivered outcome twelve months later — which is most of the explanation, and dishonesty is very little of it. A demo is optimised against the buyer's imagination; a production system is optimised against the buyer's edge cases. The two artefacts are built by different people to different specs, and only one of them is what the buyer evaluated. Every structural feature of enterprise procurement — the pilot budget that is smaller than the approval threshold, the six-week evaluation window, the reference customer selected by the vendor — pushes in the same direction.

The practical defence is cheap and almost never used: require the vendor to disclose the human-intervention rate, the escalation rate, and the evaluation set composition as contractual representations, and make the pilot's success criterion a metric computed on the buyer's data by the buyer. Presto's 70% would have surfaced in week two.

The honest vendors describe this problem in their own filings. C3.ai — which structures customer engagements as paid "Initial Production Deployments" before any subscription — warns in its Form 10-K for the fiscal year ended 30 April 2026 that "after completing an initial production deployment or trial customers do not have an obligation to continue to license our products, and we may not be able to convert initial production deployment customers into customers purchasing ongoing subscriptions." The pilot-to-production chasm is not only a buyer's complaint; it is a disclosed revenue risk on the seller's side, which tells you the conversion rate is low enough to be material.

Three questions that would have caught every vendor case in this section, asked before signature:
  • What percentage of transactions in your reference deployment required human intervention last month, and how is "intervention" defined? (Presto: ~70%, undisclosed until an SEC investigation.)
  • Show the evaluation set your accuracy figure was computed on, its size, its provenance, and the severity definition behind any error-rate claim. (Pieces: "<1 per 100,000" with no stated method; now a five-year disclosure obligation in Texas.)
  • Which components of the system do you own, and which are licensed from a third party under a contract you could lose? (Presto: all deployed units ran on a supplier's technology while filings called it "our" technology.)

11 · Success theatre — the pilot kept alive because killing it costs more

In October 2023 New York City launched an AI chatbot to help small businesses find their way through municipal regulation. In March 2024 The Markup tested it and found it advising landlords that they need not accept tenants using housing vouchers — illegal source-of-income discrimination in New York City — and telling employers they could take workers' tips.

It stayed up. For nearly two more years.

The New York City Comptroller's audit of the MyCity system found that the Office of Technology and Innovation had "very little else to show after spending more than four years and $100 million on the program." Users still could not apply for benefits through a single form. The portal "largely redirects users to pre-existing City websites that have been rebranded and redesigned as MyCity applications." On the chatbot specifically: it "appears to be unable to provide accurate or consistent information." The audit's overall conclusion — "MyCity was poorly managed from both a project management and contract oversight perspective" — was accompanied by the detail that OTI had requested a further $81 million in the 2026 budget to maintain what existed and add functionality.

The chatbot was finally shut down in early 2026, and the reason it was shut down is the point. A new administration took office; Mayor Zohran Mamdani called the bot "functionally unusable", said it was costing "around half a million dollars," and cut it as part of closing a budget gap. Reporting put the original build cost near $600,000. The bot's page now says its beta test has ended.

Nearly two years elapsed between public documentation that the system gave illegal advice and its termination. In that window the failure was known, cheap to fix by removal, and structurally protected — because the thing being protected was not the chatbot. It was the announcement. A launched initiative is a political asset whose value is realised at launch and whose costs accrue afterwards to whoever inherits it. Withdrawal converts a past success into a present admission. The rational move for the sponsor is to leave it running and let the budget carry it, which is precisely what happened until an actor with no stake in the original announcement arrived.

Success theatreA pilot maintained past the point of demonstrated failure because its termination would impose a reputational cost on its sponsor that exceeds the operating cost of leaving it running. Its diagnostic signature: the program has a communications plan and no kill criterion.

The corporate version runs on a shorter clock but the same logic. On 9 July 2024 the HR software firm Lattice announced it would become the first company to give "digital workers" official employee records — AI agents onboarded, given goals, performance metrics, system access and an accountable manager. The backlash from HR practitioners was immediate. On 12 July, three days later, CEO Sarah Franklin withdrew the feature, telling Fortune: "This innovation sparked a lot of conversation and questions that have no clear answers yet… we will not further pursue digital workers in the product."

Lattice belongs here as the control case for New York. Same failure category — a launch that met reality badly — and a total elapsed time of 72 hours, because a product company that sells to HR professionals gets its feedback from the people whose approval it needs to survive, immediately and unambiguously. New York's feedback loop ran through a press investigation, a comptroller's audit, and an election. The difference in cure time is a difference in who is allowed to say no and how fast they can say it.

Sometimes the theatre is the deliverable itself. In 2025 Australia's Department of Employment and Workplace Relations received a report from Deloitte on welfare compliance under a contract valued at about A$440,000. Academics found fabricated references — non-existent academic papers and a made-up quotation from a Federal Court judgment. The revised version disclosed that a generative AI system, Azure OpenAI, had been used in its production. Deloitte refunded A$97,000 — less than a quarter of the contract value.

The refund is the smallest number in the story. The reputational and policy cost of a government welfare-compliance report containing invented case law is not A$97,000. More instructive for anyone running an internal program: the errors were caught by an outside academic reader, not by the firm's own review. The deliverable passed every internal quality gate the engagement had. That is the exact shape of the risk when generative tools are introduced into knowledge work without a corresponding change to verification: the failure mode is not visibly bad work, it is confidently formatted work that nobody re-checks because it looks like the work that used to be checked.

12 · Where the value actually leaks — the measured evidence

Everything so far has been reconstructed from cases. What follows is the small body of properly measured evidence on what happens to value between the tool and the P&L.

Finding one: freed time is reallocated, not banked. Anders Humlum and Emilie Vestergaard linked large-scale adoption surveys to Danish administrative payroll records — 25,000 workers across 7,000 workplaces in the latest round. In workplaces combining encouraged use, enterprise tools and training, 93% of workers reported using chatbots at work, 28% daily, and 19% reported saving more than an hour a day. Those are excellent adoption numbers by any enterprise standard. The earnings and hours result: "precise null effects… at both the worker and workplace levels, ruling out effects larger than 2% two years after the launch of ChatGPT."

The outcome in that study is not what anybody said happened; it is what the payroll system recorded — the survey supplies the adoption, the administrative data supplies the result. And a precise null is a positive finding rather than an empty one: the study did not fail to detect an effect, it measured the effect, found it sitting at about zero, and had the statistical room to rule out anything much larger. "We found nothing" and "we established there is nothing above this size" are very different sentences.

The paper states the mechanism: "most chatbot users (85%) report reallocating time savings from AI chatbots to other job tasks." New tasks appeared too — content generation, oversight of AI outputs, integration work — reported by about 8% of users without employer initiatives and roughly 17% where initiatives existed.

Figure 9aWhere the value leaks — adoption succeeds, time is saved, and nothing reaches the accounts
Denmark: everything works until the last barUsed a chatbot at work93%Use one daily28%Report saving more than an hour a day19%Reallocate the saved time to other tasks85%Effect on earnings and recorded hours2% (upper bound)Denominators differ by row — see Method. Red = the bar the business case was written against.
Evidence: Humlum and Vestergaard, "Still Waters, Rapid Currents: Early Labor Market Transformation under Generative AI," NBER Working Paper 33777, May 2025, revised March 2026; latest survey round of 25,000 workers across 7,000 workplaces linked to Danish administrative payroll records.
Method: The bars are not one funnel and do not multiply. The first three are shares of workers in workplaces that combine encouraged use, enterprise chatbots and training — the best-supported settings in the sample, not the average. The fourth is a share of chatbot users. The fifth is not a share at all: it is the upper bound of the difference-in-differences confidence interval on earnings and recorded hours, drawn to the same scale so the magnitudes can be seen against one another, and its point estimate is zero.
Synthesis: Every gate an enterprise AI programme is managed against clears comfortably here, and the one nobody manages does not. Reading left to right, no mechanism existed to convert the fourth bar into the fifth.

The time was really saved. The workers really felt it. And it went into other work — some of it new work created by the tool itself. No cost line fell because no mechanism existed to make one fall. Gate 4 was never built.

Figure 9Net annual value of a 1,000-seat assistant ($000) — adoption × conversion
Weekly-active adoption of licensed seatsCapacity converted to cash20%40%60%80%0%-360-360-360-36015%-22+315+652+99030%+315+990+1,665+2,34050%+765+1,890+3,015+4,140
Evidence: Structure of the calculation only; the conversion axis is anchored on Humlum & Vestergaard's finding that 85% of users reallocate time savings to other tasks (NBER WP 33777, May 2025, rev. Mar 2026). Adoption range brackets the Danish study's 28% daily use in best-supported workplaces.
Method: Computed, not estimated. Gross theoretical pool = 1,000 seats × 150 hours saved per active user per year (3 hrs/week × 50 weeks) × $75 loaded hourly cost = $11.25m. Net value = pool × weekly-active adoption × share of freed capacity converted to cash, less $360,000 of annual licence cost ($360/seat). Cells in $ thousands. Illustrative parameters, chosen to be legible, not a forecast for any named firm; the shape of the surface, not the cell values, is the finding.
Synthesis: Adoption is the axis every program manages and conversion is the axis that decides the sign. At 0% conversion the program is negative at every adoption level, including 80%. Doubling adoption from 40% to 80% at 15% conversion adds $675k; raising conversion from 15% to 50% at 40% adoption adds $1,575k — more than twice as much movement for the axis nobody manages.

Finding two: self-reported productivity is wrong in a predictable direction. METR ran a randomised controlled trialThe design that carries the most weight in any evidence hierarchy: who gets the tool and who does not is decided by chance, so the two groups differ only in the tool. Everything else — skill, motivation, the difficulty of the work — averages out, which is what lets you say the tool caused the difference rather than merely accompanying it. with 16 experienced open-source developers across 246 real issues in repositories they knew well, averaging over 22,000 stars and a million lines of code, using Cursor Pro with Claude 3.5/3.7 Sonnet, frontier models at the time. Developers forecast a 24% speed-up. They were measured 19% slower. Asked afterwards, they estimated they had been sped up by 20%.

Two numbers, same people, same tasks: a 39-percentage-point gap between the experience and the measurement. METR is careful about generalisation, and so should anyone citing it be — 16 developers, mature codebases they already knew intimately, a setting where the AI's context disadvantage is largest. It does not show AI slows most developers. What it shows is narrower and far more damaging to standard practice: practitioner self-report is not a valid instrument for measuring AI productivity effects, even among skilled practitioners reporting on their own recent work. The overwhelming majority of enterprise AI business cases are built on exactly that instrument.

METR has since revised its own position, and the revision should be carried by anyone who quotes the 19%. On 24 February 2026 the group reported that a larger follow-on experiment begun in August 2025 "gives us an unreliable signal of the current productivity effect of AI tools" — because 30% to 50% of developers were declining to submit tasks they did not want to do without AI, which biases the estimated speed-up downward. Its own reading now: "it is likely that developers are more sped up from AI tools now — in early 2026 — compared to our estimates from early 2025," while cautioning that "our data is only very weak evidence for the size of this increase."

The sign of the 2025 effect is now the wrong thing to argue about, and this brief does not rest on it. What the trial established is the gap between what skilled people believed about their own recent work and what a clock recorded — and nothing in the 2026 update repairs that instrument. A firm that could not measure itself in 2025 cannot measure itself in 2026 either; it can only be wrong in a more flattering direction.

Finding three: the gains are real, and they are concentrated where nobody points the pilot. Brynjolfsson, Li and Raymond studied the staggered rolloutGiving the tool to different groups at different times rather than to everyone at once. The people who have not got it yet act as a comparison group for the people who have — a natural experiment you get for free from an ordinary deployment schedule, and the closest most firms will come to a controlled trial without running one. of a generative assistant across 5,179 customer-support agents: productivity up 14% on average in issues resolved per hour, "including a 34% improvement for novice and low-skilled workers but with minimal impact on experienced and highly skilled workers," plus improved customer sentiment and higher retention. The mechanism is knowledge transfer — the model disseminates the practices of the best agents to the newest.

Put finding three beside finding two. The value concentrates in the least experienced workers on the most routine tasks; the enthusiasm, the budget and the pilot sponsorship concentrate in senior people doing complex work, where the measured effect is smallest or negative. MIT NANDA observed the budget half of this directly: sales and marketing captured the largest share of GenAI allocation "because it's easier to attribute metrics," while "back-office automation often yields better ROI." Note honestly that NANDA's own text is internally inconsistent on the size of that share — its takeaway line says 50% and its body says "approximately 70 percent" — so the direction is what should be carried, not the number.

Finding four: badly-used AI has a measurable negative externalityA cost your activity imposes on somebody else, which never appears on your own ledger. Pollution is the classic case. Here it is quieter: the time you save by generating something is spent again, with interest, by the colleague who has to work out whether it is any good.. BetterUp Labs and Stanford's Social Media Lab surveyed 1,150 US full-time desk workers in September 2025 and found 40% had received "workslop" — AI-generated content that looks like work and lacks the substance to advance it — in the prior month, costing roughly two hours per incident to sort out, which they price at about $186 per employee per month, or around $9 million a year at 10,000 employees. The trust effects were worse than the time cost: recipients rated senders as less capable and less reliable.

Workslop is what individual-level productivity gains look like when they are not governed. Person A's time saving becomes Person B's time cost, plus a trust debit. Netted across the firm the tool can be productivity-negative while every individual user honestly reports a gain — and every one of those honest reports lands in the pilot's success survey.

Finding five: the same gap shows up in the accounts, and it is not a story about engineers. METR measured sixteen developers. In March 2026 the Federal Reserve Bank of Atlanta, with the Richmond Fed and Duke, ran the firm-level version — a survey of nearly 750 corporate financial executives — and found the same shape in CFO answers. Relative to firms not investing in AI, investing firms reported AI-related labour-productivity growth 2.4 percentage points higher in 2025 and expected 3.3 points higher in 2026. The productivity gains implied by those same firms' own reported inputs and outputs were 1.0 and 1.8 points. The authors name it: "a productivity paradox, in which perceived productivity gains are larger than measured productivity gains."

Be careful what this is evidence for, because it cuts against a lazy reading of this brief as well as for it. The gains here are positive, they are largest in high-skill services and finance, and the paper expects them to strengthen. Anyone arguing that enterprise AI does nothing has to explain that away, and this brief does not try to. What the paper also finds is that the effect on employment is close to nothing — firm-size-and-sector-weighted aggregate employment is expected to fall by less than 0.4% due to AI in 2026, with large firms shedding and small firms adding — while the composition shifts, routine clerical work falling by more than two percentage points of employment share over three years. Real productivity, real task reallocation, almost no headcount. That is the Danish result and the Census result arriving a third time, from a third method, in the mouths of the people who own the budget.

Figure 10aWhat CFOs report, and what their own numbers imply
Reported, 2025+2.4 ppImplied, 2025+1.0 ppReported, 2026 (expected)+3.3 ppImplied, 2026 (expected)+1.8 ppLabour-productivity growth of AI-investing firms relative to non-investing firms. Red = what executives report.
Evidence: Baslandze, Edwards, Graham, McClure, Sparks, Meyer, Waddell and Weitz, "Artificial Intelligence, Productivity, and the Workforce: Evidence from Corporate Executives," Federal Reserve Bank of Atlanta Working Paper 2026-4, March 2026, Table 2 Panel A (extensive margin of AI investment).
Method: Values as published, in percentage points of labour-productivity growth for AI-investing firms relative to non-investing firms; 2026 figures are expectations, 2025 figures are reported outturns. "Implied" is the paper's own construction from respondents' reported inputs and outputs, not an external estimate. Not comparable with METR's task-completion percentages, which are a different unit on a different population and are not plotted alongside.
Synthesis: Reported gains run roughly 2.4 times implied gains in 2025 and 1.8 times in 2026. The optimism is not confined to enthusiasts using a tool; it survives at the level of the person who signs the business case — which is precisely where a pilot's value estimate comes from.
Figure 10The measured evidence — what each study can and cannot support
StudyDesign & sampleHeadline resultWhat it establishesWhat it cannot support
Humlum & Vestergaard
NBER WP 33777, May 2025 (rev. Mar 2026)
Surveys linked to Danish administrative payroll; 25,000 workers, 7,000 workplaces; difference-in-differencesPrecise null on earnings and hours; rules out effects >2%. 85% of users reallocate time savings to other tasksThat firm-level financial impact does not follow automatically from adoption or from felt time savingsThat AI has no productivity effect; the paper documents substantial task reorganisation
METR
10 Jul 2025
RCT; 16 experienced OSS developers, 246 real issues in their own mature repos19% slower with AI; forecast +24%, self-estimated +20% afterThat practitioner self-report is an invalid measurement instrument for AI effectsGeneralisation to most developers or to unfamiliar codebases; sample is 16. METR's own 24 Feb 2026 update reports that its follow-on experiment gives "an unreliable signal" and that developers are likely more sped up in early 2026 — so the sign of the 2025 effect carries no weight here, only the perception gap
Brynjolfsson, Li & Raymond
NBER WP 31161
Staggered rollout; 5,179 customer-support agents; issues resolved per hour+14% average; +34% novice/low-skill; minimal for experiencedThat real gains exist and concentrate in routine work done by less-experienced staffThat the firm captured the gain financially — throughput is measured, P&L is not
Atlanta Fed / Richmond Fed / Duke
Working Paper 2026-4, Mar 2026
Survey of ~750 corporate financial executives; extensive- and intensive-margin regressionsReported labour-productivity growth +2.4pp (2025) and +3.3pp expected (2026) for AI investors, against +1.0pp and +1.8pp implied; aggregate employment expected to fall <0.4% in 2026That the perception-versus-measurement gap holds at the level of the budget holder, and that real gains coexist with almost no headcount effectA causal estimate; responses are executive self-report and expectations, and "implied" is a construction, not an audited outturn
BetterUp Labs / Stanford SML
Sep 2025
Survey; 1,150 US full-time desk workers40% received "workslop" in prior month; ~2 hrs per incident; ~$186/employee/monthThat ungoverned AI output transfers cost to colleagues and debits trustA firm-level net productivity figure; the cost estimate is derived from self-reported time
Evidence: As listed; each opened and read directly. Brynjolfsson et al. figures are from the NBER working-paper abstract (14%); the later Quarterly Journal of Economics version reports 15% — the working-paper figure is used here because it is the text verified.
Method: The five studies use different outcomes (earnings/hours; task completion time; issues resolved per hour; self-reported hours lost) on different populations. They are not comparable as effect sizes and are not presented as such; each row is read only for the inference in column four.
Synthesis: Taken together they describe a value chain that works at the task level and breaks at the firm level — real gains, concentrated in the least-glamorous population, reported unreliably, and reallocated rather than captured.

13 · The incentive map — everybody behaved rationally and the value still died

Almost nothing in this gallery required anyone to act in bad faith. Lay out what each participant around a pilot is actually rewarded for and most of the failures stop being surprising.

Figure 11Who is paid for what — and what that does to the value
Executive sponsorOptics
Rewarded forA launch, a board slide, being early
Not rewarded forA clean kill
ConsequenceSuccess theatre; late termination
Business unitProtection
Rewarded forHitting this year's plan
Threatened byA tool that reveals slack in its headcount
ConsequenceFreed capacity reabsorbed, never surfaced
Data / IT orgRisk
Rewarded forUptime, security, no incident
Not rewarded forBusiness outcomes it can't control
ConsequenceShips to spec; owns none of Gate 3 or 4
VendorBookings
Rewarded forA signed PoC this quarter
Paid beforeAny outcome is measurable
ConsequenceDemo-optimised scope; agent washing
End usersWorkload
Rewarded forGetting today's work done
Punished byAny tool that adds a step or a doubt
ConsequenceQuiet non-use; shadow tools instead
ProcurementUnit price
Rewarded forDiscount achieved, cycle time
Not measured onWhether the thing was used
ConsequenceSeats bought in bulk, licences idle
Risk & legalAsymmetry
Rewarded forNo adverse event on their watch
Bears none ofThe upside of shipping
ConsequenceLate, binary vetoes after sunk cost
Evidence: Behaviours inferred from the documented cases in Part I — NYC OTI (sponsor), CBA (business unit and sponsor), MD Anderson (procurement and IT), Presto and Builder.ai (vendor), MIT NANDA's shadow-AI and barrier findings (end users), Rite Aid and Mobley (risk and legal).
Method: Analytical construction, not survey data. Each "consequence" is traceable to at least one case in this brief; no frequency claim is made.
Synthesis: No participant is rewarded for the outcome the pilot exists to produce. Gate 4 sits in the gap between the business unit that must give up the capacity and the sponsor who is measured on the launch — which is to say, nowhere.

One under-quoted NANDA finding is the end-user row made visible. While only 40% of companies had bought an official LLMLarge language model — the kind of system behind a chat assistant, trained on enormous quantities of text to predict what words come next, and general enough to draft, summarise, translate or answer without being built for any one of those jobs. Its generality is what makes it easy to adopt informally and hard to govern. subscription, workers at over 90% of surveyed companies reported regular use of personal AI tools for work. Employees crossed the divide privately while the corporate program sat in pilot. The people are not resisting AI. They are resisting your AI, and they have a substitute.

The lessonIf the only person whose bonus moves when the pilot delivers value is the one who cannot personally release the headcount, the pilot has no owner at the gate that decides its fate.

14 · The taxonomy — nine failure modes, built from the cases up

Data quality, model accuracy, integration, skills: that is the list you get when a taxonomy of this subject is imposed from the top. The taxonomy that emerges from the cases above is a list of operating decisions, and hardly any of its modes would be fixed by a better model.

Figure 11aThe gallery at a glance — sixteen pilots, ambition to lesson
MD Anderson OEAHealthcare · 2013–2016
AmbitionA Watson advisor recommending oncology treatment
What happened$62m spent; built on the EHR the institution then left; never used clinically
LessonNobody owned the integration path the value ran through
Google Health ThailandHealthcare · 2018–2019
AmbitionScreen 4.5m diabetic patients past a 200-specialist bottleneck
What happened>90% lab accuracy; more than a fifth of real images refused on quality
LessonA quality threshold moves cost off the model's scoreboard onto a nurse
Epic Sepsis ModelHealthcare · 2018–2024
AmbitionWarn clinicians before sepsis onset, inside the dominant EHR
What happenedReached hundreds of hospitals; AUC 0.63 external, 14.7% sensitivity, zero-minute median lead time
LessonIn production is a deployment fact; only an outsider ever measured it
Air Canada chatbotAirline · 2022–2024
AmbitionDeflect routine fare questions from the contact centre
What happenedBot invented a bereavement-fare policy; tribunal held the airline to it
LessonYour disclaimer is an internal document; the counterparty gets the company's word
Commonwealth Bank voice-botBanking · 2025
AmbitionDeflect calls, then remove 45 contact-centre roles
What happenedCall volumes rose; overtime offered; redundancies reversed as an "error"
LessonDeflection rate and workload are different numbers and can move in opposite directions
Klarna AI assistantFintech · 2024–2026
AmbitionAutomate two-thirds of service chats, "work of 700 agents"
What happenedRoutine automation held; complex and emotive cases degraded; humans re-recruited
LessonThe automation worked; the inference from share-of-chats to headcount did not
McDonald's AOT with IBMQSR · 2021–2024
AmbitionAutomated order taking across 100+ drive-thrus
What happenedEnded Jul 2024; success metric never disclosed; relaunched 2026 at five of ~13,600 sites
LessonA test with no stated metric cannot produce an expansion decision
Taco Bell voice AIQSR · 2023–2025
AmbitionFaster, more accurate ordering and automated upsell at 500+ sites
What happenedOrder errors and viral prank orders; response was segmentation, not withdrawal
LessonWhere the technology works is a per-site, per-daypart question, not a yes/no
Just Walk Out (grocery)Retail · 2018–2024
AmbitionRemove the checkout entirely from Amazon Fresh stores
What happenedWithdrawn from Fresh for Dash Carts; retained in small-basket formats
LessonPointed at the wrong basket size — a Gate-1 error corrected late
Presto VoiceRestaurant tech · 2021–2025
Ambition"Eliminate human order taking" in the drive-thru
What happenedSEC cease-and-desist: 100% human intervention at most locations, 70% at the best
LessonAn automation rate without its denominator is a marketing number
Zillow OffersReal estate · 2018–2022
AmbitionBuy and resell homes on an algorithmic price estimate
What happened$407.9m FY21 write-downs; business wound down; ~25% of staff cut
LessonThe question is the error distribution against the margin, not the accuracy
Amazon recruiting modelTech HR · 2014–2018
AmbitionRank applicants one to five stars from a decade of résumés
What happenedLearned to penalise "women's"; neutralised, then killed before deployment
LessonThe one correct outcome in this gallery — and it is remembered as a failure
Rite Aid facial recognitionRetail · 2012–2023
AmbitionIdentify shoplifters across hundreds of stores
What happenedEight years live; FTC five-year ban after false flags, disproportionately of people of colour
LessonLate governance destroys work already paid for
NYC MyCity chatbotPublic sector · 2023–2026
AmbitionHelp small businesses navigate municipal regulation
What happenedDocumented giving illegal advice Mar 2024; killed Feb 2026 on a change of administration
LessonA launched initiative is a political asset; only an outsider can retire it
Lattice "digital workers"HR software · 2024
AmbitionGive AI agents employee records, goals and managers
What happenedWithdrawn 72 hours after announcement
LessonThe control case for New York — cure time is a function of who may say no
Deloitte Australia reportConsulting · 2025
AmbitionDeliver a welfare-compliance review under a ~A$440k contract
What happenedFabricated references and a made-up judicial quotation; A$97k refunded
LessonConfidently formatted work passes the checks built for work that used to be checked
Evidence: Each card compresses a case documented and cited in §§1–11 of this brief; the sources are the filings, orders, audits, releases and reporting listed against each case in the evidence register.
Method: "Ambition" is the sponsor's own stated purpose at launch where one was disclosed, not a third-party characterisation; where none was disclosed (McDonald's) the card says so in the outcome line. "What happened" carries only figures asserted elsewhere in this brief. "Lesson" is AIR's transferable reading, not a claim any of these organisations has made. Every case here is a named, publicly documented programme; no archetype appears in this figure.
Synthesis: Read the middle rows in sequence and the pattern is hard to miss: in eleven of sixteen the technology performed close to specification on the narrow task it was given. The failure sits in the row above it or the row below.

Sixteen cards, nine modes. Most of these programmes fail at two or three modes at once, and the modes they share turn out to be the ones decided before a model was chosen.

Figure 12Nine failure modes against three tests
Visible in the demo?Caught by model evals?Fixable by a better model?1 No value on offerNoNoNo2 Value not ownableNoNoNo3 Context mismatchNoPartlyPartly4 Silent accuracy decayNoYesPartly5 Trust & liabilityNoPartlyNo6 Unit economicsNoPartlyPartly7 Capacity never bankedNoNoNo8 Governance killNoPartlyNo9 Success theatreNoNoNo
Evidence: Modes derived inductively from the cases in §§1–13, each cited in place. "Caught by model evals" is assessed against standard offline evaluation practice (held-out accuracy, benchmark suites, red-teaming).
Method: Qualitative three-valued classification by AIR; no measurement or frequency is implied. "Partly" denotes modes where evaluation surfaces a symptom (e.g. degraded accuracy on a shifted distribution) but not the business consequence.
Synthesis: None of the nine is visible in a vendor demo. Only one — silent accuracy decay — is reliably caught by model evaluation. Six of nine cannot be fixed by a better model at any price, which is the whole argument of this brief in one grid.
Figure 13The taxonomy — mode, mechanism, the case that defines it, and how to detect it in your own program
#Failure modeCausal mechanismDefining caseDetection question
1No value on offerThe problem was never worth money, or the saving is too small to survive its own overheadLattice "digital workers"; McDonald's AOT (metric never stated)Can you name the P&L line, its owner, and the size of the move — before the model exists?
2Value not ownableThe gain is real but accrues to customers, competitors or the vendor, not to the firmRepresentative archetype, not a specific company — commodity chat deflection; generic productivity assistantsIf every competitor deploys the same tool from the same vendor, what do you still keep?
3Context mismatchThe model was optimised for conditions that do not obtain where the work happensGoogle Health Thailand (21% images rejected); MD Anderson (built on the prior EHR)Has anyone on the team watched the work being done, in the room, for a full shift?
4Silent accuracy decayLive performance diverges from validation and nothing in production surfaces itEpic Sepsis Model (AUC 0.63 external; 14.7% sensitivity in 2023 EDs)What is your production metric, measured on your data, and when was it last recomputed?
5Trust & liability collapseOutput binds the firm or loses user confidence faster than accuracy improvesMoffatt v. Air Canada; NYC MyCity chatbot; Taco BellWhat is the expected cost of one wrong answer × conversation volume — not the error rate?
6Broken unit economicsCost, latency or an irreducible human tier makes the per-transaction case negativePresto Voice (~70% human intervention); Just Walk Out in grocery basketsWhat is the fully loaded cost per transaction including escalation, review and rework?
7Capacity never bankedTime is genuinely freed and reallocated to other tasks; no cost line movesDenmark: 85% reallocate savings, null effect on earnings and hoursWho has committed, in writing and by date, to convert the freed hours into what?
8Governance killA control, legal or regulatory constraint terminates a working system after sunk costFTC/Rite Aid; EEOC/iTutorGroup; Mobley v. WorkdayWas legal, risk and compliance in the room at problem selection, or at go-live?
9Success theatreTermination costs the sponsor more than the program costs the firm, so it persistsNYC MyCity ($100m+, 4 years, chatbot killed only on a change of administration)What written criterion would kill this program, who applies it, and on what date?
Evidence: Each defining case is documented and cited in §§1–12 of this brief.
Method: Modes are inductive categories, not mutually exclusive; most real programs fail at two or three simultaneously (MD Anderson fails 3, 8 and 9). Ordering is by the gate at which the mode bites, not by frequency; no frequency data exists that would survive the comparability test in §15.
Synthesis: Modes 1, 2, 7 and 9 — four of nine, and in the case record the most expensive four — are decided entirely before any model is trained or after it has stopped mattering.

15 · The aggregate evidence, and how much of it to believe — the numbers everyone quotes

The "95% of AI pilots fail" statistic has done more to shape enterprise conversation in the past year than any case in this gallery. So handle it properly: neither repeat it nor dismiss it.

It comes from "The GenAI Divide: State of AI in Business 2025", produced by Project NANDA at the MIT Media Lab, research period January–June 2025. The executive summary: "Despite $30–40 billion in enterprise investment into GenAI, this report uncovers a surprising result in that 95% of organizations are getting zero return… Just 5% of integrated AI pilots are extracting millions in value, while the vast majority remain stuck with no measurable P&L impact."

Four things about it, in order of importance.

What it actually measured. Not pilots in general. The funnel is explicit and it splits by tool class: for general-purpose LLMs, 80% investigated → 60% piloted → 50% successfully implemented; for embedded or task-specific enterprise GenAI, 40% → 20% → 5%. The 95% is the complement of that last figure. The research note defining success is doing an enormous amount of work: "We define successfully implemented for task-specific GenAI tools as ones users or executives have remarked as causing a marked and sustained productivity and/or P&L impact."

That is a high bar and a soft instrument at once — a demanding criterion assessed by asking people. A tool delivering a genuine 3% cost reduction in one department, unremarked by an executive, scores as a failure. So does a tool that a manager enthusiastically credits with a "marked" impact that never appears in the accounts.

What the authors themselves say about it. The report's own research limitations are candid and almost never quoted: "These figures are directionally accurate based on individual interviews rather than official company reporting. Sample sizes vary by category, and success definitions may differ across organizations." The methodology is 300 publicly disclosed initiatives reviewed, 52 organisations interviewed, and 153 senior-leader survey responses "collected across four major industry conferences" — a sampling frameThe pool a survey draws its respondents from, which silently decides what the answers can mean. Ask AI questions of people who paid to attend AI conferences and you have not surveyed enterprises; you have surveyed enterprises already doing this, which is a different and much smaller country. that over-represents firms actively pursuing AI.

Its internal inconsistency. In the investment section the takeaway line states "50% of GenAI budgets go to sales and marketing" while the body of the same section says "Sales and marketing functions captured approximately 70 percent of AI budget allocation." Both cannot be right. The allocation itself came from asking executives to divide a hypothetical $100. This is not a fatal flaw in a directional finding; it is a decisive reason not to quote either number to two significant figures.

Why it survives the criticism anyway. Because it is not the only measurement pointing the same way, and the others use different methods. That convergence — not NANDA's precision — is what makes the conclusion durable.

Figure 14Two different funnels inside one statistic
General LLM: investigated80%General LLM: piloted60%General LLM: implemented50%Task-specific: investigated40%Task-specific: piloted20%Task-specific: implemented5%
Evidence: MIT NANDA, "The GenAI Divide: State of AI in Business 2025," §3.2 "The Pilot-to-Production Chasm," exhibit and accompanying research note.
Method: Values as published. Success for task-specific tools is defined by the report as tools "users or executives have remarked as causing a marked and sustained productivity and/or P&L impact" — a self-report instrument, not audited financials. The report states its own figures are "directionally accurate based on individual interviews rather than official company reporting."
Synthesis: Red marks the 5% that produces the "95% failure" headline. The headline conflates two populations: generic tools that people adopt easily and rarely change the P&L, and bespoke tools that could change the P&L and almost never ship.
Figure 15Eight headline numbers — and what each one actually counts
Source & dateFigurePopulation & methodExact definition of "failure"Tier
MIT NANDA
Jul 2025 (research Jan–Jun 2025)
95% zero return; 5% of task-specific tools reach production300 public initiatives, 52 interviews, 153 conference surveysNo "marked and sustained" productivity or P&L impact as remarked by users or executives3
Gartner
29 Jul 2024
At least 30% abandoned after PoC by end-2025Analyst forecast (not a measurement)Abandonment after proof of concept, from poor data quality, inadequate risk controls, escalating costs or unclear business value2
Gartner
25 Jun 2025
Over 40% of agentic projects cancelled by end-2027Analyst forecast; separate Jan 2025 poll of 3,412 webinar attendees on investment postureProject cancellation, same three causes2
S&P Global Market Intelligence
Mar 2025
42% abandoned most AI initiatives (17% in 2024); 46% of PoCs scrapped>1,000 enterprises, North America and EuropeAbandonment of the majority of initiatives; PoCs discarded before production2
RAND
13 Aug 2024
"More than 80% of AI projects fail"65 structured interviews with data scientists and engineersNot measured by RAND — attributed to "some estimates"; RAND's own contribution is the five root causes2
McKinsey
5 Nov 2025
39% attribute any enterprise EBIT impact; ~6% are high performers1,993 respondents, 105 countriesInability to attribute enterprise-level EBIT impact to AI; most of the 39% report under 5%2
Deloitte
Tech Trends 2026 (2025 survey)
11% of organisations run agentic AI in production; 38% pilotingDeloitte 2025 Emerging Technology Trends surveyNot "failure" — a production-adoption count, frequently misquoted as an 89% failure rate2
US Census Bureau / Federal Reserve
May 2026 / Apr 2026
19.8% of US firms use AI in a business function; 78% of the labour force works at an adopting firmBTOS (large probability sample); Fed comparison of three surveysNot a failure measure — the denominator problem itself1
Evidence: Each source opened this session — NANDA report PDF; Gartner releases as republished verbatim by THE Journal (Jul 2024) and BigDATAwire (Jun 2025), Gartner's own site being inaccessible to automated retrieval; CIO Dive's report of the S&P survey; RAND RR-A2680-1; McKinsey's Nov 2025 survey as reported by Silicon Canals; Deloitte Insights "Agentic AI strategy"; Census Bureau BTOS release (May 2026) and Federal Reserve FEDS Note (3 Apr 2026).
Method: Deliberately presented as a table rather than a chart. The eight figures use different denominators (organisations, projects, proofs of concept, initiatives, firms) and different outcomes (abandonment, non-production, no EBIT attribution, adoption), so plotting them side by side would imply a comparability that does not exist. Two are forecasts, not measurements; one (RAND's 80%) is a citation of an external estimate, not RAND's own finding.
Synthesis: The convergence is real and the precision is fake. No two of these measure the same thing, and every one of them lands between "most" and "the overwhelming majority." That is the strongest form the claim can honestly take.
Figure 16How many organisations can point to measurable financial impact? The honest answer is a range
5%MIT NANDA39%McKinsey any EBIT11%Deloitte agentic
Evidence: Low = MIT NANDA's 5% of integrated pilots "extracting millions in value" (Jul 2025). Centre = Deloitte's 11% of organisations running agentic AI in production (2025 survey, published in Tech Trends 2026). High = McKinsey's 39% attributing any enterprise-level EBIT impact (Nov 2025), of whom most report under 5% of EBIT.
Method: These are three different questions, deliberately shown as a spread rather than reconciled: "extracting millions," "in production," and "any attributable EBIT impact" are progressively weaker tests, which is why the number rises with the weakness of the test. Not a probability distribution and not an average.
Synthesis: The number you get depends almost entirely on how demanding the test is — which is the finding, not an obstacle to it. At the only test a CFO would accept (a material, attributable EBIT move), the answer is in the single digits.

One more piece of arithmetic, and it belongs to the Federal Reserve rather than to any consultancy. In April 2026 a Fed FEDS Note compared three high-quality surveys and found that Census Bureau data put firm-level AI adoption at about 18% at year-end 2025, individual work-related generative AI use at about 41%, and the Survey of Business Uncertainty at 78% of the labour force working at firms that have adopted AI. All three are correct. They count firms, people and employment-weighted exposure respectively. The note's own list of causes for the spread is instructive: "differences in sampling distributions and units of analysis… question framing, the materiality of reported usage, information asymmetries between different target respondents, and social desirability bias."

Census's May 2026 release puts the national rate at 19.8% of firms as of 3 May 2026, having hovered between 17% and 20% since December 2025 — and firms expected to be using AI in six months have hovered between 20% and 23% across the same period. The forward expectation has been running about three points above the current rate for half a year, and the current rate has barely moved. That is what a plateau looks like in survey data.

Figure 17Adoption is a large-firm phenomenon — and the national rate is under one in five
250+ employees37%100–249 employees32%All firms (national)19.8%1–4 employees<20%
Evidence: US Census Bureau, Business Trends and Outlook Survey, "Large Firms With at Least 20 Employees Biggest AI Users," collection period ending 3 May 2026.
Method: "AI use in a business function," firm-weighted, not employment-weighted; the employment-weighted figure is materially higher. The 1–4 employee bar is shown at 19% because Census reports it only as "less than 20%" — it is an upper bound, not a point estimate, and should not be read as 19.0%. Sector rates: Information 39.7%, Finance and Insurance 33.9%.
Synthesis: Even at the top of the size distribution, roughly six in ten large US firms report no AI use in any business function at all. The failure-rate debate concerns a minority of firms; the majority have not started.

Beneath that headline sits the best single piece of evidence in this entire brief, and almost nobody has read it. In 2026 Census researchers published "The Microstructure of AI Diffusion", using a dedicated AI supplement to BTOS — a nationally representative probability sampleA survey where every firm in the country has a known, non-zero chance of being picked, so the answers can be scaled up to describe the whole economy. It is the gold standard, and it is why a government statistical series and a vendor survey are not two versions of the same thing: one is measuring the population, the other is measuring whoever answered., not a conference survey — to look at three layers at once: firm-level use, deployment by business function, and worker-task use. Over the November 2025–January 2026 reference period, 18% of firms used AI in a business function, 32% employment-weighted, with 50–60% use (60–70% employment-weighted) among very large firms in Information, Professional Services and Finance.

Four findings from it displace a good deal of what circulates as fact.

Adoption is narrow even where it exists. "Among adopting firms, the scope of use remains limited: 57% of users integrate AI in three or fewer business functions, most commonly Sales and Marketing (52%), Strategy and Business Development (45%), and IT (41%)." That independently corroborates NANDA's sales-and-marketing skew — from a probability sample rather than a hypothetical $100 allocation exercise — and it settles the direction even though NANDA's own number could not be pinned down.

Augmentation, not substitution, is what is actually happening. "Most users (66%) rely on AI solely to augment tasks, while AI-related employment decreases are rare, occurring in only 2% of firms."

Figure 17aAdoption is not deployment — how narrow AI use is inside the firms that have it
Where AI sits inside the firms that have adopted itUse AI in three or fewer functions57%Sales and marketing52%Strategy and business development45%IT41%Augment tasks only, no substitution66%Report an AI-related employment decrease2%Denominators differ by row — see Method. Red = the headcount row, which is the one every business case assumes.
Evidence: US Census Bureau, Center for Economic Studies Working Paper 26-25, "The Microstructure of AI Diffusion," reference period November 2025 – January 2026; a dedicated AI supplement to the Business Trends and Outlook Survey.
Method: Values as published and not shares of one denominator: 57% is of functional adopters; the three function rows are of AI-using firms; 66% is of firms reporting task effects; 2% is of all firms. Plotted on a common axis to compare magnitudes only, not to sum or decompose. Firm-weighted throughout; employment-weighted figures differ and are higher.
Synthesis: A probability sample of the American economy says the same thing as the Danish payroll data and the CFO survey. Firms that have adopted AI use it in a handful of functions, overwhelmingly to augment, and almost never register a headcount consequence — which is Gate 4 measured from the outside rather than argued from the case record.

Diffusion runs in both directions and neither implies the other. "Worker task use sometimes occurs without formal firm-level adoption, and firm-level adoption sometimes occurs without worker task use." The first half is shadow AI measured properly; the second half is the deployment-without-behaviour-change failure measured properly.

And the finding that matters most. Regression shows a positive correlation between commercial performance and the breadth of AI integration, and it holds up. But "a distinct divergence emerges… with respect to labor outcomes. Functional breadth and operational investment are positively associated with employment decreases, whereas worker-task integration shows no significant link to headcount reduction once functional integration and operational investment are taken into account."

Giving people tools does not reduce headcount. Restructuring functions around AI, and investing operationally to do it, does. It is the American, cross-sectional, probability-sampled version of the Danish result — arrived at by a different method, in a different labour market, with different institutions — and it says the same thing. The operating change, not the tool, is the mechanism.

The lessonTwo independent national datasets, one Danish and one American, agree: worker-level AI use predicts no headcount effect, and functional restructuring does — which is Gate 4 stated as an empirical result rather than a framework.

PART II — THE DISCIPLINE

Fourteen sections, conclusion-forward. Every remedy below is matched to a numbered failure mode from Figure 13, and every one of them is an operating change that can be made before a model is chosen.

16 · Write the P&L line first — problem selection and the value hypothesis · fixes modes 1, 2

The single cheapest intervention available to an enterprise AI program costs nothing and takes an afternoon: before any tool is evaluated, write one sentence naming the account that will move, the person who owns that account, the size of the move, and the date it will be visible. If the sentence cannot be written, the pilot is not ready. The inability to write it is the diagnosis.

Test it against the cases. McDonald's could not, or would not, state the metric by which its two-year IBM test was judged — and the test ended without an expansion decision anyone could evaluate. CBA wrote a version of the sentence ("investment in technology, including AI, is making it easier and faster for customers to get help") that named no account, no baseline and no measurement, and then made a 45-person headcount decision on it, which it reversed six weeks later. Lattice's digital-worker feature had a narrative and no account at all.

A value hypothesis that passes has four components, and the fourth is the one that gets skipped.

The account. Not "efficiency." A general ledger line: contact-centre labour cost, claims leakage, days sales outstanding, first-pass yield, clinician documentation overtime. If nobody can name the account, no one will ever be able to prove the pilot worked, which means it will be judged on anecdote — and anecdote, per METR, is systematically biased upward.

The size, relative to the noise. A projected 2% improvement in a line that swings 8% quarter to quarter is unprovable and therefore unbankable, however real it is. This is the Zillow lesson generalised: compare the expected effect to the variance of the thing you are trying to move, not to zero.

The counterfactual. What happens if you do nothing, fix the process without AI, or buy the outcome? A surprising share of pilots are competing against a process fix that costs less and works better. Amazon's Dash Cart is exactly that judgment made in public: the outcome was worth having and the cheaper mechanism won.

The claim on the freed resource. Who has agreed, before the pilot starts, to give up what — a role not backfilled, a contractor budget cut, a queue that will absorb more volume without more people. This is Gate 4, moved to the front. Without it the honest projected value of a time-saving pilot is zero, and the Danish evidence says so: 85% of users reallocate savings to other tasks, and earnings and hours do not move.

Figure 18The value hypothesis test — four questions, asked before a vendor is called
  1. 01
    The account
    Name the general-ledger line and its owner. "Efficiency" is not a line. If none exists, stop.
  2. 02
    The size vs the noise
    Expected move compared to that line's historical quarterly variance. Smaller than the noise means unprovable.
  3. 03
    The counterfactual
    Do nothing / fix the process / buy the outcome. AI must beat the cheapest of these, not beat zero.
  4. 04
    The claim on the resource
    Named owner commits in advance to what is given up when capacity is freed, and by when.
Evidence: Step 4 is anchored on NBER WP 33777 (85% of users reallocate time savings; null effect on earnings and hours). Step 2 reflects Zillow's disclosed conclusion that "unpredictability in forecasting home prices far exceeds what we anticipated." Steps 1 and 3 reflect the McDonald's and Amazon Fresh cases respectively.
Method: Prescriptive framework authored by AIR; the sequencing is a design choice, and no claim is made that firms applying it show measured higher success rates — no such dataset exists.
Synthesis: Three of the four steps are finance and operations questions. None requires knowing which model you will use, which is why they can all be answered before any spend.

17 · Portfolio and kill rates — ruthless prioritisation · fixes modes 1, 9

Most enterprises set out to reduce the number of failed pilots. That is the wrong objective, and pursuing it produces exactly the pathology in §11: fewer terminations, longer zombie programs, more sunk cost.

A healthy portfolio has a high kill rate and a short time to kill. What NANDA found about speed is the sharpest operating datum in the report and has nothing to do with technology: mid-market top performers reported average timelines of 90 days from pilot to full implementation; enterprises took nine months or longer, and reported the lowest pilot-to-scale conversion rates despite leading in pilot count and staffing. More pilots, more people, more time, worse outcomes.

The mechanism is arithmetic. A pilot's expected value is roughly (probability of value) × (size of value) − (cost of running it) − (cost of the delay in learning). Enterprises optimise the third term, which is small, and ignore the fourth, which compounds. A nine-month pilot cycle gives you at most one learning iteration a year per team; a 90-day cycle gives four. Over two years that is an eightfold difference in accumulated evidence about your own operating context, which is the scarce input.

Three portfolio rules follow. Each is unpopular.

Write the kill criterion into the charter, with a date and a named person who applies it. Not "we will review progress." A number and a deadline: "if human intervention exceeds 25% of transactions at day 60, this stops." This works for a reason that has nothing to do with analytical rigour: it moves the termination decision away from the sponsor, who is punished for making it, and onto a pre-commitment made when nobody's reputation was yet attached.

Cap the number of concurrent pilots at the number of business owners you actually have. Not the number of use cases, not the budget divided by unit cost. Gate 4 requires an owner with authority over the resource; if you have six such people, you can run six pilots. NANDA's enterprise paradox — pilot volume up, conversion down — is what happens when this cap is ignored.

Report the kill rate to the board as a health metric, not a failure metric. A program that has terminated nothing in eighteen months is not succeeding — it has stopped measuring.

A portfolio with no kills is a portfolio with no measurement.
On pilot governance

18 · Sponsorship and decision rights — org design · fixes modes 1, 7, 9

RAND's contribution to this literature is often mis-cited for the 80% figure it borrowed from elsewhere. Its actual finding, from 65 structured interviews with data scientists and engineers of five-plus years' experience, is that the leading root cause of AI project failure is that "industry stakeholders often misunderstand — or miscommunicate — what problem needs to be solved using AI. Too often, trained AI models are deployed that have been optimized for the wrong metrics or do not fit into the overall business workflow and context."

Its first recommendation follows: "Ensure that technical staff understand the project purpose and domain context… Misunderstandings and miscommunications about the intent and purpose of the project are the most common reasons for AI project failure." Its second is the one no enterprise wants to hear: "Choose enduring problems. AI projects require time and patience… leaders should be prepared to commit each product team to solving a specific problem for at least a year. If an AI project is not worth such a long-term commitment, it most likely is not worth committing to at all."

That seems to contradict §17's argument for 90-day cycles. It does not. RAND's year is a commitment to the problem. The 90 days is a cycle on the solution. A team that owns "reduce claims-handling cycle time" for a year and runs four experiments against it accumulates domain knowledge; a team that owns "deploy the vendor's claims assistant" for nine months accumulates a vendor relationship. Almost every enterprise AI operating model in the field commits to solutions and rotates through problems, which is precisely backwards.

On structure, NANDA is direct: "The dominant barrier to crossing the GenAI Divide is not integration or budget, it is organizational design. Our data shows that companies succeed when they decentralize implementation authority but retain accountability." Decentralised authority with retained accountability is a specific arrangement, not a slogan: the business unit chooses and owns the use case and the value commitment; a central function owns the platform, the evaluation harness, the model contracts and the veto. Central teams that own use-case selection produce a queue. Business units that own the platform produce nine incompatible ones.

One further point on sponsorship, drawn from the incentive map. The sponsor should not be the person who benefits from the announcement. Where possible, make the accountable executive the one who must give up the resource — the operations leader whose headcount line falls, not the technology leader whose budget rises. This single reassignment collapses most of modes 1, 7 and 9 at once, because it puts the person with the strongest reason to be sceptical in charge of the go/no-go.

19 · Buy the workflow, build the edge — sourcing · fixes modes 3, 6, 9

The build-versus-buy answer in the data is lopsided. In NANDA's sample, "external partnerships with learning-capable, customized tools reached deployment ~67% of the time, compared to ~33% for internally built tools," with pilots built through strategic partnerships "2x as likely to reach full deployment" and employee usage rates "nearly double for externally built tools." The report is careful — self-reported outcomes, 52 organisations, correlation not causation — and Deloitte's independent 2025 survey reports the same 2× relationship for agentic pilots.

The internal build's cost curve is back-loaded and invisible at approval time. What gets approved is the build. What kills it is year two: the evaluation harnessThe standing test rig that scores a model against a fixed set of cases with known right answers, so you can tell whether this week's version is better or worse than last week's. Without one you are not managing a system, you are reacting to complaints — and it is invariably the line cut first from a build budget, because at launch there is nothing yet to compare against. nobody budgeted, the model version upgrade, the departure of the one engineer who understood the retrieval layerThe machinery that finds the right internal documents and feeds them to the model before it answers, so the answer is grounded in your policies and records rather than in whatever the model absorbed during training. Most of the difficulty in a corporate assistant lives here, not in the model., the integration that breaks when the source system is upgraded. MD Anderson is this failure in its purest form — the asset did not decay, its environment moved, and no one owned keeping up.

The rule that follows is a decomposition:

Buy the workflow. If the process is one thousands of firms run similarly — contract review, call summarisation and routing, document intake, code assistance, expense classification — buy it, and buy it from a vendor whose product is that workflow rather than a platform you must assemble it on. NANDA's own list of categories that succeeded is exactly this: "voice AI for call summarisation and routing, document automation for contracts and forms, code generation for repetitive engineering tasks." Its list of failures is equally instructive: "complex internal logic, opaque decision support, or optimization based on proprietary heuristics."

Build the edge. Build only where the asset is genuinely yours and durable — your proprietary data, your pricing logic, your risk model, the thing a competitor buying the same vendor product cannot replicate. This is failure mode 2 answered in the sourcing decision: if every competitor can buy your advantage, you have bought a cost, not an advantage, and you should buy it as cheaply as possible.

Never build the plumbing. Retrieval, evaluation harnesses, observabilityInstrumentation that lets you see what a live system is actually doing — what it was asked, what it answered, where it slowed down, when its behaviour drifted. The Epic sepsis case is what an unobserved deployment looks like from the inside: running fine, by all appearances., guardrail layersThe filters wrapped around a model that block certain inputs and outputs — the topics it may not discuss, the claims it may not make, the data it may not repeat. They sit outside the model rather than inside it, which is why they can be changed without retraining anything., model routingSending each request to whichever model suits it — a cheap fast one for the easy majority, an expensive one for the hard remainder — instead of paying top rate for everything. Plumbing, and the reason a stated price per query tells you very little about a real bill.. These are commodity, they are where internal builds sink their year-two budget, and the buy-versus-build calculation on them has been settled for two years.

One caution on the 67/33 figure, since it will be quoted. It is self-reported, from 52 organisations, in a report whose author acknowledges "the correlation between external partnerships and success does not necessarily prove causation" — and there is an obvious selection effectWhen the way cases end up in each group, rather than the treatment itself, produces the difference you are measuring. Here: firms build in-house precisely when nothing can be bought, so the "build" column is stocked with the hardest problems before anyone writes a line of code., since organisations attempt internal builds precisely where no product exists, which is also where problems are hardest. The direction is well-supported; the ratio is not a constant of nature.

Figure 19What actually crosses to production, and what stalls in pilot
Crossed to production
  • Voice AI for call summarisation and routing
    Bounded task, immediate visible output, low configuration burden, a human still owns the decision
  • Document automation — contracts and forms
    High-volume structured extraction where the error is caught downstream by an existing check
  • Code generation for repetitive engineering
    The user is also the reviewer; verification is instant and free
  • Ambient clinical documentation
    Attacks the task practitioners most resent; adoption is voluntary and self-sustaining
Stalled in pilot
  • Complex internal logic and orchestration
    Deep enterprise specificity; the configuration burden exceeds the value before go-live
  • Opaque decision support
    Users cannot verify the output cheaply, so they do not act on it — deployment without behaviour change
  • Optimisation on proprietary heuristics
    Requires codifying knowledge the organisation has never written down
  • Broad-scope, complex-execution agents
    "Fails" in NANDA's own scope/complexity matrix; Gartner's agentic cancellation forecast points the same way
Evidence: MIT NANDA §5.2.1 (successful and struggling categories; the narrow/broad × simple/complex matrix) and §6.1 (deployment rates by sourcing); Gartner, 25 Jun 2025 (agentic cancellation forecast); Permanente Medical Group ambient-scribe results (Kaiser Permanente Division of Research, 2025) for the fourth winning row.
Method: A qualitative two-bucket sort of categories named in the sources, not a ranked or measured list. The "why" column is AIR's causal reading, not the sources' wording.
Synthesis: The winners share one property the losers lack — cheap verification. Where a user can tell in seconds whether the output is right, adoption and value follow; where verification is expensive, deployment happens and behaviour does not change.

20 · Production engineering — evals, monitoring, reliability, cost · fixes modes 3, 4, 6

The Epic Sepsis Model makes a boring proposition unarguable: a deployed model without a live, on-your-own-data performance measurement is an unmonitored liability that looks like an asset. It ran for years. Its degradation relative to the vendor's cited range was found by academics, not by the hospitals paying for it.

Four engineering commitments follow, each mapped to a case above.

Compute the metric on your data, at your threshold, on a schedule. A vendor's accuracy figure is a claim about their evaluation set. Michigan's AUC of 0.63 against a cited 0.76–0.83 is the size of the gap that is possible. Texas made this an enforceable obligation for one vendor in one sector; buyers should impose it contractually everywhere.

Measure decision value, not just discrimination. The sepsis model's 14.7% sensitivity is bad. Its zero-minute median lead time is disqualifying, and it is a different kind of measurement — it asks whether the output arrives in time to change anything. Every predictive deployment needs the equivalent: for a churn model, how long before the churn; for a fraud model, before or after the funds move; for a maintenance model, before or after the failure window closes.

Instrument the human tier as a first-class metric. Presto's ~70% human-intervention rate was, in the SEC's account, invisible to investors until an investigation. Whatever your equivalent — escalation rate, override rate, share of outputs edited before use, images rejected — that number is the honest measure of automation achieved, and it should sit on the same dashboard as accuracy, computed weekly. Google Health's 21% rejection rate is the same statistic wearing different clothes; had it been a tracked KPI rather than a finding, the deployment design would have changed.

Model the fully loaded cost per transaction, including rework. Inference is the cheapest line and falling. Escalation handling, review labour, incident response and the legal review of what the system may say are the expensive lines and they are not falling. Gartner's $5–20 million range for business-model-innovation deployments is mostly not compute. A pilot business case that shows only model cost is not a business case.

Figure 20Four production metrics, each mapped to the failure it would have caught
MetricDefinitionCadenceFailure it would have caught
Own-data performanceThe vendor's headline metric, recomputed on your population at your operating thresholdMonthly, and on every model version changeEpic Sepsis Model — AUC 0.63 on external data against a cited 0.76–0.83
Decision lead timeTime between the output and the moment the decision must be madeMonthly, reported as a median with intervalEpic Sepsis Model — median lead time of zero minutes across 145,885 encounters
Human intervention rateShare of transactions requiring a person to enter, correct, validate or escalateWeekly, on the same dashboard as accuracyPresto Voice — ~70% of orders on the proprietary system; Google Health Thailand — 21% of images rejected
Fully loaded cost per transactionInference + escalation labour + review + rework + incident and legal overheadQuarterly, against the pre-agreed baselineJust Walk Out in large-basket grocery; Gartner's $5–20m deployment cost range
Evidence: Wong et al., JAMA Internal Medicine (21 Jun 2021) and Ostermayer et al. (2023 cohort) for rows 1–2. Wong's published abstract carries the 38,455 hospitalisations, the AUC of 0.63 (95% CI 0.62–0.64) and the finding that the model "did not identify 1709 patients with sepsis (67%)" — from which the 33% sensitivity follows; the 12% positive predictive value and the vendor-cited 0.76–0.83 range appear in the paper body, and the PPV reconciles arithmetically (≈842 true positives against 6,971 alerts at a score of ≥6 = 12.1%); SEC Rel. 33-11352 ¶26 and Beede et al. via MIT Technology Review (2020) for row 3; Retail Dive (2024) and Gartner (Jul 2024) for row 4.
Method: Prescriptive mapping by AIR. Cadences are recommendations, not observed practice. Each "failure it would have caught" is a counterfactual claim: the metric was measurable at the time and was not being tracked by the deploying organisation.
Synthesis: None of the four is technically difficult and none requires new tooling. All four were absent in the cases above, which is the point — this is a measurement discipline problem wearing an engineering costume.

21 · Attribution a board can audit — value measurement · fixes mode 7

Modes 3, 4 and 6 are engineering problems with engineering answers. Mode 7 — capacity freed and never banked — is a measurement and authority problem, and it defeats programs where everything else went right.

Multiplying users by hours saved by loaded cost, where hours saved comes from a survey, produces a number that is both large and false. METR's developers, reporting on their own recent work in a controlled setting, were off by 39 percentage points in the flattering direction. There is no reason to expect a claims processor's estimate of time saved to be better calibrated than a senior engineer's.

Three methods do work, in ascending order of rigour and descending order of convenience.

The holdout. Withhold the tool from a randomly selected, comparable group for a defined period and compare the outcome metric — issues resolved per hour, cases closed, cycle time — not the self-report. This is what Brynjolfsson and colleagues effectively exploited with a staggered rollout across 5,179 agents, and it is why their 14% average and 34% novice effect are believable in a way that vendor case studies are not. Most enterprises can do this and decline to, because it delays universal rollout by a quarter and someone has to explain to the withheld group why.

The pre-registered baseline. Before deployment, agree with the finance function on the account, the measurement window, the seasonality adjustment and the definition of the counterfactual. Write it down. Pre-registration is not a statistical nicety. After deployment every party's incentive is to redefine success, and a document written before anyone's reputation was attached is the only defence.

The capacity conversion ledger. The one nobody keeps. For every unit of capacity a pilot frees, record what happened to it: converted to headcount not backfilled, converted to volume absorbed without hiring, converted to a service level a customer pays for, or reallocated to other work. The fourth category is the honest destination of most of it, and naming it is what converts an inflated business case into a real one. The Danish study's 85% is the base rateWhat normally happens, before you account for anything special about your own case. It is the number a forecast should start from and argue away from, rather than the number people reach for only after their own estimate has already been proved wrong.; a program that assumes better than that needs to explain the mechanism.

Figure 9 showed the arithmetic, and it inverts standard practice. Raising conversion from 15% to 50% at fixed 40% adoption adds about $1.58 million on those parameters; doubling adoption from 40% to 80% at fixed 15% conversion adds about $675,000. Enterprises spend nearly all their change-management energy on adoption and almost none on conversion, and conversion is worth more than twice as much per unit of movement. At zero conversion the program is negative at every adoption level including 80%, which is the mathematical statement of why so many well-adopted tools show up nowhere in the accounts.

Figure 21The attribution ceiling
Orgs attributing any enterprise EBIT impact to AI39%Of the 88% using AI in a function; most of the 39% put it below 5% of EBIT (McKinsey, Nov 2025)
Evidence: McKinsey State of AI survey, 1,993 respondents across 105 countries, fielded mid-2025, published 5 Nov 2025, as reported by Silicon Canals; McKinsey's own page was not retrievable by automated request this session.
Method: Denominator is all surveyed organisations, 88% of which report AI use in at least one function. "Any enterprise-level EBIT impact" is a weak test — most of the 39% place the figure below 5% of EBIT, and roughly 6% meet McKinsey's high-performer threshold. Because the figure is a report of a report, confidence is lower than for the primary-document sources in this brief.
Synthesis: Six in ten organisations using AI cannot attribute any enterprise-level profit effect to it. That says less about the technology's power than about the absence of an attribution method.

22 · Adoption in fact — change management · fixes modes 3, 5, 7

The best-documented adoption success in enterprise AI is dictation.

The Permanente Medical Group deployed ambient AI scribes — systems that listen to a clinical encounter and draft the note — across Northern California. In the first year, physicians used the technology more than 2.5 million times, and analysis found it saved nearly 16,000 hours of documentation time. Of 102 adult and family medicine physicians surveyed, two-thirds used it five or more days a week and 63% used it in every in-person visit. Crucially: "We found the highest adoption rates in departments that typically suffer from the highest levels of documentation and burnout," and the users who benefited most were the highest-volume users, whose time savings "substantially surpassed the time savings among their peers who used the technology infrequently or not at all."

Four properties made that work, and each is an operating choice available to any organisation.

It attacked the task the users most hated, so adoption did not need to be mandated. It produced output the user verifies in seconds — the physician reads the note they were going to write anyway — so trust was established transactionally rather than institutionally. It was voluntary, so non-adoption was informative rather than hidden. And no headcount reduction was attached to it, which removed the single largest reason for users to sandbagQuietly underperform on purpose — use the tool sparingly, report modest gains, keep the old process running alongside it. Rarely a conspiracy and almost never visible in a metric, because a person protecting their job looks exactly like a person who finds the tool unhelpful. a deployment.

The last of those should be uncomfortable. If a tool's stated purpose is to reduce headcount, the people who must adopt it are being asked to build the case for their own redundancy, and they will not do it well. CBA announced 45 redundancies and then discovered the workload had not fallen. Klarna's most consequential retreat was not technical but positional — from replacement to augmentation, with a guaranteed human path. The organisations getting adoption in fact are, so far, overwhelmingly the ones that separated the tool from the headcount question in time.

Adoption alone is not the goal. BetterUp's workslop finding is the counterweight: 40% of desk workers received AI-generated content that looked like work and was not, costing about two hours each time and debiting the sender's credibility. High adoption of an ungoverned generative tool can be net-negative at the firm level while every individual reports a gain. The governance that prevents this is unglamorous and cheap — a norm that the person who generates output owns its verification, and an expectation that AI-assisted work is disclosed to the colleague who must build on it.

NANDA found that while only 40% of companies had bought an official LLM subscription, workers at over 90% of surveyed companies used personal AI tools for work. That is shadow AIEmployees using AI tools they bought or signed up for themselves, on company work, outside any policy or contract. It is what adoption looks like when nobody has to be persuaded — and it runs on personal accounts, which is why it is invisible to the programme and visible only in surveys., and the security lens gets it backwards. It is free, high-quality market research about which tasks your people find AI genuinely useful for — a revealed-preferenceWhat people's behaviour shows they want, as opposed to what they say they want in a survey. Behaviour costs something and answers do not, which is why the tool an employee pays for out of their own pocket is better evidence than the one they rated highly in a pilot questionnaire. dataset most organisations are trying to suppress rather than read. The tasks where shadow use is heaviest are the tasks where an official deployment will get adoption in fact.

The lessonWhere people already use a tool without permission, adoption is solved and only value capture remains — start there, not with the use case that looked best in the vendor's deck.

23 · Governance as accelerator — fixes mode 8, and makes 5 survivable

The standard complaint is that risk and legal slow AI down. The case record says something more precise: late governance slows AI down, and it does so by destroying work that has already been paid for. Rite Aid deployed for eight years and then had the capability removed by consent orderA settlement a regulator writes and a court enforces. The company admits no wrongdoing and still accepts binding obligations — here, a ban with a term of years — that it cannot later argue its way out of. Cheaper than losing a trial, and considerably more restrictive than winning one.. iTutorGroup shipped screening logic and paid $365,000 plus five years of EEOC monitoring. Workday is three years into litigation over a product function.

Three questions move from go-live to problem selection, where answering them is nearly free.

What can this system say or decide on our behalf, and what is the ceiling on the resulting obligation? This is the Air Canada question. It has a numeric answer — expected cost of an erroneous representation multiplied by volume — and it belongs in the business case, not the risk register.

What claims will we make about this system's performance, and can we substantiate them by the method we will publish? This is the Pieces and Presto question, and it now has regulators attached in at least two jurisdictions. Internally, the same discipline kills mode 9: a program that must state its metric methodology in writing cannot survive on a vibe.

Who is a protected party in this decision, and what is our evidentiary record if challenged? This is Mobley. If the system touches hiring, credit, insurance, housing or healthcare access, the record you keep from day one is the defence, and it cannot be constructed retroactively.

Two dated obligations now make the timing question concrete rather than theoretical. New York City's Local Law 144 has, since enforcement began on 5 July 2023, prohibited employers and employment agencies from using an automated employment decision tool "unless the tool has been subject to a bias audit within one year of the use of the tool, information about the bias audit is publicly available, and certain notices have been provided to employees or job candidates." A bias audit is not something you can produce retrospectively for a system already in use; the record either exists or it does not.

And the EU AI Act became applicable on 2 August 2026, three weeks before this brief's evidence cutoff, having entered into force on 1 August 2024. Prohibited practices and AI-literacy obligations applied from 2 February 2025; general-purpose model obligations from 2 August 2025. Crucially for planning, the timetable then moved — and it moved as enacted law, not as a proposal. The "AI Omnibus" simplification package, put forward by the Commission in November 2025, was published in the Official Journal on 24 July 2026 and entered into force on 27 July 2026, one month before this brief's cutoff. It defers the Annex III high-risk obligations — the sensitive-area list covering employment, credit, education and essential services — to 2 December 2027, and Annex I product-embedded systems to 2 August 2028. The prohibitions, the AI-literacy duty, the general-purpose model obligations and the Article 50 transparency and content-labelling duties are unaffected and already live.

Read that as an operator rather than a lawyer. An organisation planning against the original August 2026 date now has roughly sixteen further months on its highest-risk use cases, and it is a settled sixteen months rather than a hoped-for one — which is exactly the window in which the evidentiary record a bias claim will later be judged against can still be built cheaply. An organisation reading only the deferral will miss that the duties that bite first, on disclosure and prohibited practice, never moved at all. The deferral bought time on the paperwork, not on the conduct.

A decision boundary cleared in advance — "the assistant may quote published policy verbatim and may not construct an answer about entitlement" — lets a team ship in weeks without a bespoke legal review per feature. That is the acceleration, and it runs backwards from expectation. The slow path is the one where the boundary is negotiated after the demo has been shown to the board, when the sunk cost makes every restriction a fight.

There is a live counter-argument, and it is a fair one. Heavy pre-deployment governance can be an excuse for institutional paralysis, and some organisations use "responsible AI" processes to avoid deciding anything. The test that distinguishes the two: a governance function that produces reusable, written decision boundaries is an accelerator; one that produces case-by-case reviews with no accumulating precedent is a tax. Ask how many of last quarter's reviews created a rule that removed the need for a future review. If the answer is none, the complaint about slowness is correct.

24 · Sector overlays — the same failure, wearing local clothes

The nine modes are general. What varies by sector is which mode binds first and how expensive it is when it does.

Figure 22Sector overlays — the binding constraint and the local aggravator
SectorBinding modeLocal aggravatorCase in this briefWhere value has landed
Healthcare4 — silent accuracy decay; 3 — context mismatchClinicians cannot verify a prediction cheaply, and alert fatigue forces thresholds that destroy sensitivityEpic Sepsis Model; MD Anderson OEA; Google Health Thailand; Pieces TechnologiesDocumentation, not diagnosis — 2.5m ambient scribe uses at TPMG in one year
Financial services7 — capacity never banked; 5 — trust and liabilityRegulated advice and dispute exposure; deflection metrics that diverge from workloadCBA voice-bot; Klarna; SEC AI-washing actionsRoutine contact automation with a guaranteed human path; fraud and document work
Retail & QSR6 — unit economics; 5 — trustFranchise heterogeneity, accents and dialects, viral failure at consumer scaleMcDonald's AOT; Taco Bell; Just Walk Out; PrestoSmall-basket, high-friction formats; back-of-house forecasting and scheduling
Public sector9 — success theatre; 8 — governance killLaunch is the political deliverable; termination is an admission; audit cycles run in years. US federal agencies must now name a Chief AI Officer and apply minimum risk-management practices to "high-impact AI" under OMB M-25-21NYC MyCity ($100m+, chatbot live ~2 years after documented illegal advice); Deloitte Australia reportConstrained, verifiable transactions — the MyCity childcare portal did digitise a real application
Manufacturing & industrials3 — context mismatch; 1 — no value on offerSensor and process data collected for control, not for learning; site-by-site variation defeats one modelNot directly evidenced in this brief — see "not verified"NANDA records maintenance pilots with no major supply-chain shift; treat claims here sceptically
Technology7 — capacity never bankedGains are real and land in individual throughput, where the firm has no mechanism to collect themMETR RCT; workslop; Salesforce support reduction from ~9,000 to ~5,000Support automation with a measured baseline; code assistance where the user is the reviewer
Evidence: Cases as cited in Part I. Salesforce figures from CNBC's report of Marc Benioff's 2 Sep 2025 remarks — a chief executive's public remark, not an audited disclosure, and carried here on the same self-report caveat this brief applies to Klarna's figures. NANDA's sector table records "documentation/transcription pilots; clinical models unchanged" for healthcare and "maintenance pilots; no major supply chain shifts" for advanced industries.
Method: The "binding mode" column is AIR's judgment from the case record, not a measured ranking; sectors are not equally represented in the public evidence, and manufacturing is materially under-evidenced here — that row is flagged rather than filled.
Synthesis: The value that has landed, in every sector, sits in the same place: high-volume, bounded tasks where a human verifies the output in seconds. Nothing on the right-hand column is a decision-making system.

25 · The maturity model — self-assess honestly · covers all nine modes

Five levels, six dimensions. Score yourself at the lowest level you can honestly claim across the six, not the highest: the binding constraint is what determines the outcome. An organisation at Level 4 on engineering and Level 1 on value attribution is a Level 1 organisation with expensive infrastructure — which describes a large share of the programs behind the statistics in §15.

Figure 23The value maturity model — score at your lowest honest level, not your highest
DimensionL1 · TheatreL2 · ExperimentingL3 · ShippingL4 · CapturingL5 · Compounding
Problem selectionUse cases arrive from vendors and conferencesInternal list, ranked by enthusiasmWritten value hypothesis with a named accountHypothesis includes the counterfactual and the size-vs-noise testProblems owned for a year; solutions cycled quarterly against them
SponsorshipSponsor benefits from the announcementTechnology leader accountableBusiness leader accountable for deliveryAccountable executive is the one who gives up the resourceDecentralised authority, retained central accountability, veto exercised
Production engineeringVendor's accuracy figure is the metricOwn-data evaluation at go-live onlyScheduled own-data evaluation on your thresholdDecision lead time and human-intervention rate tracked weeklyFully loaded cost per transaction, with automatic rollback triggers
Value attributionUsers × hours saved × loaded cost, from a surveyPre/post comparison, no controlPre-registered baseline agreed with financeHoldout group; measured outcome, not self-reportCapacity conversion ledger reconciled to the general ledger
AdoptionSeats provisioned counted as adoptionWeekly-active users trackedAdoption voluntary; non-use treated as a signalShadow-AI usage read as demand data and acted onVerification norms in force; workslop measured and falling
GovernanceReviewed at go-live, or after an incidentChecklist applied per projectRisk and legal present at problem selectionWritten decision boundaries reusable across projectsKill criteria pre-committed with dates and named owners; kill rate reported to the board
Evidence: Level descriptors are constructed from the practices whose presence or absence is documented in the cases and studies of Part I — RAND's five root causes and five recommendations; NANDA's organisational-design and sourcing findings; the Epic, Presto and Google Health measurement gaps; NBER WP 33777 on capacity reallocation; BetterUp/Stanford on workslop; the FTC, EEOC, Texas AG and CRT actions on governance timing.
Method: Analytical instrument, not a validated assessment. No claim is made that firms at higher levels show measured higher returns — no dataset supports that, and any vendor maturity model claiming otherwise should be asked for it.
Synthesis: The L1→L2 move is cheap and the L3→L4 move is where almost every program stalls, because L4 requires an executive to give something up. That is an authority problem, and no amount of engineering maturity substitutes for it.

26 · The first 90 days — inheriting a pile of stalled pilots

Assume the realistic situation: you have arrived, there are somewhere between nine and forty active AI initiatives, nobody can tell you what any of them are worth, and the board wants a number by the next meeting. The instinct is to build a strategy. Do the inventory instead — the strategy falls out of it, and the inventory is defensible in a way the strategy is not.

Figure 24Ninety days — inventory, kill, prove, then propose
  1. 01
    Days 1–15 · Inventory
    One page per initiative: the account it moves, its owner, spend to date, run-rate, the last date its metric was computed on your data, and the human-intervention rate. Initiatives whose page cannot be completed are your answer to step 2.
  2. 02
    Days 15–30 · Kill
    Terminate everything with no named account or no resource claim. Announce the kill rate as a governance result, not a failure. Redirect the run-rate, not the headcount.
  3. 03
    Days 30–45 · Instrument
    On the survivors, stand up the four metrics of Figure 20 and a pre-registered baseline agreed in writing with finance. Nothing new is approved until these exist.
  4. 04
    Days 45–75 · Prove one
    Pick the single highest-volume bounded task with cheap verification and heavy existing shadow use. Run it with a holdout. The objective is one auditable number, not a portfolio.
  5. 05
    Days 75–90 · Propose
    Take the board the kill rate, the one measured result with its method, and a capped portfolio sized to the number of business owners who will sign a resource claim. Not a roadmap.
Evidence: Sequencing draws on NANDA's finding that top mid-market performers convert pilots in ~90 days against nine months or more for enterprises, and on RAND's recommendation to commit to enduring problems rather than to technologies.
Method: Prescriptive plan authored by AIR. Day ranges are a design choice; the ordering — inventory before strategy, kill before build, prove one before proposing many — is the load-bearing element.
Synthesis: The deliverable at day 90 is one auditable number and a shorter list. A new strategy at day 90 is the same artefact that produced the pile you inherited.

27 · Alternatives considered and rejected — the explanations that do not survive the cases

Four rival explanations for the value gap deserve a fair hearing. Each has real evidence behind it. None survives as the primary cause.

"The models aren't good enough yet." The strongest version: most of these failures date from a period of rapidly improving capability, and a sufficiently capable system would have handled Thailand's poor-quality images, Taco Bell's noisy drive-thru, and the sepsis signal. There is something to this — capability genuinely constrained several cases. But it cannot explain MD Anderson, where the binding constraint was an EHR migration; or CBA, where the bot worked and the workload rose; or Denmark, where the tools worked, the time was saved, and no earnings moved; or NYC, where the failure was documented for two years and the system stayed up. And the capability improvement is not in doubt: Stanford's 2026 AI Index records performance on SWE-bench VerifiedA standard test in which a model is given a real, previously reported bug in a real open-source codebase and has to produce a fix that passes the project's own tests. It is a hard benchmark and a narrow one — a graded exam, not a day at work — which is why saturating it says a great deal about capability and very little about value. rising "from 60% to near 100% in a single year," agent task success on OSWorldA benchmark that scores an AI system on operating a real computer the way a person does, clicking through actual applications, files and browsers to finish a task end to end. It tests the "agent" claim rather than the writing one. jumping "from 12% to ~66%," US private AI investment of $285.9 billion in 2025, and organisational adoption at 88%. Capability, money and adoption all moved hard. The value-realisation numbers did not move with them. If capability were the binding constraint, that combination should be impossible. Rejected as primary cause — high confidence

"The data isn't ready." Gartner names poor data quality first among abandonment causes and Deloitte reports nearly half of organisations citing data searchability and reusability as obstacles. Data readiness is a real and expensive constraint. But it functions more often as the acceptable public explanation than the operative one, because it is the only failure a technology leader can announce without implicating a business decision. Note that in the cases where the data was genuinely the problem — Amazon's résumé model learning from a decade of male-skewed hiring — the failure was correctly diagnosed and the project correctly killed. Real but secondary — moderate confidence

"Enterprises are too slow and risk-averse." NANDA demolishes this directly: enterprises "lead in pilot count and allocate more staff to AI-related initiatives," 90% "have seriously explored buying an AI solution," and they report the lowest pilot-to-scale conversion. The problem is that pilot volume without owned value hypotheses produces a queue, not a portfolio. Rejected — high confidence

"It's too early; the J-curve takes a decade." This is the strongest counter-argument in the set and it should not be dismissed. The J-curve is a shape, and worth having clearly in mind. A genuinely general technology first costs more than it returns, because the spending goes into rewiring the organisation around it — new processes, new skills, new plant — and the return only arrives once the rewiring is finished. Measured productivity therefore sags before it climbs, and the path drawn out looks like the letter. Electrification took roughly forty years to show in US productivity statistics; the productivity payoff from computing arrived long after the investment. Humlum and Vestergaard's own framing is compatible with it — "technological change reshapes work well before it surfaces in earnings or hours," and they document substantial task reorganisation with no earnings effect, which is precisely what an early J-curve looks like. If this is right, most of what this brief calls failure is transition cost.

Three things constrain it. First, it is unfalsifiable on any near horizon, which makes it useless as a management input — a sponsor cannot fund on it and a board cannot audit it. Second, the historical analogues involved complementary capital and organisational redesign taking decades, which is an argument for the operating discipline in this Part, not against it. Third, and decisively for the practical question: the J-curve explains why aggregate statistics lag. It does not explain why one organisation's program lands and its competitor's does not, and that variance — visible in every study here — is where operators actually work. Partly true, and not actionable — moderate confidence

The J-curve may explain the decade. It does not explain why the firm next door banked the saving and you didn't.
On the strongest counter-argument

28 · Second-order effects — what follows if this is right

If the operating diagnosis in this brief is right, the consequences run further than the case record shows.

Cheaper models make the operating gap wider, not narrower. When inference cost falls, the binding constraint moves to integration, evaluation, monitoring, escalation and legal review — costs that scale with the number of workflows touched and with headcount, not with tokens. Falling model prices therefore increase the number of technically viable pilots faster than they increase the number of organisations capable of capturing value from one. Expect the ratio of pilots to value-realising deployments to get worse as capability gets cheaper, which will be widely misread as evidence that the technology disappoints.

The advantage accrues to firms with strong management accounting, not strong data science. The scarce capability in this brief is the ability to define an account, agree a baseline, run a holdout, and make an executive give up a resource on a date. That capability lives in finance and operations. The corollary is that the firms best positioned are not the most technically sophisticated but the most operationally disciplined — which is a very different list, and one that public-market narratives about AI adoption largely ignore.

Vendor liability is migrating toward the vendor, which will reprice the software. This prediction has survived a year it could easily have failed. Workday has now been through three rounds of pleading and the "we are merely a software supplier" defence has not carried; the ADEA disparate-impact theory survived dismissal again in March 2026, and in June 2026 the court declined to accept that California's law stops at the applicant's state line when the screening happens in California. None of that is a merits finding, and it may yet become one in Workday's favour — but the cost of defending the theory is already being paid, and that is what prices software. Mobley's agency theory, if it holds, makes an AI vendor performing a core business function liable alongside its customer. Vendors will respond by narrowing warranties, raising prices in regulated functions, or withdrawing from them. Buyers should expect the "AI does the deciding" product category to become materially more expensive in hiring, credit, insurance and healthcare — and should read a vendor's willingness to accept contractual accuracy representations as the single best signal of whether their claims are real.

The measurement discipline that fixes AI value will expose everything else. A capacity conversion ledger applied honestly to an AI program is applicable to any productivity investment, and most enterprises have never had one. The first organisation to run it rigorously will discover that a decade of process-improvement programs also freed capacity that was never banked. That discovery is politically explosive and is one reason the ledger does not get built.

Delegation is rising faster than the controls around it. Anthropic's Economic Index reports that "directive" conversations — where a user hands over a complete task rather than iterating — rose from 27% to 39% of usage, and that lower-adoption regions skew toward delegation while high-adoption regions skew toward augmentation. The same report notes 40% of US employees using AI at work, up from 20% two years earlier. If delegation is growing fastest where usage is least mature, the aggregate volume of unverified AI output entering enterprises is growing faster than the verification capacity around it. That is the workslop mechanism with a trend line attached, and it argues for the verification norms in §22 becoming urgent rather than optional.

The macro data will keep saying "not yet", and that will be misread as "not ever." The Federal Reserve's own mid-2026 assessment, a FEDS Note of 17 July 2026, finds that "labor market impacts remain concentrated and have not yet broadened in the aggregate" and reads the evidence as "a buildout phase rather than the onset of broad-based displacement." That is the correct reading of the aggregate series and it is the wrong input for an operator, for the same reason the J-curve is: a national statistic averages the firms that built the conversion mechanism with the many that did not. Expect the gap between the two groups to widen while the average stays flat, and expect the flat average to be quoted at you in budget season.

A backlash is coming and it will overshoot. The gap between what was promised in 2024–25 and what the accounts show in 2026–27 is large enough to produce a governance reaction — budget consolidation, centralised approval, a freeze on new initiatives. The firms that suffer most from the freeze will be the ones that never built attribution, because they will have no defensible evidence to bring to it. Attribution is not a reporting nicety; it is the survival mechanism for a program in a downcycle.

29 · Falsifiers and dated predictions — what would change this view

Current view: the enterprise AI value gap is mostly an operating failure — value hypothesis, decision rights, attribution and capacity conversion — not a technology one.
  • A large, well-identified study (holdout or staggered-rollout design, not self-report) finds firm-level financial effects that scale with model capability rather than with operating practice — i.e. the same organisational design produces materially better value purely because the model improved. That would move the cause back to technology.
  • Aggregate value realisation rises sharply — say McKinsey's "any enterprise EBITEarnings before interest and taxes — operating profit, the money the business itself makes before the financing and the tax authorities take their turns. It is the number a chief executive is judged on, which is why "did it move EBIT" is the version of the value question that survives contact with a board. impact" figure moving from 39% to above 60%, with the sub-5% share falling — without any observable change in attribution practice or decision rights. That would suggest the constraint was capability all along.
  • Agentic deployments demonstrably close Gate 4 automatically by executing end-to-end processes with no freed-capacity conversion step, so that cost lines fall without an executive decision. Gartner's cancellation forecast and Deloitte's 11%-in-production figure both argue against this today; a reversal by 2027 would be decisive.
  • Replication of the Danish null result fails in a comparable economy with comparable data — for example, linked administrative data elsewhere showing earnings or hours effects well above the 2% bound — which would undercut the capacity-conversion mechanism at the centre of this brief. Note that the Census Bureau's 2026 microstructure study already provides one independent US corroboration, so this falsifier now requires overturning two datasets built on different methods.
Figure 25Five dated predictions, each falsifiable against a named source
PredictionHorizonSettled byConfidence
US firm-level AI adoption (Census BTOS, firm-weighted) remains below 30% — the plateau at 17–20% through mid-2026 does not break upwardBy Dec 2027Census Bureau BTOS releasesHigh
Gartner's forecast that over 40% of agentic AI projects are cancelled by end-2027 is met or exceeded, with "unclear business value" the most-cited cause rather than model capabilityBy Dec 2027Gartner follow-up research; Deloitte production-adoption trackingModerate
At least one further US state AG or federal agency action turns on an AI vendor's accuracy or automation-rate claim rather than on downstream consumer harm, extending the Pieces/Presto patternBy Dec 2027State AG and FTC/SEC enforcement announcementsModerate
McKinsey's share of organisations reporting any enterprise EBIT impact stays below 55%, and the share reporting above 5% of EBIT stays in single digitsNext two annual surveys, to end-2027McKinsey State of AI surveyModerate
At least one additional large enterprise publicly reverses an AI-justified headcount reduction after workload fails to fall, following the CBA patternBy Dec 2027Company statements, union disputes, labour tribunal filingsModerate
Evidence: Baselines are Census BTOS (17–20% Dec 2025–May 2026; 19.8% at 3 May 2026); Gartner (25 Jun 2025); Texas OAG (Sep 2024) and SEC Rel. 33-11352 (Jan 2025); McKinsey (Nov 2025, 39% / mostly under 5%); ABC News on CBA (Aug 2025).
Method: Each prediction names the specific published series that would settle it, so it can be checked without re-deriving this brief's reasoning. Confidence reflects both the strength of the base rate and the number of independent things that must hold.
Synthesis: Four of five predictions are that measured reality stays roughly where it is. That is the substance of the disagreement with the consensus, which expects each of these to break upward on capability improvement alone.

Evidence register — dated and tiered; every item opened and read for this brief

Tier 1 = primary disclosures, filings, court and regulator records, first-party post-mortems and official statistics. Tier 2 = reputable analyst, academic and established trade reporting. Tier 3 = single-source, self-reported or estimated, used only where flagged. Event dates and publication dates are distinguished where they differ.

RegisterSources, tiered and dated
SourceDateTierWhat it carries here
Zillow Group Q3 2021 results release (Exhibit 99.1 to Form 8-K)2 Nov 20211Wind-down announcement; Barton quote; $304m Q3 write-down; $421.6m Homes pre-tax loss; $240–265m expected Q4 loss; ~25% workforce reduction
Zillow Group Form 10-K, FY2021 — Item 1A risk factors and MD&AFeb 20221$407.9m FY21 inventory write-downs; $71.2m impairment and restructuring; the "submarket nuances" pricing-model risk factor; the board's stated wind-down rationale
SEC Order, In re Presto Automation Inc., Rel. 33-1135214 Jan 20251~70% human-intervention rate; reliance on off-site agents in the Philippines and India; undisclosed third-party technology; cease-and-desist, no civil penalty
SEC press release 2024-36, first AI-washing actions18 Mar 20241Delphia $225,000 and Global Predictions $175,000 penalties; Gensler quote
SEC Order, In re Delphia (USA) Inc., Rel. IA-657318 Mar 20241Aug 2019–Aug 2023 misstatements; "no such data was being used"; continuation of false statements after a 2021 corrective undertaking
Moffatt v. Air Canada, 2024 BCCRT 149 (full reasons)Event Nov 2022; decision 14 Feb 20241"A remarkable submission"; responsibility for all website information; $650.88 damages plus $36.14 interest and $125 fees; failure to produce the tariff
FTC press release, Rite Aid facial-recognition orderConduct 2012–2020; order 19 Dec 20231Five-year ban; false flags disproportionately affecting people of colour; obligation to discontinue uncontrollable automated systems
FTC, Operation AI Comply25 Sep 20241DoNotPay "robot lawyer" complaint — no testing against a human-lawyer standard, no attorneys retained; five actions on deceptive AI claims
EEOC press release, iTutorGroup consent decree11 Sep 20231$365,000; software programmed to auto-reject women 55+ and men 60+; 200+ applicants; five years of monitoring
Texas Attorney General, Pieces Technologies settlementSep 20241"<1 per 100,000" severe-hallucination claim found likely inaccurate; four Texas hospitals; five-year metric-methodology disclosure obligation
NYC Comptroller, audit of OTI's MyCity system20251$100m+ over four years; chatbot "unable to provide accurate or consistent information"; further $81m requested for FY2026; poor project and contract management
US Census Bureau, Business Trends and Outlook Survey releaseCollection to 3 May 20261National AI use 19.8%; 17–20% range Dec 2025–May 2026; 37% at 250+ employees; Information 39.7%, Finance 33.9%
US Census Bureau CES Working Paper 26-25, "The Microstructure of AI Diffusion"Reference period Nov 2025–Jan 2026118% of firms (32% employment-weighted); 57% of adopters use AI in ≤3 functions; Sales and Marketing 52%; 66% augment only; AI-related employment decreases in just 2% of firms; worker-task use shows no significant link to headcount reduction once functional integration and operational investment are controlled for
European Commission, AI Act regulatory framework and application timelineIn force 1 Aug 2024; applicable 2 Aug 20261Prohibitions and AI literacy from 2 Feb 2025; GPAI obligations from 2 Aug 2025; Annex III high-risk extended to 2 Dec 2027 and Annex I to 2 Aug 2028 under the AI Omnibus agreed Nov 2025
NYC Department of Consumer and Worker Protection, Local Law 144 on automated employment decision toolsEnforcement from 5 Jul 20231Bias audit within one year of use, published summary, and candidate notice required before an AEDT may be used
C3.ai Form 10-K, FY ended 30 Apr 2026 — risk factorsFiled 24 Jun 20261Paid "Initial Production Deployment" engagement model; disclosed risk that trial customers may not convert to ongoing subscriptions — the pilot-to-production chasm as a seller-side revenue risk
Federal Reserve FEDS Note, "Monitoring AI Adoption in the US Economy"3 Apr 20261~18% of firms adopted at year-end 2025; ~41% individual work-related GenAI use; 78% of the labour force at adopting firms; sources of divergence between surveys
Klarna press release, AI assistant month-one results27 Feb 202432.3m conversations; two-thirds of chats; "equivalent work of 700 full-time agents"; 2 vs 11 minutes; 25% fewer repeat inquiries; estimated $40m profit improvement — all self-reported and unaudited
Francisco Partners, completion of the IBM Watson Health assets acquisitionJun 20221Watson Health's data and analytics assets divested and relaunched as Merative
JPMorganChase, Chairman and CEO letter to shareholders, 2025 Annual ReportApr 20261"AI will affect virtually every function, application and process"; adoption pace framed as faster than electricity or the internet — used as an expectations datum, not a value datum
OMB Memorandum M-25-21, "Accelerating Federal Use of AI"Apr 20251Chief AI Officer requirement; minimum risk-management practices for "high-impact AI"; used as the public-sector governance anchor in §24
Stanford HAI, 2026 AI Index Report20262SWE-bench Verified from 60% to near 100% in a year; OSWorld agent success 12% → ~66% with ~1 in 3 still failing; $285.9bn US private AI investment in 2025; 88% organisational adoption — the capability counter-evidence in §27
Anthropic Economic Index, September 2025 reportSep 20253"Directive" (full-delegation) conversations rising 27% → 39%; 40% of US employees report using AI at work, up from 20% in 2023; automation-versus-augmentation skew by region — first-party platform telemetry, not a representative survey
RAND, "The Root Causes of Failure for Artificial Intelligence Projects," RR-A2680-113 Aug 2024265 interviews; five root causes led by misunderstood problem framing; five recommendations including "choose enduring problems"; the >80% figure attributed to external estimates
METR, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity"10 Jul 2025216 developers, 246 issues; 19% slower; +24% forecast, +20% self-estimate after; authors' own generalisation caveats
Humlum & Vestergaard, NBER Working Paper 33777May 2025, rev. Mar 2026225,000 workers / 7,000 workplaces; null effects ruling out >2%; 85% reallocate time savings; 93%/28%/19% adoption and saving figures in best-supported workplaces
Brynjolfsson, Li & Raymond, "Generative AI at Work," NBER WP 311612023 (QJE 2025)25,179 support agents; +14% issues per hour; +34% novice; minimal for experienced; improved sentiment and retention
Wong et al., "External Validation of a Widely Implemented Proprietary Sepsis Prediction Model," JAMA Internal Medicine21 Jun 2021 (cohort Dec 2018–Oct 2019)238,455 hospitalisations; AUC 0.63; sensitivity 33%; PPV 12%
Ostermayer et al., external validation of the Epic sepsis model in two county emergency departmentsCohort Jan–Dec 20232145,885 encounters; sensitivity 14.7%; PPV 7.6%; median lead time 0 minutes (80% CI −6h42m to +12h)
Kaiser Permanente Division of Research, ambient AI scribes at The Permanente Medical Group2025 (first year of deployment)1>2.5m uses; ~16,000 documentation hours; adoption highest in highest-burnout departments; 102 physician survey responses
BetterUp Labs and Stanford Social Media Lab, "workslop" researchSep 202521,150 US desk workers; 40% received workslop in the prior month; ~2 hours per incident; ~$186 per employee per month
MIT NANDA, "The GenAI Divide: State of AI in Business 2025"Research Jan–Jun 20253The 95% figure and its exact success definition; the 80/60/50 and 40/20/5 funnels; 67% vs 33% buy-vs-build; 90 days vs nine months; shadow-AI 40% vs 90%; the report's own stated limitations and its internal 50%/70% inconsistency
Deloitte Insights, "Agentic AI strategy," Tech Trends 20262026 (2025 Emerging Technology Trends survey)211% running agentic AI in production; 38% piloting; 14% ready to deploy; 48%/47% data searchability and reusability obstacles; partnerships twice as likely to reach deployment
Gartner forecast on GenAI abandonment, as republished by THE JournalRelease 29 Jul 20242≥30% abandoned after PoC by end-2025; the four stated causes; Sallam quotes; $5m–$20m deployment cost range
Gartner forecast on agentic AI cancellations, as republished by BigDATAwire25 Jun 20252>40% cancelled by end-2027; "agent washing"; only ~130 of thousands of agentic vendors judged real; Jan 2025 poll of 3,412 attendees
CIO Dive on S&P Global Market Intelligence's enterprise AI survey14 Mar 20252>1,000 enterprises; 42% abandoning most initiatives (17% in 2024); 46% of PoCs scrapped; cost, privacy and security as top obstacles
Silicon Canals report of McKinsey's State of AI surveySurvey mid-2025; McKinsey published 5 Nov 202521,993 respondents, 105 countries; 88% using AI; 7% fully scaled; 39% attributing any enterprise EBIT impact, most under 5%; ~6% high performers
The Register on the University of Texas System audit of the OEA projectAudit Nov 2016; report 20 Feb 20172$62m total; $39.2m to IBM; fees set just below Board-approval thresholds; $11.59m gift deficit; never in clinical use; ClinicStation/Epic integration gap
ABC News on Commonwealth Bank's reversal of AI-attributed redundanciesCuts Jul 2025; reversal 21 Aug 2025245 roles; the "error" admission; rising call volumes; FSU dispute at the Fair Work Commission; A$10.25bn FY25 cash profit
CX Dive on Klarna's return to human customer service20252Siemiatkowski on always guaranteeing a human path; the "Uber-type" flexible agent model
CIO Dive on McDonald's ending the IBM automated order-taking testJun 20242Partnership since 2021; test purpose stated as operational savings and speed; McDonald's declined to state whether it succeeded or by what metric
Restaurant Business on the AOT wind-down memo17 Jun 20242Mason Smoot's system message; more than 100 restaurants; shut off no later than 26 Jul 2024
Restaurant Business on the ArchIQ / "Archy" test2 Jun 20262Five US locations; Google partnership; ~13,600 US restaurants; dialect and format heterogeneity as the stated obstacle
TechCrunch on Taco Bell's voice-AI reassessment30 Aug 20252500+ drive-thrus; 18,000-water-cup prank order; Dane Mathews' "active conversation"; the coaching and segmentation response
Retail Dive on Just Walk Out's removal from Amazon Fresh20242Replacement by Dash Carts; Amazon's rebuttal of the human-reviewer characterisation, quoted in full in §7
MIT Technology Review on Amazon's scrapped recruiting model10 Oct 2018 (development from 2014)2Five-star ranking; penalties on "women's" and all-women colleges; project killed after loss of confidence in neutrality
MIT Technology Review on the Google Health diabetic-retinopathy deployment in Thailand27 Apr 2020 (field study 2018–19, Beede et al., CHI 2020)211 clinics; >90% lab accuracy; more than a fifth of images rejected; upload delays; the nurse's "10 patients in two hours" quote; 4.5m patients to ~200 specialists
Holland & Knight on the preliminary collective certification in Mobley v. WorkdayOrder 16 May 20252Nationwide ADEA collective; applicants 40+ since 24 Sep 2020; rejection of the "just a software provider" defence; Mobley's 100+ applications
The Markup on NYC's MyCity chatbot giving illegal advice29 Mar 2024 (launch Oct 2023)2Advice that landlords may refuse housing-voucher tenants and employers may take tips
The Markup on the chatbot's termination30 Jan 2026 (shut down Feb 2026)2"Functionally unusable"; ~$500,000 annual cost; ~$600,000 reported build cost; termination as a budget measure by a new administration
Fortune on Lattice withdrawing its "digital workers" featureAnnounced 9 Jul 2024; withdrawn 12 Jul 20242Three-day reversal; Franklin's "questions that have no clear answers yet" statement
CFO Dive on Deloitte's partial refund to the Australian governmentOct 20252A$97,000 refunded on a ~A$440,000 contract; fabricated references; Azure OpenAI use disclosed in the revised report
The Register on Builder.ai's insolvency21 May 20252>$500m raised; apps built by the team rather than by the advertised automation; prior scrutiny as Engineer.ai
CNBC on Salesforce's customer-support headcount reduction2 Sep 20252Support workforce from ~9,000 to ~5,000; Benioff's "I need less heads" framing
BBC News on the DPD chatbot failureJan 20242Failure introduced by a system update to a chatbot run "successfully for a number of years"; 800,000 views in 24 hours; component disabled
NPR on NEDA taking down the Tessa chatbot8 Jun 20232Chatbot dispensing dieting and calorie-counting advice to an eating-disorder population; withdrawal
Presto Automation Form 10-K, FY ended 30 Jun 2023Filed 11 Oct 20231Company's own claim of "approximately 95% non-intervention rate at certain locations"; Hi Auto subcontract disclosed — the filing against which the SEC order should be read
European Commission, "AI Omnibus enters into force"In force 27 Jul 2026 (OJ 24 Jul 2026)1Annex III standalone high-risk obligations deferred to 2 Dec 2027 and Annex I embedded systems to 2 Aug 2028; prohibitions, AI-literacy, GPAI and Article 50 transparency duties unchanged
Mobley v. Workday, order on the Second Amended Complaint (N.D. Cal., Dkt 267)6 Mar 20261FEHA counts dismissed with leave to amend for want of a California nexus; ADEA disparate-impact claim survives; AARP granted leave to file amicus
Civil Rights Litigation Clearinghouse docket, Mobley v. Workday, 3:23-cv-00770Docket to 20261Case chronology: Jul 2024 agency ruling; 16 May 2025 preliminary certification; Jul 2025 HiredScore scope order; Jan and Mar 2026 amendments and orders
Federal Reserve Bank of Atlanta Working Paper 2026-4, "Artificial Intelligence, Productivity, and the Workforce"Mar 20261~750 corporate financial executives; reported labour-productivity growth +2.4pp (2025) and +3.3pp expected (2026) against +1.0pp and +1.8pp implied; aggregate employment expected to fall <0.4% in 2026; routine clerical share down >2pp over three years
Federal Reserve FEDS Note, "The AI Buildout and the Economy"17 Jul 20261"Labor market impacts remain concentrated and have not yet broadened in the aggregate"; the evidence read as "a buildout phase rather than the onset of broad-based displacement"
"Multicenter Prospective Validation of an Updated Proprietary Sepsis Prediction Model," JAMA Network Open27 Feb 20262Epic Sepsis Model v2 across 4 US health systems and 227,091 encounters; AUROC 0.82–0.92; PPV 0.13–0.26 at 60% sensitivity; median lead time 1.9–10.3 h; authors recommend local validation before deployment
METR, "We are Changing our Developer Productivity Experiment Design"24 Feb 20262The follow-on experiment "gives us an unreliable signal"; 30–50% of developers declined to submit tasks without AI; authors' own view that developers are likely more sped up in early 2026
Duane Morris analysis of the 22 Jun 2026 order on the Third Amended Complaint in Mobley24 Jun 20262Granted in part and denied in part; extraterritoriality argument rejected — wrongful conduct within California is not "extraterritorial" regardless of applicants' locations
STAT on internal IBM documents concerning Watson for Oncology25 Jul 2018 (decks Jun–Jul 2017)2"Multiple examples of unsafe and incorrect treatment recommendations"; customers describing output as "often inaccurate"; training on a small number of synthetic rather than real patient cases
Beede et al., "A Human-Centered Evaluation of a Deep Learning System Deployed in Clinics for the Detection of Diabetic Retinopathy," CHI 20202020 (field study 2018–19)2Eleven clinics in Thailand; the study's own framing of "tensions between the model's thresholds for data quality, and the quality of data that arise from an imperfect, resource-constrained environment"
Evidence: 67 distinct sources, each opened and read directly for this brief. Several filings are cited for more than one distinct section (Zillow's Q3 release and its FY2021 10-K risk factors and MD&A are separate documents; the SEC's Presto order and its Delphia order are separate orders; Presto's own Form 10-K is a separate document from the order about it). Every URL cited in this brief was re-opened and confirmed live at the evidence cutoff.
Method: Tier assignment follows the definitions above. MIT NANDA and the Klarna release are tiered 3 despite being first-party documents because both rest wholly on self-report — NANDA on interview recall, Klarna on unaudited internal estimates — and both are treated in the text as directional rather than precise.
Synthesis: The load-bearing conclusions of this brief rest on Tier 1 and Tier 2 sources. Every use of a Tier 3 figure is accompanied in the text by its definitional limits.

Not verified / not load-bearing — stated plainly so no conclusion rests on it

Sources that could not be retrieved this session, and how that was handled. Gartner's own website, McKinsey's, BCG's, the ACM Digital Library, JAMA Network and the Boston Consulting Group's January 2026 AI Radar release all refused automated retrieval. Where a Gartner figure is used, it is cited to a trade publication that reproduced the release verbatim (THE Journal for the July 2024 forecast, BigDATAwire for the June 2025 agentic forecast), and labelled as such. McKinsey's survey figures are cited to a trade report of the survey, not to McKinsey, and the confidence attached to them is correspondingly lower. Beede et al.'s CHI 2020 paper is unreachable through the ACM Digital Library, but the authors' own publication page is not, and it is cited directly in §6 alongside MIT Technology Review's contemporaneous account, which quotes the study and its participants. Three further primaries refused automated retrieval and are named rather than quietly replaced: Reuters' original report of the Amazon recruiting model (401), which MIT Technology Review's account stands in for; CanLII and the BC tribunal's own site for Moffatt (403), for which a full PDF of the reasons is cited instead; and the Texas Attorney General's release (intermittently 402/200), which was reachable and is cited. BCG's 2026 figures were not obtainable and no BCG number appears in this brief.

Numbers circulating widely that this brief declines to use. Several figures in current circulation could not be traced to a retrievable primary and are therefore excluded: a "74% of enterprises have rolled back a customer-facing AI agent" statistic attributed to a vendor report; an "89% agent failure rate" attributed to Deloitte, which appears to be a misreading of Deloitte's 11%-in-production figure (the complement of "in production" is not "failed"); and various 2026 "AI ROI" percentages attributed to composite surveys whose underlying instruments are not published. Gartner has since published a figure of at least 50% of GenAI projects abandoned after proof of concept by the end of 2025 — an outturn rather than the 2024 forecast, and one that would, if citable, confirm this brief's own reading. It is not used, for a narrow reason worth stating precisely: every Gartner surface carrying it refuses automated retrieval, so the figure could not be read in a Gartner document this session. Only the July 2024 ≥30% forecast, verified through a verbatim trade republication, appears here. Readers should treat this brief's abandonment evidence as conservative on that account.

Case details deliberately not asserted. Widely repeated specifics of the MD Anderson contract — an original six-month, $2.4 million scope, twelve amendments, $39.2 million paid to IBM and roughly $21–23 million to PwC — appear in secondary accounts that could not be opened this session (The Cancer Letter is paywalled; the JNCI article and Medscape refused automated retrieval). The audit itself — since removed from utsystem.edu — has now been retrieved from the Internet Archive's capture and is cited directly in §1. It confirms the contract history: a six-month, $2.4 million original IBM agreement extended twelve times to $39.2 million in total fees, plus approximately $21.2 million of OEA-related PwC fees. Every audit quotation in §1 was checked against the retrieved text verbatim; the $62 million combined total is asserted as before. Epic's subsequent overhaul of the model is verified and is now treated in §6 on its own evidence: the updated version was prospectively validated across four US health systems and 227,091 encounters and published in February 2026. The argument in §6 concerns version one as externally validated in 2019 and 2023 and as deployed for roughly six years, and it is not weakened by the improvement — but the improvement is real and is stated rather than omitted. What remains unverified is any claim about how many hospitals have moved to the updated model or retrained it locally; no number is asserted. Amazon's characterisation of the human role in Just Walk Out is disputed by Amazon and both positions are given in §7; no claim about the true intervention rate is made. The Klarna "$40 million profit improvement" and "700 full-time agents" figures are Klarna's own estimates, never audited or restated in a financial filing, and no conclusion here rests on them. McDonald's never disclosed the success metric for the IBM test, so no inference is drawn about whether it met one.

Named cases versus archetypes. Every case treated at length in Part I is a named, publicly documented programme with a dated source: MD Anderson, Watson for Oncology beyond it, Google Health Thailand, the Epic Sepsis Model, Air Canada, DPD, NEDA, Commonwealth Bank, Klarna, Taco Bell, McDonald's, Just Walk Out, Presto, Zillow, Amazon's recruiting model, iTutorGroup, Workday, Rite Aid, Pieces, Builder.ai, Delphia, Global Predictions, DoNotPay, NYC MyCity, Lattice, Deloitte Australia, Salesforce and The Permanente Medical Group. Figure 11a contains no archetype. Exactly one entry in this brief is a composite rather than a company: failure mode 2 in Figure 13, "value not ownable," whose defining case is labelled in the figure as a representative archetype, not a specific company — commodity chat deflection and generic productivity assistants. It is drawn from the documented pattern across the cases above and no company is implied by it. Every other row in that figure names a real programme.

Under-evidenced areas. Manufacturing and industrials are materially under-represented in the public case record used here; the sector row in Figure 22 is flagged rather than filled, and the corresponding claims are held at low confidence. Public-sector evidence outside New York City and the Australian Deloitte engagement was not developed — Australia's Robodebt Royal Commission report could not be retrieved this session and is not cited. No European or Asian enterprise case is developed at length; the Danish labour-market study is the only non-Anglophone primary. Readers should treat the geographic generality of the taxonomy as an inference, not a demonstration.

Where this brief is most likely to be wrong. The capacity-conversion mechanism rests substantially on one very good study of one small, high-trust, high-union-density labour market. Denmark's institutions make headcount conversion harder than in the United States, which could mean the 85% reallocation figure is an upper bound rather than a base rate. The Census Bureau's 2026 microstructure paper is an independent US corroboration by a different method — cross-sectional regression on a probability sample rather than difference-in-differences on linked payroll — and it points the same way, which raises confidence materially. It does not eliminate the risk: both are observational, neither identifies a causal effect of an operating change on headcount, and a well-identified US study finding otherwise would require the weight placed on Gate 4 to fall.

ClosingThe models will keep getting better and it will not be enough, because the gate that decides the money was never a model gate — and nobody has been made responsible for it.

This is an internal educational and analytical reference on why enterprise AI pilots fail to deliver value. It is not investment advice, not legal advice, and not an endorsement or criticism of any vendor, employer or public body beyond what the cited public record supports. Litigation described here that has not reached judgment — including Mobley v. Workday — is at a procedural stage only, and no finding of liability is implied. Evidence cutoff: 26 August 2026.

Argue with this brief.

Every brief is built to be pushed on — corrections, counter-evidence and questions all land in the live thread, and they sharpen the next revision.

Discuss on X →

Read next

Primer
Understanding Enterprise AI Risk — The Adoption Problems Hiding Behind the Hype
Primer
The Political Economy of Latin America
The Primer Desk.
Powered by Atlas Intelligence Research (AIR)
Independent research. Sourced to primary documents, published only after passing an internal verification gate. Nothing here is investment advice.
@primerdesk on X →