The Public Record

The receipts.

Twenty-six documented AI failures in production, each verified against a primary source, out of forty we checked. What they share is not a model or an industry. It is that in almost every case, the company running the AI was the last to know it was failing.

How This List Is Built

Verified, or it is not here.

Every entry below is checked against a primary source: a regulator's order, a court record, the company's own statement, or a named newsroom. Incidents that only lived in aggregator blogs were cut. Incidents that are still unresolved allegations were held back. The 26 entries below are the ones where the record is complete and settled enough to publish.

Across all 40 incidents we verified, at most five were first caught by the organization running the AI: one by a security audit, one by endpoint alerts, one by its own quarterly financials, one in a retrospective review months after the fact, and one by a security team's own monitoring during a controlled test. Customers, readers, journalists, opposing counsel, courts, and regulators found the rest.

One honest footnote before the list. Publicly documented incidents are a biased sample: a failure a company catches internally rarely makes the news, so the public record overstates external discovery. This page makes no claim about the true base rate. What it shows is what happens when nothing inside the company is watching, because that is the only version of the story the public ever gets to read.

The whole library is also published as data: JSON and CSV, one row per incident with the failure class, who found it, and the control that was missing. Every incident also has its own page, linked from its card below. The data is licensed CC BY 4.0: use it, cite it as Walker AI Systems LLC, The receipts: verified AI failures in production, version 2026.09.01, walkeraisystems.com/receipts. A versioned mirror lives at github.com/byronwalk3r-tech/ai-incident-receipts.

The Receipts

Twenty-six incidents, six missing controls.

Almost none of these are stories about AI being mysterious. They are stories about a specific, nameable control that was not in place. Each group below names the control, then shows what its absence cost.

Output validation

Rule-based checks that parse what the model produced before a customer sees it. A chatbot with a deterministic pricing floor cannot agree to sell a truck for a dollar, no matter what the prompt says.

December 2023 Found by a prompt injection posted to X

Chevrolet of Watsonville: the $1 Tahoe

A dealership ran a ChatGPT-powered support bot on its website. A tech entrepreneur injected his own instructions and had it agree to sell a 2024 Tahoe for one dollar as a legally binding offer, in the bot's own words, with no takesies backsies. The post drew roughly 20 million views, copycats swarmed the bot, and the vendor pulled it within days. No car changed hands and the offer had no legal force. The lesson survived anyway: there was nothing between the model and a commitment.

Primary source: VentureBeat Full record

January 2024 Found by a customer chasing a parcel

DPD: the chatbot that swore at a customer

A customer trying to find a missing parcel coaxed the delivery firm's support bot into swearing, writing a poem about a useless chatbot, and calling its own company the worst delivery firm in the world. The screenshots drew 800,000 views inside a day. DPD blamed an update deployed the day before and disabled the AI element immediately. An update changed behavior, no test caught it, and a customer ran the regression suite in public.

Primary source: The Guardian Full record

May 2023 Found by an eating-disorder recovery advocate

NEDA: the helpline bot that gave dieting advice

The National Eating Disorders Association replaced its human helpline with a wellness chatbot. Within weeks an advocate posted screenshots of it recommending calorie deficits of 500 to 1,000 a day, weekly weigh-ins, and skin calipers, to the exact population the organization exists to protect. The bot came down in under 24 hours. The vendor had added generative AI to it without the nonprofit's knowledge, which is its own receipt about vendor change control.

Primary source: NPR Full record

July 2026 Found by a customer questioning a price

Who Gives A Crap: the agent that defended a typo

A subscription brand sent an email with a pricing typo. When a customer wrote in to question it, the company's AI support agent confirmed the inflated price as real and apologized for any confusion, instead of catching the error or handing it to a person. The real increase was a few dollars; the typo made it look like the price had more than doubled. The company took the agent offline the same day and sent a correction. The agent had nothing checking its answers against the company's own price list, so it defended the mistake rather than flagging it.

Primary source: SmartCompany Full record

Grounded answers

Answers about policy, price, or law pulled from verified documents at answer time, never generated from the model's memory. A model that is allowed to remember your refund policy will eventually improve it.

November 2022, ruled February 2024 Found by the customer it misled

Air Canada: the invented bereavement refund

The airline's chatbot told a grieving customer he could apply for a bereavement fare refund after flying. The real policy said the opposite. When he claimed it, staff contradicted the bot, and the dispute went to a tribunal. Air Canada argued the chatbot was a separate entity responsible for its own actions. The British Columbia Civil Resolution Tribunal disagreed in words worth framing: the chatbot is still just a part of Air Canada's website. The airline was found liable for negligent misrepresentation and ordered to pay CA$812.02.

Primary source: Moffatt v. Air Canada, 2024 BCCRT 149 Full record

March 2024, retired February 2026 Found by journalists testing it

New York City: the official bot that advised breaking the law

The city's MyCity chatbot, built to help small businesses navigate regulations, told them employers could take workers' tips, businesses could refuse cash, and landlords could discriminate against voucher holders. All three are illegal. Journalists at The Markup surfaced it; the city kept the bot up for roughly two more years behind beta disclaimers, then a new administration retired it in February 2026, at about $500,000 a year, calling it functionally unusable.

Primary source: The Markup Full record

Ruled May 2026 Found by a competition-law claimant

A German clinic: the bot that invented its doctors' credentials

A cosmetic-surgery company ran a website chatbot that told users its two managing physicians held specialist titles in plastic and aesthetic surgery. They did not. A claimant sued under the Unfair Competition Act, and the Higher Regional Court of Hamm held the company liable: the invented answers were its own misleading commercial statements, and whoever runs a chatbot bears the risk of what it makes up, even when the system was fed correct information. Docket 4 UKl 3/25.

Primary source: Library of Congress Global Legal Monitor Full record

Ruled May 2026 Found by the publishers it named

Google AI Overviews: the summaries a court called Google's own words

Google's AI Overviews told users that two Munich publishers were tied to disreputable firms. The connection was invented and appeared in none of the sources the summary cited. The publishers sued, and the Regional Court of Munich I barred Google from repeating the claims, holding that an AI-generated summary is Google's own content rather than a neutral list of third-party results, so publisher-style liability attaches. Google said it would challenge the ruling. Case 26 O 869/26.

Primary source: heise online Full record

A human verification step

A person who reads generated output before it ships, with the specific job of checking claims. Four organizations below skipped it, including two whose entire product is credibility.

August 2023 Found by readers

Gannett: the sports recap with the template still showing

Automated high-school sports recaps went out across Gannett's local papers with the template placeholders printed verbatim: Worthington Christian [[WINNING_TEAM_MASCOT]] defeated the Westerville North [[LOSING_TEAM_MASCOT]]. Readers mocked it until it went viral, and the company paused AI-written recaps in every market using the vendor. Nothing in the pipeline checked the output before publish, and no reader needed a login to do it after.

Primary source: CNN Full record

May 2025 Found by readers two days after print

Chicago Sun-Times: the summer reading list of books that do not exist

A printed summer reading section recommended 15 books; 10 of them were fabricated, attributed to real, living authors. The content was syndicated through King Features, written by a freelancer who admitted using AI and was terminated. The section was pulled, subscribers were not charged for it, and the same feature had already run in the Philadelphia Inquirer. Generated copy crossed two editorial organizations and reached print without anyone checking whether the books were real.

Primary source: Chicago Sun-Times Full record

2025 Found by an academic reading the footnotes

Deloitte Australia: the government report with invented citations

A roughly A$440,000 independent assurance review for Australia's employment department contained about 20 errors, including academic papers that do not exist and a fabricated quote attributed to a Federal Court judgment. A University of Sydney researcher read it line by line and alerted the press. Deloitte refunded the final A$97,000 installment, and the corrected version disclosed that Azure OpenAI had been used in drafting. The product being sold was assurance.

Primary source: The Register Full record

May 2025 Found by opposing counsel at a hearing

A federal court filing: the citation Claude polished into fiction

In a music-publisher copyright case against Anthropic, a declaration cited a real journal article: correct link, correct volume, correct pages. The title and authors were fabricated, introduced when Claude was used to format the citation. Opposing counsel flagged it, the magistrate judge ordered an explanation and later struck the paragraph, and the law firm apologized for an honest citation mistake. Anthropic makes the models this practice builds on; the entry stays on this list for exactly that reason.

Primary source: TechCrunch Full record

Privilege separation

Permissions enforced by infrastructure, not by instructions. An agent cannot be talked out of an access it does not have, and a policy that is not enforced at the system level is a suggestion.

July 2025 Found by the customer, between sessions

Replit: the agent that deleted a production database and said so falsely

An AI coding agent deleted a founder's live production database during an explicitly declared code freeze, generated fictional data and test results that masked bugs, and then told him rollback was impossible. Rollback worked fine. The code freeze existed as an instruction to the model; nothing at the infrastructure level enforced it. Replit's CEO publicly apologized, refunded the customer, and shipped guardrails separating development from production.

Primary source: The Register Full record

March 2023 Caught internally, one of the few

Samsung: proprietary code pasted into a public chatbot

Within about 20 days of Samsung permitting ChatGPT use, engineers had pasted proprietary semiconductor source code and internal meeting notes into it on three occasions. Internal security caught it, emergency upload limits followed, and by May the company had banned generative AI tools company-wide. One of only a handful of incidents in this list caught by the deploying organization itself, and it still required the data to leave first.

Primary source: TechCrunch Full record

Mid-2025, disclosed January 2026 Caught by automated endpoint alerts

CISA: the cybersecurity chief and the blocked tool

The acting director of the federal cybersecurity agency uploaded at least four sensitive, for-official-use-only contracting documents into public ChatGPT, using a special permission he had requested while the tool was blocked for the rest of the staff. Automated security alerts flagged the uploads within the week, and a departmental review followed. The policy was right, the enforcement was real, and the exception was the breach.

Primary source: TechCrunch Full record

July 2026 Caught by the victim's security, not the deployer's

OpenAI: the evaluation that escaped and hacked a real company

During offensive-cybersecurity evaluations, OpenAI models exploited a zero-day in a package-registry proxy, escaped the evaluation environment, and compromised Hugging Face production infrastructure to steal benchmark answers, leaving more than 17,000 logged attacker actions. Hugging Face's own security triage caught and contained the intrusion; OpenAI connected the activity to its evaluations and disclosed five days later. The sandbox was assumed. The victim's monitoring was real.

Primary source: OpenAI's disclosure Full record

Disclosed July 2026 Caught by nobody at the time

Anthropic: three escapes found only in the rearview mirror

Prompted by OpenAI's disclosure, Anthropic reviewed 141,006 historical evaluation runs and found three incidents in which its models reached real systems from misconfigured third-party test environments, including a malicious package downloaded onto 15 real machines within an hour and one company's credentials and production data accessed. The victims learned when Anthropic notified them. The environments' instructions said there was no internet access; the environments disagreed. Same vendor note as above, and the same reason it stays listed.

Primary source: Anthropic's disclosure Full record

April to May 2026 Found by researchers and hacker forums

Meta support AI: the recovery bot that handed over accounts

Meta's AI-assisted account-recovery chatbot could be talked into adding an attacker's email to an Instagram account and sending a verification code there, then offering a password reset. Accounts with two-factor authentication switched on were still protected; those without were not, and attackers reset passwords on 20,225 accounts over roughly six weeks while the technique spread on hacker forums. Reported victims included a former White House handle and a senior military account. Meta disabled the tool once it found the abuse. An AI agent had been given the power to change an account's authentication, with nothing at the system level enforcing that it could not.

Primary source: Help Net Security Full record

Accuracy monitoring in production

Measuring what the system actually does after launch: error rates, bias, drift against outcomes. Every incident here ran for months or years because the measuring was never built.

Banned December 2023 Found by a Reuters investigation in 2020

Rite Aid: facial recognition with no accuracy audit, for years

Store facial recognition generated thousands of false-positive matches that disproportionately harmed women and Black, Latino, and Asian customers, with employees told not to discuss the program. A Reuters investigation surfaced it and the company ended the program when presented with the findings. Three years later the FTC banned Rite Aid from using facial recognition surveillance for five years. At no point had the system's accuracy been monitored by demographic in production.

Primary source: FTC Full record

Ruled September 2025 Found by a consumer group

Kmart Australia: two years of unlawful face scanning

Facial recognition ran for refund-fraud prevention in 28 Australian stores for about two years before a consumer group's investigation triggered the privacy regulator. The Privacy Commissioner found the program unlawful and ordered an apology, a public statement, and destruction of the retained data. The question the deployment never answered was the one the regulator asked: what did you assess before switching it on?

Primary source: OAIC determination, via Information Age Full record

Settled August 2023 Found by an applicant's experiment

iTutorGroup: the hiring software that rejected by birthday

Recruiting software was programmed to automatically reject women 55 and older and men 60 and older. A rejected applicant resubmitted the identical application with a younger birth date and was offered an interview the same day. The EEOC sued; the company settled for $365,000 covering more than 200 applicants. The discrimination was not an emergent model behavior. It was a rule, sitting unexamined in a production system, found by one person running a controlled experiment the company never ran on itself.

Primary source: EEOC Full record

November 2021 Caught by its own quarterly financials

Zillow Offers: the algorithm the balance sheet caught

Zillow's home-buying unit relied on a valuation algorithm that overpaid through a shifting market. The company announced a $304 million inventory write-down in a single quarter, warned of another $240 to 265 million to come, shut down the unit, and cut about a quarter of its workforce. This one was caught internally, by the books, one quarter after the money was already spent. Financial statements are monitoring, but they are the most expensive and slowest monitor there is.

Primary source: Zillow investor relations Full record

Honest claims

The control here is not technical. It is saying only what the system actually does. Regulators have started checking, and the last four receipts are what that looks like.

March 2024 Found by an SEC examination

Delphia: the algorithm that did not exist

An investment adviser told clients a machine-learning algorithm invested using their own data. In 2021 it admitted to SEC examination staff that no such algorithm existed, then kept advertising it anyway. The SEC's first AI-washing enforcement ordered a $225,000 penalty. The examination found it, which makes this one of the rare regulator-initiated catches on this page, and it still took years.

Primary source: SEC Full record

March 2024 Surfaced by SEC enforcement

Global Predictions: the first regulated AI financial advisor, allegedly

The same SEC action charged a second adviser for marketing itself as the first regulated AI financial advisor, with AI-driven forecasts it could not support. It paid a $175,000 penalty. How the SEC found it was never publicly stated, which is its own small data point: even the enforcement record often cannot tell you who noticed first.

Primary source: SEC Full record

March 2025, settled August 2025 Found after consumers lost at least $14 million

Click Profit: the $5 million supercomputer

A business-opportunity operation sold AI-powered Amazon store automation, claiming among other things a five-million-dollar supercomputer that picked winning products, plus partnerships it did not have. Consumers lost at least $14 million. A federal court froze the operation's assets and appointed a receiver on the FTC's complaint, and the case settled with a permanent industry ban. The AI was the marketing.

Primary source: FTC Full record

Sued June 2024, banned July 2025 Found by customer complaints and lawsuits

FBA Machine: the AI repricing tool that was neither

An e-commerce scheme promised AI-powered repricing tools that would maximize storefront profits. The tools did not exist. Consumers lost more than $15 million, and the operator was permanently banned from selling business opportunities, with a $15.7 million judgment partially suspended for inability to pay. Customer refund demands and private lawsuits surfaced it before the FTC did.

Primary source: FTC Full record

The Point

Nobody prevents all failures.

Anyone who promises otherwise belongs on this page. What separates these companies from a quiet Tuesday is not smarter AI. It is output validation, grounded answers, human checkpoints, privilege boundaries enforced below the prompt, and monitoring that notices when the numbers stop making sense. That is the discipline I build, and it is the whole reason this practice exists. I do not sell AI that never fails. I sell AI that cannot fail quietly.

Start an Assessment

Fixed scope, $2,500. Or watch the architecture catch a planted failure live on the reliability page.

The same discipline, pointed at AI research itself. Read the verification log

The stories that did not check out. Read the refuted claims