Who's accountable when your AI model is wrong?
When your AI is wrong, "the computer did it" won't save you. Name the accountable person before you need them, not after.
Good morning everyone,
In 2022, a grieving passenger named Jake Moffatt asked Air Canada’s chatbot about bereavement fares. The bot invented a policy; Moffatt booked on it; the airline refused to honor it. In front of the tribunal, Air Canada argued that its own chatbot was, in effect, a separate legal entity responsible for its own actions. The tribunal member called that “a remarkable submission”, ordered the airline to pay CAD $812.02, and handed every AI team a sentence worth printing out: the company is responsible for all the information on its website, chatbot included. In this article we are going to deep dive into the question behind that case: when your AI is wrong, who answers? And more usefully: how do you design that answer before the incident, instead of discovering it in court? Let’s go.
Why “the model made a mistake” fails everywhere
Picture a general contractor handing over the keys to a finished house. If the wiring is dangerous because the electrician they subcontracted botched it, the homeowner does not go chasing the electrician; they hold the contractor who signed off on the build and handed it over. The contractor can pursue that subcontractor afterwards, but their accountability to the homeowner is total, full stop. Courts on both sides of the Atlantic are treating AI exactly this way. Air Canada owned its chatbot’s words. And Europe has now said the same thing out loud: in May 2026 the Higher Regional Court of Hamm ruled that a German clinic whose chatbot had invented professional credentials for its directors could not hide behind the bot, because a chatbot is “part of the company’s business organization”, not a third party, so its statements are the company’s own, even when the code and the training data were perfectly sound. In Mobley v. Workday, a US federal court let discrimination claims proceed against the vendor as an “agent” of employers, meaning vendor and deployer can be on the hook at the same time. And Europe’s top court ruled in the SCHUFA case (SCHUFA is Germany’s main credit bureau, whose scores banks routinely follow) that a score a human merely rubber-stamps counts as an automated decision under GDPR, with all the obligations that triggers. The supplier contract is your recourse. It is not your defense.
The organizational version of the problem is older than the technology. In 1996, Helen Nissenbaum named it the “problem of many hands”: data scientists, vendors, product managers and executives each contribute a piece, so when harm occurs, everyone points sideways. Add the machine as a scapegoat (”the computer did it”) and you do not get transferred accountability. You get an accountability vacuum.
ℹ️ Small glossary
Provider / deployer = EU AI Act roles: the one who builds or places the system on the market vs the one who uses it. Different legal duties attach to each.
Human in the loop (HITL) = a person positioned to review or approve AI outputs before they take effect.
Automation bias = the human tendency to rubber-stamp whatever the system suggests.
Kill switch = pre-agreed authority to degrade or stop an AI system without escalation.
Blameless post-mortem = incident review focused on fixing the system, not naming a culprit.
What it costs when nobody owns it
Two case studies mark the ends of the spectrum. On the commercial end, Zillow’s home-buying algorithm systematically overpaid while humans overrode conservative estimates to hit growth targets. Nobody owned model health as a full-time answerable role, and the unit lost roughly $881M in 2021 before shutting down, taking about 2,000 jobs with it. The algorithm did not get fired; employees did.
On the human end, the Dutch tax authority’s fraud-scoring system helped wrongly accuse roughly 26,000 parents of benefits fraud. Families were ruined, more than a thousand children entered foster care, the entire cabinet resigned in January 2021, and the data-protection authority issued a then-record €3.7M fine. Notice the pattern in both: accountability arrived, years late, at maximum blast radius. The choice is never accountability or no accountability. It is accountability by design, or accountability by explosion.
The twist: regulators already wrote your org chart
Here is what changed. The EU AI Act does not just regulate models; it assigns roles. Providers carry documentation, testing and logging duties (Article 16); deployers must use systems as intended, monitor them, and assign oversight to people “who have the necessary competence, training and authority” (Article 26); and Article 25 hides a trap: substantially modify or rebrand a high-risk system and congratulations, you are now the provider. Meanwhile the EU quietly withdrew its AI Liability Directive in 2025, which means civil liability falls back on 27 national regimes plus the revised Product Liability Directive. Less harmonization is more reason to design your own accountability, not less: you cannot rely on a single predictable external rule to sort it out for you.
The good news: you do not need a governance office. NIST’s AI Risk Management Framework makes Govern (roles, responsibilities, lines of accountability) the foundation, and ISO/IEC 42001 asks top management to assign AI responsibilities; both are management-system requirements, not headcount requirements. For a mid-size team, the whole spirit fits on one page per system.
The framework: a one-page accountability map
Accountability decided after the incident is like arguing over who should have driven, after the crash. The point is to name the designated driver while everyone is still sober. Here is the map; one name per cell, never a committee:
Then tier the ceremony by risk, so governance does not smother adoption. Not every AI system deserves the full apparatus. The trick is to ask two questions about a wrong output, how consequential is it, and how easily undone, and let the answers drop the system into one of three tiers:
Tier 1: it touches customers or makes real decisions. A chatbot that quotes prices or policy, a CV screener, a credit or insurance-claim decision. A wrong output reaches the outside world and is hard to walk back. ⇒ Full accountability map, a written incident runbook, and a quarterly review.
Tier 2: it advises a human who then decides. An internal tool that flags which invoices look risky or which sales leads look hot, for a person to act on. A wrong output can still be caught before it bites, but only if that person is doing real oversight (see the next section). ⇒ A named owner, agreed thresholds, and monitoring.
Tier 3: it drafts and assists, low stakes. A tool that suggests email copy or summarizes a document, always with a human editing before anything leaves the building. ⇒ A named owner and a feedback channel. That is genuinely enough; do not bury it in ceremony.
When you are unsure, tier up one level for the first few months, then relax once the system has earned your trust with evidence.
And classify every system in provider/deployer vocabulary even outside the EU: did we build it (provider-like duties) or deploy someone else’s (fit-for-purpose checks, monitoring, a vendor contract with indemnities)? If you fine-tuned or white-labeled a vendor model, re-read the Article 25 trap above.
Human in the loop: real oversight, or just theater?
The seductive wrong answer to all of this is “don’t worry, we put a human in the loop.” It sounds responsible. Very often it is theater. Ben Green’s review of oversight policies is blunt: across study after study, “people are unable to perform the desired oversight functions”, and an oversight rule on paper mostly ends up legitimizing a shaky system while letting everyone involved shirk the blame.
Madeleine Clare Elish gave the trap its perfect name: the moral crumple zone. In a car crash, the crumple zone is the part engineered to fold and absorb the impact so the passengers survive. In an AI system, it is the human positioned to absorb the blame when the machine fails, even though they never truly had control. The starkest example is the 2018 Uber self-driving car that struck and killed a pedestrian in Tempe, Arizona. The safety driver, sitting behind the wheel while the software did the driving, was the one criminally charged; the company was not. One underpaid human became the crumple zone for an entire system’s failure.
So what separates real oversight from theater? A genuine overseer needs four things, and missing any single one quietly turns them into a crumple zone:
Competence ⇒ they actually understand what the system does and how it tends to fail.
Time ⇒ they can review before the decision takes effect, not rubber-stamp four hundred an hour.
Information ⇒ they can see why the model suggested what it did, not just the bare output.
Authority ⇒ they are genuinely allowed, and expected, to overrule it, without begging a manager first.
Contrast two pictures. An airline captain does not hand-fly most of the flight, yet no crash investigation ever ends with “the autopilot did it”, because the captain is trained, drilled on exactly when to take over, and equipped with alarms and a clear handover procedure. That is oversight with all four ingredients present. A bank clerk clicking “approve” on whatever a scoring model returns, with no time, no explanation and no real permission to say no, is a crumple zone with a login. The fastest test on your own systems: look at the override rate. If your human “in the loop” overrules the model 0.1% of the time, they are not overseeing anything. That is automation bias with an audit trail, and courts are increasingly seeing straight through it.
Honest nuance
Three caveats to keep this framework honest. First, over-assigning accountability kills adoption: five sign-offs per model means nobody ships, and shadow AI (zero accountability) proliferates instead. That is what the tiers are for. Second, causality is genuinely shared: in Zillow’s case the model erred and humans overrode toward growth targets and leadership set the incentives. The map does not pretend one person caused the failure; it guarantees one person must respond per decision. Causal contribution is always plural; answerability must never be. Third, the new AI insurance wave (Lloyd’s syndicates covering chatbot errors, AIUC certifying and insuring agents) transfers financial risk, not answerability; insurers actually push the other way, conditioning coverage on named owners and monitoring evidence. Insurance is the seatbelt, not the driver. And pair the map with blameless post-mortems, or people will game metrics and hide incidents; the owner answers for the response, not for the model’s every error.
Is this a big thing?
Yes. The recap:
“The model did it” has lost in court ⇒ Air Canada owns its chatbot’s words, Workday can be liable as the employer’s agent, and rubber-stamped scores count as automated decisions.
Accountability arrives either by design or by explosion ⇒ because no one owned the failure ahead of time, Zillow’s losses reached about $881M and roughly 2,000 jobs, and the Dutch benefits scandal collapsed an entire national government. The bill always arrives, and it is far larger late than early.
Regulators now assign roles, not just rules ⇒ provider vs deployer duties land Aug 2, 2026 in the EU, and modifying a vendor system can silently make you the provider.
A human in the loop is not a defense ⇒ without authority, training, time and real override rates, your reviewer is a moral crumple zone.
One page per system is enough ⇒ Before/During/After, one name per cell, ceremony tiered by risk. Insurers and enterprise procurement will ask for exactly this map.
The one-line takeaway: the question “who’s accountable?” only has a good answer if it was written down before it was needed. Name the designated driver while everyone is still sober.
What do you think?
Try it this week: pick your riskiest AI system and fill the nine cells of the map in one meeting. Which cell had no obvious name?
Sources
Cases & regulation
ABA Business Law Today (2024). BC Tribunal confirms companies remain liable for AI chatbot information (Moffatt v. Air Canada). https://www.americanbar.org/groups/business_law/resources/business-law-today/2024-february/bc-tribunal-confirms-companies-remain-liable-information-provided-ai-chatbot/
Holland & Knight (2025). Federal court allows collective action over alleged AI hiring bias (Mobley v. Workday). https://www.hklaw.com/en/insights/publications/2025/05/federal-court-allows-collective-action-lawsuit-over-alleged
DLA Piper (2026). German court addresses chatbot liability (Higher Regional Court of Hamm, judgment 12 May 2026, I-4 UKl 3/25). https://www.dlapiper.com/en-us/insights/publications/2026/06/german-court-addresses-liability
EU AI Act, Articles 16 / 25 / 26 / 14. https://artificialintelligenceact.eu/article/26/
IAPP (2025). European Commission withdraws AI Liability Directive. https://iapp.org/news/a/european-commission-withdraws-ai-liability-directive-from-consideration
Clifford Chance (2022). Dutch government fraud scandal leads to record GDPR fine. https://www.cliffordchance.com/insights/resources/blogs/talking-tech/en/articles/2022/04/dutch-government-fraud-scandal-leads-to-record-breaking-gdpr-fin.html
Research
Nissenbaum (1996). Accountability in a Computerized Society. https://nissenbaum.tech.cornell.edu/papers/accountability.pdf
Green (2022). The Flaws of Policies Requiring Human Oversight of Government Algorithms. https://arxiv.org/abs/2109.05067
Elish (2019). Moral Crumple Zones. https://estsjournal.org/index.php/ests/article/view/260
NIST. AI RMF ↔ ISO/IEC 42001 crosswalk. https://airc.nist.gov/docs/NIST_AI_RMF_to_ISO_IEC_42001_Crosswalk.pdf
Industry
Full Stack Economics (2021). How Zillow’s homebuying scheme lost $881 million.
The Decoder (2025). Lloyd’s insurers launch first AI chatbot error policies. https://the-decoder.com/lloyds-insurers-launch-first-ai-chatbot-error-policies/
Fortune (2025). AIUC emerges from stealth with $15M to insure AI agents. https://fortune.com/2025/07/23/ai-agent-insurance-startup-aiuc-stealth-15-million-seed-nat-friedman/


