Never automate three things. The step where a professional's judgement is the deliverable. A process whose underlying logic is already wrong — automation scales brokenness rather than repairing it. Any output that becomes a promise the organisation must honour. Everything else turns on two questions: is being wrong recoverable inside the organisation, and does the output create a commitment somebody outside can act on?
Ross Jones — Founder, The Hopium Lab. Last modified 22 July 2026. Twenty-four shipped products across fifteen sectors, every one with an MCP server and a CLI.
Hasn't somebody already written the negative version?
Yes, and pretending otherwise is the quickest way to lose the only readers worth having. Zapier, Retool and CIO have each published a variant: don't automate one-off decisions, processes with no digital record, or the thing that defines your core promise. The advice is right and nearly identical everywhere.
Not one of those pages names a case. Abstraction is worthless in the meeting where somebody has already signed the licences. What follows is function-level: one named failure mode per function, a primary source wherever the claim is load-bearing, and a correction wherever the story everyone repeats is wrong.
What did the worst automated-decision disasters actually run on?
Mostly not AI. Miscategorising them hands a hostile reader the whole argument.
Australia's Robodebt scheme was rule-based data matching plus income averaging — no machine learning anywhere. The Royal Commission reported on 7 July 2023 that averaging annual ATO income across fortnights was inconsistent with social security legislation, and therefore unlawful. AU$746m was wrongfully recovered from roughly 381,000 people and later refunded, and around AU$1.75bn in raised debts was ultimately wiped. The purest illustration of brokenness at scale — automating a process whose underlying logic is already wrong.
The EEOC's iTutorGroup settlement is cited everywhere as the first AI hiring discrimination case. The EEOC's own release of 11 September 2023 never says AI, machine learning or algorithm. It says the company "programmed their tutor application software to automatically reject female applicants aged 55 or older" — and men aged 60 or older — a hard-coded date-of-birth filter, caught because one applicant sent two identical applications differing only in birth year. $365,000 and five years of monitoring.
Moffatt v. Air Canada, 2024 BCCRT 149 is a British Columbia Civil Resolution Tribunal decision of 14 February 2024, and the tribunal defined the thing neutrally: an automated system responding to a user's prompts and input. No finding about the underlying technology. Damages: CA$650.88.
Three of the most-cited automated-decision disasters of the last decade were never found to involve machine learning at all. The danger was never the model class. It was automating a decision nobody could contest.
Which processes should each function stop automating?
The split is the same in every function: automate the work, never the commitment.
| Function | Automate freely | Never run unattended | Named failure mode |
|---|---|---|---|
| Sales | Enrichment, routing, transcription, pipeline roll-ups | Any quote or term a customer can hold you to | Promise leakage |
| Support | Triage, order status, retrieval of documented policy | Answering where documentation is silent; granting exceptions | Promise leakage |
| Finance | Reconciliation matching, anomaly flagging, coding suggestions | Figures that enter a filing; arithmetic done in prose | AI washing |
| HR | Scheduling, job-ad drafting, policy FAQ | Screening, ranking, rejection, promotion, termination | Null compliance |
| Legal | Clause extraction, first-pass review against a real corpus | Citation generation; the advice itself | Judgement substitution |
| Ops | Bounded forecasting, routing, capacity planning | Eligibility or penalty decisions with no contest path | Brokenness at scale |
The finance row is the counter-intuitive one. The regulated risk there is not the model being wrong — it is AI washing, advertising a capability the system does not have. SEC press release 2024-36 of 18 March 2024 settled the first such actions against two registered investment advisers — $225,000 and $175,000, one marketed as the first regulated AI financial adviser. The SEC found neither had the capability advertised; both settled without admitting or denying the findings. No number that reaches a filing should be computed by a language model in prose — finance and legal sit on one rule, the model may propose, never compute.
What makes an output a promise the organisation must honour?
Anything a customer can reasonably act on. Promise leakage — an automated interface creates a commitment the organisation never authorised.
Air Canada's chatbot told a grieving passenger he could apply for bereavement fares retroactively. Company policy forbade it. Air Canada argued the chatbot was a separate legal entity responsible for its own actions; the tribunal's reply — "This is a remarkable submission" — is the only sentence anyone needs from the decision. The tribunal found negligent misrepresentation, resting on a duty of care arising from the commercial relationship between service provider and consumer. Small jurisdiction, limited precedential weight, technology-agnostic principle.
Anysphere's Cursor support bot went further in April 2025. It explained a session-management bug by asserting a one-active-session-per-subscription limit that did not exist, and users cancelled before the company could correct it.
Automation is a promise-making machine. Whatever the interface says, the organisation has said — and inventing the policy is worse than getting it wrong, because there is nothing to correct it against.
Where is the failure rate actually measurable?
In legal work, where courts write the failures down. Damien Charlotin of HEC Paris maintains a database of decisions in which a judge explicitly found, or clearly implied, reliance on AI-hallucinated material: 719 by January 2026, 1,227 by early April 2026, 1,598 as of 9 June 2026 — roughly six new cases a day across that window, and stale by the time you read it.
Judgement substitution — automating the step where a professional's judgement is the deliverable. In Mata v. Avianca (S.D.N.Y., 22 June 2023) the aggravating factor was never using ChatGPT: Judge Castel sanctioned two attorneys and their firm $5,000 jointly for conscious avoidance and false statements after the six fabricated citations were questioned. The model produced the draft; the lawyers declined to do the one thing they were being paid for.
The sharpest example of judgement substitution is a consultancy audit. A roughly A$440,000 assurance review, delivered to an Australian government department in July 2025, contained fabricated academic references and an invented quote attributed to a Federal Court judgment with the judge's name misspelled. Deloitte repaid the final contract instalment; the corrected version disclosed Azure OpenAI GPT-4o had been used. The subject of the review was the department's automated compliance system.
Running the gate below across a live process is a one-day exercise. I sell it as a fixed-scope go/no-go; the procedure is free either way.
Does regulation settle which processes to avoid?
No. Regulation tells you which processes carry paperwork, a different question entirely. EU AI Act Annex III reads like a function-by-function list: 4(a) recruitment and candidate filtering, 4(b) promotion, termination and task allocation, 5(a) eligibility for essential public benefits, 5(b) creditworthiness with an express carve-out for fraud detection, 5(c) life and health insurance pricing, 5(d) emergency call triage. Article 14 requires design permitting effective oversight by natural persons during use.
The deadline moved. Final Council approval on 29 June 2026 defers stand-alone Annex III high-risk obligations from 2 August 2026 to 2 December 2027, and embedded Annex I systems to 2 August 2028. Anyone still writing a countdown to next month has not read the Omnibus. Sixteen extra months means the commercial argument has to stand on its own — which it does. The logging mechanics sit in a separate note.
Where a rule already bites, the result is what Wright and colleagues named null compliance — the audit obligation is met on paper and produces no accountability. New York City's Local Law 144 mandates bias audits for automated employment decision tools. Their study of 391 employer observations found 18 posting the required audit (~5%) and 13 the required notice (~3%), with almost every published audit reporting no adverse impact. The New York State Comptroller concluded on 2 December 2025 that enforcement had been ineffective, finding at least 17 potential non-compliances despite penalties of up to $1,500 per violation per day.
Buying the screen from a vendor does not move the liability, either. In Mobley v. Workday (N.D. Cal.), Judge Rita Lin granted preliminary ADEA collective certification on 16 May 2025 covering applicants aged 40+ since 24 September 2020, on the theory that the vendor acts as its client-employers' agent and so falls within the statutory definition of employer. Allegations at the certification stage. Nobody has been found liable. The vendor is nonetheless inside the blast radius.
Isn't a post built on failures just sampling on the outcome?
Yes, and the honest correction changes the claim. Successful automations do not generate royal commissions. McDonald's ended its IBM drive-thru order-taking pilot in June 2024 after viral misorder videos, then launched a Google Cloud voice assistant in June 2026, now running at US sites in English and Spanish. Cite the 2024 failure as proof the process cannot be automated and an informed reader points out the company re-automated it.
Amazon's 2014 résumé-rating tool cuts the same way. Reuters reported in October 2018 that it penalised résumés containing "women's" — and that it never left trial. A tool caught and killed before deployment argues for pre-deployment testing, not against automation.
Zillow Offers is told wrong most often of all. Stanford GSB analysis attributes the $304m Q3 2021 write-down and roughly $569m total to business strategy rather than algorithmic inadequacy: a late entrant buying non-cookie-cutter homes where valuation models are weakest, then bidding above its own model's estimate to win volume. Override drift — humans override the model precisely on the cases it was built to decide — is a governance failure wearing an algorithm's coat.
The defensible claim is about the shape of failure, never its frequency.
What four questions settle it?
Recoverability, promise, process integrity and judgement — run per process step rather than per system. A gate that flags everything is worthless, so the worked examples clear one live automation and block three famous ones.
THE FOUR-QUESTION AUTOMATION GATE — v1.0
Run per step. Any RED = the model may propose, a human decides.
Q1. RECOVERABILITY
If this step is wrong 1 time in 100, can the organisation
undo it within its own control, at its own cost, before
anyone outside notices?
[ ] GREEN reversible internally
[ ] RED lands on a customer, a filing, or a named person
Q2. PROMISE
Does the output create a commitment a third party can
reasonably act on?
[ ] GREEN informational, no reliance
[ ] RED price, entitlement, policy, eligibility, deadline
Q3. PROCESS INTEGRITY
Is the current process written down, lawful, and contestable
by the person it affects — TODAY, before automation?
[ ] GREEN documented + contest path exists
[ ] RED lives in someone's head, or has no appeal route
Q4. JUDGEMENT
Is the judgement in this step the thing being paid for?
[ ] GREEN judgement sits upstream or downstream
[ ] RED the judgement IS the deliverable
NAMED FAILURE MODES (identical wording wherever used)
Promise leakage an automated interface creates a commitment
the organisation never authorised
Brokenness at scale automating a process whose underlying logic
is already wrong
Judgement substitution automating the step where a professional's
judgement is the deliverable
Null compliance the audit obligation is met on paper and
produces no accountability
AI washing advertising a capability the system does
not have
Override drift humans override the model precisely on the
cases it was built to decide
WORKED EXAMPLES
Drive-thru order capture, voice pilot, 2026
Q1 GREEN (customer corrects at the window)
Q2 GREEN (order confirmed before it binds)
Q3 GREEN Q4 GREEN -> AUTOMATE
Airline chatbot describing fare policy
Q2 RED (representation binds the carrier) -> RETRIEVE ONLY
Welfare debt calculation by income averaging
Q3 RED (unlawful logic, no contest path) -> FIX FIRST
Legal brief citations
Q4 RED (the judgement is the product) -> HUMAN SIGNS
BEFORE YOU SHIP
[ ] Name the person who signs each RED step.
[ ] Write the sentence a customer could quote back at you.
[ ] Measure how often the human overrides the model, by case
type. A rising rate on one type is override drift.
Q3 does most of the work and gets skipped most often. A process nobody can write down is not a candidate for automation — it is a candidate for being written down, which is why most pilots die long before the model is at fault.
Take one process this week. Ask whether the output is something a customer could hold you to. If it is, the model drafts and a named human signs — and the discovery is usually how few steps in the pipeline were ever decisions at all.
Ross Jones, Founder, The Hopium Lab. Last modified 22 July 2026. Engineering commentary, not legal advice; case descriptions come from the primary sources linked above, and allegations are identified as such.