Navigation
What We AreThe BrainPortfolioThe Lab's LabBuilt For YouThe WhiteboardServices & Prices
Let's Talk →
← The Whiteboard

Which business processes should not use AI?

Plenty of people say don't automate everything. Nobody names a case, because naming one means checking it. Six functions, five named failure modes plus the one nobody names, four questions that decide it — and the corrections to three disasters everybody miscites as AI.

Never automate three things. The step where a professional's judgement is the deliverable. A process whose underlying logic is already wrong — automation scales brokenness rather than repairing it. Any output that becomes a promise the organisation must honour. Everything else turns on two questions: is being wrong recoverable inside the organisation, and does the output create a commitment somebody outside can act on?

Ross Jones — Founder, The Hopium Lab. Last modified 22 July 2026. Twenty-four shipped products across fifteen sectors, every one with an MCP server and a CLI.

Hasn't somebody already written the negative version?

Yes, and pretending otherwise is the quickest way to lose the only readers worth having. Zapier, Retool and CIO have each published a variant: don't automate one-off decisions, processes with no digital record, or the thing that defines your core promise. The advice is right and nearly identical everywhere.

Not one of those pages names a case. Abstraction is worthless in the meeting where somebody has already signed the licences. What follows is function-level: one named failure mode per function, a primary source wherever the claim is load-bearing, and a correction wherever the story everyone repeats is wrong.

What did the worst automated-decision disasters actually run on?

Mostly not AI. Miscategorising them hands a hostile reader the whole argument.

Australia's Robodebt scheme was rule-based data matching plus income averaging — no machine learning anywhere. The Royal Commission reported on 7 July 2023 that averaging annual ATO income across fortnights was inconsistent with social security legislation, and therefore unlawful. AU$746m was wrongfully recovered from roughly 381,000 people and later refunded, and around AU$1.75bn in raised debts was ultimately wiped. The purest illustration of brokenness at scale — automating a process whose underlying logic is already wrong.

The EEOC's iTutorGroup settlement is cited everywhere as the first AI hiring discrimination case. The EEOC's own release of 11 September 2023 never says AI, machine learning or algorithm. It says the company "programmed their tutor application software to automatically reject female applicants aged 55 or older" — and men aged 60 or older — a hard-coded date-of-birth filter, caught because one applicant sent two identical applications differing only in birth year. $365,000 and five years of monitoring.

Moffatt v. Air Canada, 2024 BCCRT 149 is a British Columbia Civil Resolution Tribunal decision of 14 February 2024, and the tribunal defined the thing neutrally: an automated system responding to a user's prompts and input. No finding about the underlying technology. Damages: CA$650.88.

Three of the most-cited automated-decision disasters of the last decade were never found to involve machine learning at all. The danger was never the model class. It was automating a decision nobody could contest.

Which processes should each function stop automating?

The split is the same in every function: automate the work, never the commitment.

FunctionAutomate freelyNever run unattendedNamed failure mode
SalesEnrichment, routing, transcription, pipeline roll-upsAny quote or term a customer can hold you toPromise leakage
SupportTriage, order status, retrieval of documented policyAnswering where documentation is silent; granting exceptionsPromise leakage
FinanceReconciliation matching, anomaly flagging, coding suggestionsFigures that enter a filing; arithmetic done in proseAI washing
HRScheduling, job-ad drafting, policy FAQScreening, ranking, rejection, promotion, terminationNull compliance
LegalClause extraction, first-pass review against a real corpusCitation generation; the advice itselfJudgement substitution
OpsBounded forecasting, routing, capacity planningEligibility or penalty decisions with no contest pathBrokenness at scale

The finance row is the counter-intuitive one. The regulated risk there is not the model being wrong — it is AI washing, advertising a capability the system does not have. SEC press release 2024-36 of 18 March 2024 settled the first such actions against two registered investment advisers — $225,000 and $175,000, one marketed as the first regulated AI financial adviser. The SEC found neither had the capability advertised; both settled without admitting or denying the findings. No number that reaches a filing should be computed by a language model in prose — finance and legal sit on one rule, the model may propose, never compute.

What makes an output a promise the organisation must honour?

Anything a customer can reasonably act on. Promise leakage — an automated interface creates a commitment the organisation never authorised.

Air Canada's chatbot told a grieving passenger he could apply for bereavement fares retroactively. Company policy forbade it. Air Canada argued the chatbot was a separate legal entity responsible for its own actions; the tribunal's reply — "This is a remarkable submission" — is the only sentence anyone needs from the decision. The tribunal found negligent misrepresentation, resting on a duty of care arising from the commercial relationship between service provider and consumer. Small jurisdiction, limited precedential weight, technology-agnostic principle.

Anysphere's Cursor support bot went further in April 2025. It explained a session-management bug by asserting a one-active-session-per-subscription limit that did not exist, and users cancelled before the company could correct it.

Automation is a promise-making machine. Whatever the interface says, the organisation has said — and inventing the policy is worse than getting it wrong, because there is nothing to correct it against.

Where is the failure rate actually measurable?

In legal work, where courts write the failures down. Damien Charlotin of HEC Paris maintains a database of decisions in which a judge explicitly found, or clearly implied, reliance on AI-hallucinated material: 719 by January 2026, 1,227 by early April 2026, 1,598 as of 9 June 2026 — roughly six new cases a day across that window, and stale by the time you read it.

Judgement substitution — automating the step where a professional's judgement is the deliverable. In Mata v. Avianca (S.D.N.Y., 22 June 2023) the aggravating factor was never using ChatGPT: Judge Castel sanctioned two attorneys and their firm $5,000 jointly for conscious avoidance and false statements after the six fabricated citations were questioned. The model produced the draft; the lawyers declined to do the one thing they were being paid for.

The sharpest example of judgement substitution is a consultancy audit. A roughly A$440,000 assurance review, delivered to an Australian government department in July 2025, contained fabricated academic references and an invented quote attributed to a Federal Court judgment with the judge's name misspelled. Deloitte repaid the final contract instalment; the corrected version disclosed Azure OpenAI GPT-4o had been used. The subject of the review was the department's automated compliance system.

Running the gate below across a live process is a one-day exercise. I sell it as a fixed-scope go/no-go; the procedure is free either way.

Does regulation settle which processes to avoid?

No. Regulation tells you which processes carry paperwork, a different question entirely. EU AI Act Annex III reads like a function-by-function list: 4(a) recruitment and candidate filtering, 4(b) promotion, termination and task allocation, 5(a) eligibility for essential public benefits, 5(b) creditworthiness with an express carve-out for fraud detection, 5(c) life and health insurance pricing, 5(d) emergency call triage. Article 14 requires design permitting effective oversight by natural persons during use.

The deadline moved. Final Council approval on 29 June 2026 defers stand-alone Annex III high-risk obligations from 2 August 2026 to 2 December 2027, and embedded Annex I systems to 2 August 2028. Anyone still writing a countdown to next month has not read the Omnibus. Sixteen extra months means the commercial argument has to stand on its own — which it does. The logging mechanics sit in a separate note.

Where a rule already bites, the result is what Wright and colleagues named null compliance — the audit obligation is met on paper and produces no accountability. New York City's Local Law 144 mandates bias audits for automated employment decision tools. Their study of 391 employer observations found 18 posting the required audit (~5%) and 13 the required notice (~3%), with almost every published audit reporting no adverse impact. The New York State Comptroller concluded on 2 December 2025 that enforcement had been ineffective, finding at least 17 potential non-compliances despite penalties of up to $1,500 per violation per day.

Buying the screen from a vendor does not move the liability, either. In Mobley v. Workday (N.D. Cal.), Judge Rita Lin granted preliminary ADEA collective certification on 16 May 2025 covering applicants aged 40+ since 24 September 2020, on the theory that the vendor acts as its client-employers' agent and so falls within the statutory definition of employer. Allegations at the certification stage. Nobody has been found liable. The vendor is nonetheless inside the blast radius.

Isn't a post built on failures just sampling on the outcome?

Yes, and the honest correction changes the claim. Successful automations do not generate royal commissions. McDonald's ended its IBM drive-thru order-taking pilot in June 2024 after viral misorder videos, then launched a Google Cloud voice assistant in June 2026, now running at US sites in English and Spanish. Cite the 2024 failure as proof the process cannot be automated and an informed reader points out the company re-automated it.

Amazon's 2014 résumé-rating tool cuts the same way. Reuters reported in October 2018 that it penalised résumés containing "women's" — and that it never left trial. A tool caught and killed before deployment argues for pre-deployment testing, not against automation.

Zillow Offers is told wrong most often of all. Stanford GSB analysis attributes the $304m Q3 2021 write-down and roughly $569m total to business strategy rather than algorithmic inadequacy: a late entrant buying non-cookie-cutter homes where valuation models are weakest, then bidding above its own model's estimate to win volume. Override drift — humans override the model precisely on the cases it was built to decide — is a governance failure wearing an algorithm's coat.

The defensible claim is about the shape of failure, never its frequency.

What four questions settle it?

Recoverability, promise, process integrity and judgement — run per process step rather than per system. A gate that flags everything is worthless, so the worked examples clear one live automation and block three famous ones.

THE FOUR-QUESTION AUTOMATION GATE — v1.0
Run per step. Any RED = the model may propose, a human decides.

Q1. RECOVERABILITY
    If this step is wrong 1 time in 100, can the organisation
    undo it within its own control, at its own cost, before
    anyone outside notices?
    [ ] GREEN  reversible internally
    [ ] RED    lands on a customer, a filing, or a named person

Q2. PROMISE
    Does the output create a commitment a third party can
    reasonably act on?
    [ ] GREEN  informational, no reliance
    [ ] RED    price, entitlement, policy, eligibility, deadline

Q3. PROCESS INTEGRITY
    Is the current process written down, lawful, and contestable
    by the person it affects — TODAY, before automation?
    [ ] GREEN  documented + contest path exists
    [ ] RED    lives in someone's head, or has no appeal route

Q4. JUDGEMENT
    Is the judgement in this step the thing being paid for?
    [ ] GREEN  judgement sits upstream or downstream
    [ ] RED    the judgement IS the deliverable

NAMED FAILURE MODES (identical wording wherever used)
  Promise leakage      an automated interface creates a commitment
                       the organisation never authorised
  Brokenness at scale  automating a process whose underlying logic
                       is already wrong
  Judgement substitution automating the step where a professional's
                       judgement is the deliverable
  Null compliance      the audit obligation is met on paper and
                       produces no accountability
  AI washing           advertising a capability the system does
                       not have
  Override drift       humans override the model precisely on the
                       cases it was built to decide

WORKED EXAMPLES
  Drive-thru order capture, voice pilot, 2026
    Q1 GREEN (customer corrects at the window)
    Q2 GREEN (order confirmed before it binds)
    Q3 GREEN  Q4 GREEN            -> AUTOMATE
  Airline chatbot describing fare policy
    Q2 RED  (representation binds the carrier) -> RETRIEVE ONLY
  Welfare debt calculation by income averaging
    Q3 RED  (unlawful logic, no contest path)  -> FIX FIRST
  Legal brief citations
    Q4 RED  (the judgement is the product)     -> HUMAN SIGNS

BEFORE YOU SHIP
  [ ] Name the person who signs each RED step.
  [ ] Write the sentence a customer could quote back at you.
  [ ] Measure how often the human overrides the model, by case
      type. A rising rate on one type is override drift.

Q3 does most of the work and gets skipped most often. A process nobody can write down is not a candidate for automation — it is a candidate for being written down, which is why most pilots die long before the model is at fault.

Take one process this week. Ask whether the output is something a customer could hold you to. If it is, the model drafts and a named human signs — and the discovery is usually how few steps in the pipeline were ever decisions at all.

Ross Jones, Founder, The Hopium Lab. Last modified 22 July 2026. Engineering commentary, not legal advice; case descriptions come from the primary sources linked above, and allegations are identified as such.