Ask an AI vendor for their eval set and their last three failed releases, for p95 latency measured on your corpus with your permission filters applied, for a failure they found themselves in production, for their model-deprecation exposure and whose retirement calendar governs it, and for a decision reproduced from six months ago without calling the model. Certifications are not architecture questions.
Ross Jones — Founder, The Hopium Lab. Last modified 22 July 2026.
Engineering commentary, not legal advice. Written from the vendor side of the table — I answer these questionnaires as well as send them.
Which of the standard vendor questions are theatre?
The ones every credible vendor answers correctly. A theatre question is one where every serious candidate gives the same acceptable answer, so the answer separates nobody — a week of procurement for zero bits of information at selection time.
The published checklists are made of them. DeepInspect's 27 questions cover provenance, access, data flow, logging and regulatory mapping. worqlo's 40 span seven headings, roughly four of them commercial. FairNow's — inside AuditBoard since October 2025 — reduced to ten governance questions with model answers attached. Between them: nothing on eval sets, latency percentiles, retrieval design, model pinning or reproducing a past decision. The closest is worqlo's "what are your false positive and false negative benchmarks?", two summary numbers with no stated dataset. Every vendor has those on a slide.
| The standard question | What the buyer thinks it tests | What it actually tests | Ask instead |
|---|---|---|---|
| Are you SOC 2 Type II? | The system is safe | A change-management process exists | Show me the last change that made the model worse |
| False positive / negative rates? | Accuracy | They own a slide | On which labelled set, what size, built by whom, changed when? |
| What is your uptime SLA? | Reliability | Commercial exposure, not behaviour | p95 and p99 on our corpus with our ACLs applied |
| Which LLM do you use? | Technical depth | Nothing. The answer is a logo | Which exact model string, and whose retirement calendar governs it? |
| Where is our data stored? | Control | Procurement's checklist is done | Reproduce a six-month-old decision from stored records only |
SOC 2 attestation reports and ISO/IEC 42001:2023 certification do real work on access control, change management and incident response, and no regulated buyer skips either. Vocabulary matters: SOC 2 yields an attestation report, never a certificate; ISO 42001 yields certification.
Every serious candidate has the certificate. A control that every candidate passes carries no information at selection time — it is a filter you already applied.
What should you ask about the eval set?
Ask for the set, and when they cannot hand it over, ask for everything about it except the data.
A competent vendor will say the set is built on customer data and cannot be shared. Correct answer, not evasion, and treating the refusal as a red flag makes you look green. Ask instead for size, who labelled it, inter-annotator agreement, the rubric, the regression threshold that blocks a release, who signs off, and the last release that failed. All shareable, none of it leaks a customer.
The tell is speed. A team that runs an eval before every release answers instantly, with a number and a grievance. A team without one answers in methodology nouns — rigorous, comprehensive, continuously monitored — and offers to follow up.
A vendor who cannot state the size of their eval set does not have one. Nobody forgets the number of the thing they run before every release.
What latency number should you ask for, and at what volume?
p95 end-to-end, measured on your corpus with your access-control filters applied, plus p99 if the output touches a customer.
Demos run unfiltered on a few hundred documents. Production runs at low selectivity across millions, because enterprise ACLs make most of the index invisible to most users. Filtered approximate nearest-neighbour search is a different engineering problem. arXiv:2602.11443 (11 February 2026) finds partition-based IVFFlat outperforming graph-based HNSW on low-selectivity queries, and pgvector's optimiser "frequently selects suboptimal execution plans". ACORN (SIGMOD 2024) reports 2–1,000× higher throughput at fixed recall against prior hybrid methods. Three orders of magnitude is not a config flag anybody sets correctly by accident. The unfiltered demo is the first named failure mode: recall and latency both measured with the hard case removed.
Refuse the bare threshold, including from me. A p95 figure means nothing without index type, embedding dimension, hardware, filter selectivity and whether reranking sits inside the measurement. Recall at fixed search effort degrades as the corpus grows and buys back with latency, so ask where that knob is set. The five numbers I ask for, on one labelled pair set: Recall@10, NDCG@10, MRR, p95 retrieval latency at the stated index and dimension, and cost per million tokens embedded at projected volume. Practitioner convention, not a standard. The gap is wide enough that a gateway vendor built a March 2026 press release out of nobody publishing p95 at all.
What happens when the model provider changes the model underneath them?
Faster than buyers assume, and the notice windows differ between providers by a factor of three.
OpenAI's deprecation policy commits to "at least 6 months" for generally available models, "at least 3 months" for specialised variants, and for preview models "such as 2 weeks" — all prefaced with "unless safety or compliance concerns require a faster timeline". Anthropic commits to at least 60 days for publicly released models. Observed rather than promised: claude-opus-4-20250514 was deprecated on 14 April 2026 and retired on 15 June 2026, and requests to retired models fail.
Then the follow-up nobody asks. Anthropic's dates cover Anthropic-operated platforms; partner platforms including Amazon Bedrock and Google Cloud set their own retirement schedules, so lifecycle status differs. A vendor reselling through Bedrock cannot answer from Anthropic's page, and watching them try is diagnostic. The reseller's calendar is the failure mode: retirement dates set by a platform the vendor never reads.
Expect one parry: Anthropic has committed to preserving model weights for at least the lifetime of the company. Preserved weights you cannot call do not replay a decision. Keeping the file is not answering the request.
Can they reproduce a decision from six months ago?
One question, and either the architecture survives it or the call ends early. Can you rebuild the answer without calling the model? If not, you don't have a system. You have a demo with a try/except around it. The full version is the Replay Test, and it is the single best question on this page.
The standard defence is the seed defence: temperature zero and a pinned seed. Both providers contradict the defence in their own documentation. OpenAI's cookbook promises only a "best effort to sample deterministically" and states that "determinism is not guaranteed" even with matching seed, parameters and system_fingerprint. On Anthropic's current line the vendor cannot set temperature at all — temperature, top_p and top_k return a 400 on a non-default value from Claude Opus 4.7 onward. The mechanism is batch invariance rather than floating-point folklore: Thinking Machines' Defeating Nondeterminism in LLM Inference (10 September 2025) ran 1,000 completions of Qwen3-235B-A22B-Instruct-2507 at temperature 0 on one prompt and got 80 unique outputs, diverging at token 103. Server load moves batch size; batch size moves your output. The engineering that survives it is its own post.
Six months is not a number I invented. EU AI Act Article 26(6) puts the retention floor on the deployer — you, not the vendor — at "at least six months" for logs under your control, with Article 19 mirroring it provider-side. Timing matters more than the floor. Under the Digital Omnibus agreed 6 May 2026, Annex III stand-alone high-risk obligations apply from 2 December 2027 and Annex I embedded systems from 2 August 2028, per the Commission's implementation timeline. What lands on 2 August 2026 is Article 50 transparency, GPAI, prohibitions and AI literacy. Nobody is under high-risk logging duties this August — and the vendor you sign in 2026 will still be running when the duty does. I sell the paid version of the question as a fixed-price engagement, which is why you should ask it for free first.
How do you ask a vendor to show you a failure?
Ask for a failure they found themselves, in production, and what shipped as a result.
Breaking a demo with adversarial input proves nothing, and the vendor is right to say you fed it garbage. A failure they caught, diagnosed and fixed is not fakeable, and the answer reveals whether anybody watches production at all. Barnett et al. (arXiv:2401.05856, 11 January 2024) put the point in peer-reviewable words: "validation of a RAG system is only feasible during operation", and "the robustness of a RAG system evolves rather than designed in at the start."
Hand over your five worst documents first — the crooked scan, the table spanning a page break, the superseded version beside its replacement. Certification text needs point-in-time resolution because the same clause is a different document on a different date.
A demo is a system running under conditions its author chose. No amount of polish converts that into evidence.
Who should be on the call, and what proves they wrote the retrieval layer?
The person who wrote the retrieval layer, and three answers prove authorship. "Send an engineer, not a salesperson" is already common advice, and vendors have learned to send a solutions architect — so the value sits in what you ask once they arrive: why that chunk size and what they tried before it; where the reranker cutoff sits and what it costs; what broke last quarter and in which layer. Authors answer in seconds, with a number and a regret. The solutions architect who has never opened the repository answers in architecture nouns and offers to follow up.
From my side of the table: a forty-question RFP takes me under an hour from a template, and I never open a repository to complete it. The question that has cost me a full working day came from a buyer's own engineer, asking for p95 on their corpus with their ACLs applied. Only a question that forces the vendor to measure something is doing work.
How do you run a two-week bake-off?
Two weeks, paid, same corpus, same queries, same window, thresholds written down before the first vendor is contacted.
Published guidance says four to six weeks and is right about production trials. A bake-off is not a production trial. Two weeks cannot observe drift or a silent model swap, so put both in the contract — named model strings, notice on substitution, log portability on exit. Two weeks tests whether a vendor can produce evidence under time pressure, the capability you need on the day something goes wrong. Pay for it: a free pilot buys their B team.
THE TWO-WEEK BAKE-OFF — a decision gate, not a production trial
The Hopium Lab · v1.0 · 22 July 2026 · take it, fork it, argue with it
BEFORE ANY VENDOR IS CONTACTED
[ ] Freeze one corpus. Real documents, real ACLs, real mess. No curation.
[ ] Write 120 queries: 60 routine, 40 from last year's escalations,
20 with no correct answer anywhere in the corpus.
[ ] Label them. Two annotators. Measure agreement. Publish the number.
[ ] Write the thresholds NOW. A threshold agreed after seeing results
is not a threshold, it is a negotiation.
DAYS 1-3 · ARCHITECTURE CALL (whoever wrote retrieval, not sales)
[ ] Chunk size, and what they tried first
[ ] Reranker cutoff, and its latency cost at that cutoff
[ ] Exact model string, and whose retirement calendar governs it
[ ] A production failure they found themselves, and what shipped
DAYS 4-10 · RUN IT ON YOUR CORPUS. SAME QUERIES. SAME DAY.
Record per vendor:
Recall@10 · NDCG@10 · abstention rate on the 20 unanswerable
p95 AND p99 end-to-end WITH your ACL filters applied
cost per 1,000 queries at projected volume
index type · embedding dimension · hardware · reranking in or out
(a latency number without those four variables is not a number)
DAYS 11-14 · THE EVIDENCE TEST
[ ] Take a query answered on day 4. From stored records only, with no
call to any model provider, reproduce the output byte-for-byte.
[ ] Hand the record to somebody outside the evaluation. Can they verify it?
FAIL — any single one ends the evaluation
Abstains on fewer than 18 of the 20 unanswerable queries
p95 measured without your ACLs, or refuses to measure with them
Names the provider but cannot name the exact model string
Cannot reproduce day 4 without calling the model. No partial credit.
Nobody on the architecture call has written any of the code
Eval-set size is unknown to everyone present
Two failure modes are left to name. The borrowed benchmark: recall quoted from a public dataset that resembles nothing you own. The pilot that proves the vendor can run a pilot: their engineers tune against your queries all fortnight, and you learn what a staffed system does rather than a shipped one. Cap the hands-on hours in the order form.
Every product I ship carries an MCP server and a CLI, so a buyer can run these questions against my systems instead of taking my answers on trust. Demand the same. A vendor whose architecture only demonstrates itself through a screen-share has told you where the seams are.
Ross Jones, Founder, The Hopium Lab. Last modified 22 July 2026. Engineering commentary, not legal advice, and not an assurance service — my tooling is first-party by definition.