Navigation
What We AreThe BrainPortfolioThe Lab's LabBuilt For YouThe WhiteboardServices & Prices
Let's Talk →
← The Frameworks

The Grounded Report

a reliable multi-source reporting architecture, Pour five sectors into one model and it drifts, invents numbers, and cannot tell you why it said anything. This is the architecture for a report you can actually stand behind.

The shape of it
Probabilistic understanding, the front
the line, interpretation ends and governance begins
Deterministic governance, the shell

The Grounded Report is an architecture for multi-source reporting that does not hallucinate. Every figure enters as a sourced data call rather than a guess, every claim is settled by a narrow judge that fires only on evidence rather than one prompt asked to assess everything at once, and the finished report is a replayable projection of the underlying record, so each statement carries its lineage back to the thing that produced it. It is the difference between a report that sounds authoritative and one you can defend line by line.

The diagram above walks the whole pipeline. Tap any stage to open it.

Ross Jones, Founder, The Hopium Lab. Last modified 4 August 2026.

What problem does The Grounded Report solve?

The specific, repeatable failure of trying to generate a serious multi-source report with a single model call.

Take the common example: a corporate responsibility or ESG report that has to cover five sectors at once, financial, environmental, governance, social, and geopolitical, drawing on filings, prices, disclosures and internal documents. The obvious approach is to hand all of it to one capable model and ask for the report. It works in the demo and falls apart in production, in three predictable ways: the model drifts across the sectors, it invents figures that look plausible, and when you ask why it concluded something, there is no answer, only the output.

Those are not prompt-quality problems you can fix with better wording. They are structural, and they need an architectural answer.

Why does one big prompt drift?

Because a single pass is being asked to make dozens of judgements at once, with nothing anchoring any of them and a working memory it is free to overwrite.

When you ask one prompt to score forty things across five sectors, each individual judgement gets a sliver of the model's attention and none of them is tied to specific evidence. The model fills the gaps the way language models fill gaps, fluently, which is exactly the mechanism that produces an invented number stated with total confidence. Add a long context of source documents and the problem compounds, because the model now has more to lose track of. This is a case where the honest move is to notice you are asking the model to do something it structurally cannot, and to tell whether you are looking at a retrieval failure or a genuine hallucination before you trust a word of it.

What does decomposition look like?

Forty small evidenced decisions instead of one giant guess.

Each claim the report might make gets its own narrow judge. That judge answers one question, fires only when it can cite a direct quote or a specific figure, returns a confidence, and records what it saw. A claim with no supporting evidence produces no output rather than a confident fabrication. This is the single highest-leverage move in the whole architecture, and it is the same discipline that makes an LLM usable as a judge at all: narrow scope, evidence required, confidence recorded.

How do the numbers stay real?

They enter as data calls, not inferences.

A stock price, a revenue figure, an emissions number: these are looked up from a source and attached to the record as data. The model is never asked to recall or estimate them, because a model doing arithmetic or remembering a figure is a model making something up with good grammar. Never let the model do the arithmetic is not a style preference, it is the line between a report that is right and one that is merely convincing.

How does it find the links between sectors?

The cross-sector connections are edges in a graph projected from the record, not a paragraph the model was asked to write.

The most valuable part of a serious report is usually the connection nobody spelled out: a geopolitical event touching a supply line touching a currency exposure touching a number in the financials. A single-pass model either misses these or invents them. Here they are a traversal over a graph built from the evidence, so a claimed link is one you can follow back to the two records that co-occurred to produce it. The analysis becomes something you can point at.

Why can you trust it a week later?

Because the report is a deterministic projection of an immutable record, not a one-time generation.

The judges' outputs feed a pure, replayable engine. Aggregations, weightings and section scores are computation over the evidence, so the same inputs produce the same report every time. Correct a single input, and the whole thing re-derives with no stale figures left behind. Every claim keeps its lineage, so the report is reproducible and survives the replay test: rebuild it from the record and you get the same document, which is the only real basis for trusting it later.

Where does The Grounded Report not apply?

It does not apply where there is genuinely nothing to anchor a judgement to.

Some reports are pure narrative or forecast, a view of where a market is heading that no current evidence can confirm or deny. Forcing an evidence-anchored judge onto a genuinely speculative claim just produces false precision, and the honest move is to mark that section as opinion rather than dress it as fact. It also earns little on a one-off report you will never reproduce or defend, where the whole apparatus is overhead. And it assumes you can get at governed sources: where the inputs are somebody else's, unverifiable, or arriving too fast to structure, the foundation is missing.

Use it wherever a report has to be defended, reproduced, or trusted by someone who was not in the room. That describes most reporting that actually matters.

The transferable core

THE GROUNDED REPORT, THE TEST
The Hopium Lab · v1.0 · 4 August 2026 · take it, fork it, argue with it

1. CAN EVERY NUMBER NAME ITS SOURCE?
   [ ] Did each figure enter as a data call, or did the model produce it?
   [ ] Could you click any number and see where it came from?

2. IS EACH CLAIM A SEPARATE, EVIDENCED DECISION?
   [ ] Is one prompt scoring many things at once, or many judges scoring one thing each?
   [ ] Does a claim with no evidence produce nothing, or a confident guess?

3. DOES THE REPORT REBUILD THE SAME WAY TWICE?
   [ ] Is the report a projection of a record, or a one-time generation?
   [ ] Correct one input: does everything downstream re-derive, or drift?

THE TEST: if you cannot answer "why does it say this?" for any sentence in
the report, you did not write a report. You printed a very confident guess.

Ross Jones, Founder, The Hopium Lab. The Grounded Report shares its spine with The Compounding Stack and The Trustworthy Record: a probabilistic front, a hard line, a deterministic shell.