Navigation
What We AreThe BrainPortfolioThe Lab's LabBuilt For YouThe WhiteboardServices & Prices
Let's Talk →
← The Whiteboard

How do you measure something without teaching people to fake it?

Name the quality in the question and you have contaminated the answer. The rubric design that separates a genuine signal from a well-trained echo — and why unprompted evidence is worth more than the same words after a prompt.

You measure a quality by asking content-free questions and scoring what the subject supplies unprompted, never by naming the quality you are looking for. A question containing the word "resilient" teaches the subject that resilience is wanted and hands them the vocabulary to perform it. Every rubric therefore needs three parts: what genuine evidence looks like in their own words, how to elicit it without leading, and the false positives that look identical to the real thing.

Ross Jones — Founder, The Hopium Lab. Last modified 22 July 2026.

Drawn from the extraction design in Living is Learning, a system that observes a learner's development from ordinary conversation. The pattern applies anywhere a system scores a human quality from what a human said.

Why does naming the quality destroy the measurement?

Naming the quality converts an observation into an instruction, and people are extremely good at following instructions.

Ask "why is that the logical answer?" and you have told the subject that logic is the currency here. Their next answer will be shaped like logic whether or not any reasoning occurred. Ask "weren't you brave to keep going?" and you have supplied both the verdict and the word — the only remaining move is agreement. What comes back is not evidence about the person. It is evidence about the question.

This is the same failure as a leading question in a courtroom or a badly written survey item, and it is endemic in AI evaluation, where a judge prompt routinely names the property it is scoring and then congratulates itself on finding it.

A leading question hands the subject your vocabulary. A clean question makes them produce their own — and only the second one is data.

What does genuine evidence look like?

Genuine evidence is specific, unprompted, and expressed in the subject's own construction rather than yours.

Four worked rubrics, from a taxonomy of human development. The pattern transfers; the dimensions are illustrative.

QualityEvidence, in their wordsA clean elicitationNever ask
Critical thinkingDistinguishes a guess from a fact unprompted — "I think it's this but I'm not sure"; generates a counter-case; asks for grounds; revises a position and says they are revising"What made you land on that one?" · "Was there another way it could have gone?""Why is that the logical answer?" — names the quality. "Don't you think X is wrong?" — supplies the counter-case
ResilienceRe-attempts differently after a setback; reframes failure as information — "so that's why it didn't work"; names a hard feeling without it collapsing into self-worth"What did the part that didn't work tell you?" · "What did you do after it went wrong?""Weren't you brave to keep going?" — sycophantic, and supplies the verdict
ArticulationOrders an explanation so a listener could follow; repairs their own wording toward precision; teaches you something and you genuinely follow it"Can you walk me through how that worked?" · "If I'd never seen it, how would you tell me?""Explain it clearly and in order" — instructs the form you were trying to observe
CuriosityGenerates their own next question unprompted; follows a tangent and connects it back; returns to a topic days later"Is there anything in that you're still wondering about?" · "What would you poke at next?""Wasn't that fascinating?" — supplies the affect. "What else should you learn about this?" — turns it into homework

The right-hand column is the one worth stealing. Most rubric design stops at defining the target; almost none of it enumerates the questions that would manufacture a false positive.

Why weight unprompted evidence higher?

Unprompted evidence is higher-quality because the subject chose to produce it, which means it survived their own filter for what was worth saying.

A question generated immediately after a prompt about that topic is weaker evidence of curiosity than a question launched from nowhere, and the scoring should say so explicitly rather than treating both as one unit of curiosity. Compliance produces the same surface behaviour as interest. The difference is entirely in what preceded it, which means the scoring has to see the conversational context, not just the utterance.

The same logic applies to timing. A topic returned to across days — detected by reactivation across sessions rather than from a single transcript — is far stronger evidence than enthusiasm within one conversation, because enthusiasm in the moment is cheap and returning is not.

What are the false positives that look identical?

Every quality has a counterfeit that scores the same on a naive rubric, and naming them is most of the work.

Fluent but empty verbosity is the counterfeit of articulation. Length correlates with the appearance of explanation, so a rubric must score listener-followable ordering rather than word count. Related: correctly recalled jargon with no own-words restatement, which reads as expertise and demonstrates only recall.

Mechanical repetition is the counterfeit of resilience. Attempting the identical thing again is persistence; resilience requires the second attempt to be different. Two others sit alongside it — the adult-pleasing "I'll keep trying!" with no actual re-attempt, and masking, where "it's fine" accompanies visible disengagement and only cross-checking sentiment catches it.

Reflex hedging is the counterfeit of critical thinking. Low confidence produces the same linguistic markers as genuine epistemic care. So does contrarianism with no counter-case attached, and parroting back a counter-example you supplied three turns earlier — which is the most common one, because the subject is being cooperative.

Compliance-curiosity is the counterfeit of curiosity, and topic-hopping is the trickiest of all: it looks like breadth of interest and is sometimes avoidance, distinguishable only by cross-checking against how and why previous topics stopped.

A rubric that cannot name its own false positives is not measuring the quality. It is measuring the surface behaviour most correlated with the quality — which is precisely what a motivated subject optimises.

How does this apply to LLM-as-judge?

It applies directly, because a judge prompt is an elicitation and inherits every failure mode above.

Naming the property in the judge prompt biases the judge toward finding it, in the same way naming it to a person does. Asking a single judge to score many properties at once produces contamination between them — a response that scores well on clarity drifts upward on accuracy. And a one-to-ten quality score has no defined false positive, which is why those scores cluster at seven and eight and why disagreement between runs is impossible to locate.

The fix is the same three-part structure: define what the evidence looks like as a quotable span, define the elicitation so the judge is asked what it observed rather than whether the property is present, and enumerate the counterfeits so the judge has something to reject. One narrow judge per property, each anchored to a quote, beats one judge scoring everything — because a judgement you cannot locate in the text is a judgement you cannot check.

What does a complete rubric look like?

A complete rubric has four parts, and most published ones have one.

RUBRIC DESIGN — FOUR PARTS, NOT ONE
The Hopium Lab · v1.0 · 22 July 2026 · take it, fork it, argue with it

For every quality you intend to measure:

1. EVIDENCE — in their words, not yours
   [ ] Three concrete behaviours, quotable, observable in a transcript
   [ ] At least one that requires the subject to have initiated it

2. ELICITATION — content-free
   [ ] Two questions that could be asked about literally any topic
   [ ] Check: does the question contain the name of the quality? Rewrite it
   [ ] Check: does the question supply the answer's shape, affect or verdict?

3. FALSE POSITIVES — the counterfeits
   [ ] What scores identically but is not the thing?
   [ ] What does compliance look like here?
   [ ] What does the anxious, eager-to-please version look like?

4. WEIGHTING — provenance matters
   [ ] Unprompted > prompted. Say by how much
   [ ] Across sessions > within a session
   [ ] Own construction > echo of a phrase you used earlier

THE TEST: could a subject who wanted to score well, and who had read your
rubric, produce a perfect false positive? If yes, part 3 is incomplete.

That final test is the whole discipline. Assume your rubric is public and your subject is motivated — because in an evaluation the model has effectively read your rubric, and in a performance review the human definitely has.

Where does this not apply?

It does not apply where the thing being measured is genuinely a task outcome rather than a quality.

If you want to know whether the code compiles, whether the invoice total matches, or whether the answer is the correct one, ask directly and score the result. Content-free elicitation is expensive and it is only worth its cost when the property lives in how someone responded rather than what they produced. Applying it to a factual check is theatre.

It also assumes enough conversational surface to observe. A single-turn interaction gives you very little unprompted material and no cross-session reactivation, so the weighting scheme collapses to almost nothing — which is worth knowing before designing an evaluation around one exchange.

Ross Jones, Founder, The Hopium Lab. Last modified 22 July 2026. The dimensions above are illustrative; the pattern is the point.