“The users like it” is not a finding
A demo, a positive quote and a shipped changelog entry are all compatible with a surface nobody trusts — the only way out is to name, before you build, which question each number is going to answer.
The review gate demos well. The per-field confidence bands read clearly, the source-image crop lands next to the extracted value, the raw OCR is one disclosure away, and nothing in the flow makes a reviewer hunt for anything. Somebody in the room says the users like it. Somebody else says it feels much faster than the old process. The work ships.
Six months later a different portfolio company asks whether it is worth adopting, and the honest answer available to you is: it demoed well, it shipped, and nobody has complained. Every one of those is true. Not one of them is a finding. All three are equally compatible with a review surface that reviewers have quietly stopped reading.
This lesson is about the gap between those two things, and about the one move that closes it, which happens before any instrumentation exists: naming the question each number is going to answer rather than shipping first and mining the logs afterwards.
What “the users like it” is actually a measurement of
It is not nothing. In the Google HEART framework it has a name — Happiness, one of five categories, covering the attitudinal side of user experience: satisfaction, visual appeal, likelihood to recommend, perceived ease of use. The other four are Engagement (depth and frequency of interaction), Adoption (new users in a period), Retention (users from an earlier period still present later) and Task success (efficiency, effectiveness, error rate).
Happiness is a legitimate category. The problem is that it is a category, offered as if it were a conclusion, and that on a review gate it is the category least connected to whether the thing works. A reviewer can be perfectly satisfied with an interface that is causing them to approve incorrect extractions quickly and pleasantly. Satisfaction and correctness are not in tension — they are simply unrelated, and only one of them is what a portfolio company is buying.
The HEART authors make the sharper version of this point against the metrics that are easiest to collect, which they group as PULSE — page views, uptime, latency, seven-day active users, earnings. Their objection is not that these are wrong but that they are either very low-level or indirect metrics of user experience, and their worked example is the exact ambiguity this course keeps returning to: page views for a feature may rise because the feature is genuinely popular, or because a confusing interface leads users to get lost in it, clicking around to figure out how to escape. Same number, two readings, and the flattering one is the default.
Goals, then signals, then metrics — in that order
The transferable part of that paper is a three-rung ladder, and the only thing that makes it work is the direction of travel.
- Goals. What is this surface for? Not what it does — what it is supposed to achieve. Stated in a sentence a non-designer would accept.
- Signals. How would success or failure show up, in behaviour or attitude? Still not a number; a description of what would be different in the world.
- Metrics. The countable version of a signal, defined well enough to put on a dashboard.
Two instructions from the paper are worth quoting because they are the ones people skip. At the goals stage: do not get too distracted by worrying about whether or how it will be possible to find relevant signals or metrics. Name the question before checking what is instrumented, because a question filtered through existing instrumentation is just a description of your logging. And at the signals stage: choose signals sensitive and specific to the goal — they should move only when the user experience is better or worse, not for other, unrelated reasons.
That second one is a high bar, and the five metrics this course is built on all fail it to some degree. Correction rate moves when the agent changes, when the reviewer changes, when the input distribution changes, and when the interface makes approving cheaper. Knowing that in advance is the difference between a metric you can defend and a metric you can only present.
The organisational reason the ladder exists is worth naming too. The authors are blunt that the difficulty is not analytical: product teams have not always agreed on or clearly articulated their goals, which makes defining related metrics difficult. Ninety portfolio products with ninety unstated goals will produce ninety differently-shaped numbers under one shared metric name. Writing the goal down is the intervention.
The goal is not more trust
There is a specific way this goes wrong on an agentic surface, and it is worth heading off before you write a single goal statement. The obvious goal for a review gate is “make reviewers trust the agent.” That goal is wrong, and the human-factors literature has been explicit about why for two decades.
Lee and See define trust as the attitude that an agent will help achieve an individual’s goals in a situation characterised by uncertainty and vulnerability, and then spend the paper on the observation that reliance fails in two directions, not one. Misuse is relying on automation where its assumptions do not hold. Disuse is rejecting capabilities it actually has. Both are inappropriate reliance. Their design recommendation is a single line and it belongs on the wall: design for appropriate trust, not greater trust.
They give the vocabulary for saying what “appropriate” means, which is more useful than the slogan:
- Calibration — the correspondence between a person’s trust and the system’s actual capabilities. Overtrust is trust exceeding capability; distrust is trust falling short of it.
- Resolution — how finely trust tracks capability. Poor resolution means large changes in capability produce small changes in trust.
- Specificity — whether trust attaches to a particular component or moment, or spills across the whole system indiscriminately.
Read the metric set through that lens and it reframes. A falling correction rate is only good if capability rose to meet it — that is calibration. Per-field confidence is an attempt at specificity: it exists so a reviewer can distrust one field without distrusting the document. If a review gate’s confidence display makes no difference to which fields get checked, it has failed at specificity, and no aggregate metric in the set will tell you that.
The question each number is answering
Run the ladder on a review gate and the five metrics stop being a list and start being answers to particular questions:
- Correction rate — how often does the human disagree with the agent, and where? It is the only direct measure of disagreement you get for free from the workflow itself.
- Acceptance rate — how much of the agent’s output survives contact with a human untouched? Which is the actual question behind “is this saving anyone time.”
- Time-to-trust — how long before a reviewer settles into a stable relationship with the agent? Explicitly this course’s own construct, and the definitions lesson does not let that slide.
- Mid-review abandonment — how often is the gate impossible, or not worth, finishing? A completion metric that exists to catch cases where the surface itself is the obstacle.
- Escalation rate — how often does a reviewer decide they are not the right person to decide? The only one of the five that measures a reviewer knowing their own limits, and the only one that cannot exist without an affordance built for it.
Note what none of them answers: whether the agent is correct. Correctness needs ground truth, and a review gate does not have any — it has a human’s judgement, which is a different and weaker thing. That limitation is not a gap to be closed later; it is a permanent property of instrumenting an interface rather than evaluating a model, and pretending otherwise is how a UI metric gets quoted as an accuracy figure.
The NIST AI Risk Management Framework is unusually direct about this class of problem for a standards document. It states plainly that there is a current lack of consensus on robust and verifiable measurement methods for risk and trustworthiness, and that measurement approaches can be oversimplified, gamed, lack critical nuance, become relied upon in unexpected ways, or fail to account for differences in affected groups and contexts. Its MEASURE 1.1 requirement is the one to steal outright: the characteristics that will not — or cannot — be measured must themselves be documented. A spec that lists only what it measures is only half a spec.
The moment a metric becomes a target
Marilyn Strathern’s 1997 paper on audit in the British university system gives the formulation everyone quotes: “When a measure becomes a target, it ceases to be a good measure.” Every metric in this course is a candidate target, and correction rate is the obvious one: it can be moved decisively without touching the agent, simply by making approval cheaper than scrutiny.
The provenance is its own small lesson. Strathern does not claim the line as hers — she attributes the naming to Hoskin, who named it after Charles Goodhart’s observation about monetary control. Goodhart’s 1975 original could not be opened for this course, so the sentence circulating as “Goodhart’s law” is quoted here from Strathern, who is quoting an attribution. The most-repeated sentence in metrics criticism is itself a number nobody checked the source of.
Check your recall
Answer from memory — no scrolling back.
Retrieval check
Your spec measures correction rate, acceptance rate and escalation rate. What does NIST’s MEASURE 1.1 requirement say you still owe?
Check your answer
A written statement of what will not, or cannot, be measured. In this case: agent correctness, since the gate has a human’s judgement rather than ground truth; errors of omission, since a field the reviewer never looked at leaves no trace distinguishable from a field that was checked and found right; and anything about reviewers who never reached the gate at all.
The reason this is a requirement rather than good manners is the companion line in the same framework: an inability to measure a risk does not imply the system poses either a high or a low risk. An unmeasured thing reads as a zero on a dashboard. Documenting the gap is what stops the absence being read as a result — and in a portfolio setting, where your spec is handed to teams who were not in the room, it is the only defence you have against your own silence being quoted back as evidence.
Hands on
Write the goals before the metrics
Done when: INSTRUMENTATION.md opens with a Goals section containing two or three goal statements for a real review gate, each with its signals, plus a written list of what this spec will not be able to measure — and no metric names appear anywhere in the goals.
- Open
learning/agent-evaluation/INSTRUMENTATION.mdand add a section 0 — Goals above the events table. The file ships with four tables and no goals on purpose; supplying them is the first thing the course asks of you. - Write two or three goals for the HouseWarm review gate specifically. Each must be a sentence about what the surface is for that a non-designer would accept, and each must be capable of failing. If a goal cannot be false, it is a mission statement.
- Under each goal, write its signals: how success or failure would show up in behaviour. Still no numbers, still no metric names. Then check each signal against the specificity test — would it move for reasons unrelated to the experience getting better or worse? Write the honest answer next to it, because for most of these the answer is yes.
- Now write one goal that is deliberately about appropriate reliance rather than more of it — something that would be violated by a reviewer approving too readily and by a reviewer correcting fields the agent got right. Both failure modes named in the same sentence.
- Add a short Will not measure list in section 4, per NIST’s MEASURE 1.1. At minimum it has to include agent correctness and errors of omission, with one line each on why the interface cannot see them.
- Bring the goals into the chat. I will try to satisfy each one with a reviewer who has stopped reading. Any goal I can satisfy that way needs rewriting before you go near a metric.
What this does not cover
Goals and signals are not yet numbers, and this lesson deliberately kept them apart — a signal that arrives already wearing a numerator is a signal shaped by what was convenient to log.
The definitions lesson does the descent to the bottom rung: all five metrics pinned to a numerator, a denominator, a window and a population, tightly enough that an engineer three time zones away computes the same number from the same spec. It also confronts the one term in the set that has no standard definition at all, and decides what to do about that in public. After it comes the interpretation lesson, which is where the ambiguity that has been circling this whole page — the same number, two readings, the flattering one by default — gets treated as the subject rather than the warning.
Read this next — primary source
Measuring the User Experience on a Large Scale: User-Centered Metrics for Web ApplicationsKerry Rodden, Hilary Hutchinson & Xin Fu (Google), CHI 2010 — free author copy hosted by Google; a practitioner note, not an independent study
This lesson takes the Goals-Signals-Metrics ladder and the PULSE critique. The full paper is six pages and worth all of them for the worked examples — Gmail, Google Maps, Google Finance — which show the ladder failing as often as it works, and for its two explicit caveats: not every metric category belongs on every product, and these metrics are for evaluating launched products, not a substitute for formative research. It is also the clearest short statement anywhere of why a metric detached from a stated goal is worse than no metric, which is the argument this whole course is built on.
Stuck, curious, or think this lesson is wrong? Ask your teaching agent. The lessons are the scaffold; the conversation is where the learning gets unstuck.