Picking a measurement you can actually get
The best measurement you cannot obtain is worth less than the mediocre one already instrumented — how to choose against what the company can currently observe rather than against what a research paper would want.
Your three cases have a question mark in the measurement field. This lesson removes it, and the removal is less satisfying than you want it to be, because the measurement that would settle your argument is usually not one your company can produce.
The failure this prevents is specific and common. You ship the change, wait a quarter, go looking for the number, and find that nobody instrumented the thing you needed — or that they did, but only started after the change went out, so there is no before. You now have shipped work and no evidence, which is the position from which people start making claims they cannot support.
The ladder
Sort every candidate measurement by how hard it is to obtain, not by how convincing it would be. This course’s own ordering, five rungs, cheapest first:
- Already reported to someone outside your team. It appears in a board deck, a monthly business review, a customer success dashboard. Somebody else already owns it and already defends its definition. This is the cheapest measurement in existence and it is almost always undervalued because it is not yours.
- Instrumented and queryable, but nobody looks. The events are in the warehouse. Getting the number is a query and a conversation with whoever owns the schema.
- Obtainable by asking a named person. A support lead can tell you what the top ticket categories were last quarter. Not a system, a person. Slower, but real, and it usually beats the query because the person can also tell you what the number means.
- Requires new instrumentation. Somebody has to build logging, and it has to be built and running before your change ships or you have no baseline.
- Requires an experiment. A holdout, a staged rollout, randomised assignment. This is the only rung that produces a real counterfactual, and it needs a person with authority to agree that some customers get the worse version on purpose.
The point of the ladder is not that low rungs are better. It is that rungs four and five have a deadline attached and the first three do not. Instrumentation and experiments are decisions you must make before you ship. Everything else you can decide afterwards. So the question to ask while a case is still a case is not “what is the best measurement” but “is there anything on this list I lose forever if I do not act now.”
Measure the near end of the chain
Here is the move that does the most work, and it follows directly from the mechanism field you already wrote.
Your mechanism is a chain: something changes in the interface, an operator behaves differently, an error does not propagate, a support contact does not happen, a renewal conversation goes differently. The far end of that chain is the valuable one. It is also the one with the most other causes attached to it, and the one you are least likely to be able to measure cleanly.
The near end is the opposite on both counts. It is less impressive and far more attributable, because there are fewer competing explanations for it. If the review gate flags low-confidence fields and an operator corrects them, then the count of corrections made at the gate is a measurement of the mechanism’s first link — and that number did not exist before the gate, because the gate is what creates it.
A claim built on the near link sounds smaller and holds up better: the gate intercepted this many low-confidence fields last quarter, this share of them were corrected rather than approved as-is, and each of those is an error that previously reached a downstream system. Every clause is countable. The financial value of it is still an inference — you have not shown what a caught error is worth — but you have moved the inference to the end of the sentence where the listener can see it and price it themselves, rather than burying it in the middle where it looks like a fact.
Take the baseline before you need it
A measurement you can get after the change but not before is not a measurement. It is a level, and a level tells you nothing about a change.
This is the most retrievable-in-a-hurry idea in the lesson, so it is worth saying in the most annoying possible way: the cheapest thing you will ever do for a value case is write down the current number before you ship. Not measure it properly. Write it down. Ask the support lead what their top three ticket categories are today and put the answer in the case file with the date on it. That single line, taken in five minutes, is the difference between an argument and an anecdote three months later.
Where this course parts company with the standard advice
Nielsen Norman Group, which sells UX training and consulting and is therefore an interested party in this question, published a useful piece arguing that design teams overthink return-on-investment maths. Its second myth is that an ROI calculation has to be perfectly accurate, and the article’s answer is direct: “We want ROI calculations to be as accurate and realistic as reasonably possible. But, ultimately, these are only estimates.” Its third myth is that the calculation has to account for every detail; the advice there is to do “only as much work on these calculations as is necessary.”
Both are right, and this course agrees with them. An estimate is a legitimate artifact. Refusing to estimate because you cannot be precise is its own kind of dishonesty, and it is how design functions talk themselves out of the conversation entirely.
The part to be careful with is what happens to an estimate as it travels. You write “roughly, on these assumptions, this is worth about X.” It goes into a deck as “worth X.” It is repeated in a review as “delivered X.” Nobody lied at any step. The assumptions simply fell off, because assumptions are the part of a sentence that does not survive being retold.
The defence is to make the assumptions structurally inseparable from the number rather than adjacent to it. Not a footnote. Inside the sentence: if a caught error is worth one support contact, and a support contact costs what our own support lead says it costs, then this quarter’s catches are worth about X. Someone repeating that sentence has to carry the conditions with it or visibly cut them out.
Where people get burned
The specific version of this you are most likely to commit: taking an estimate you made before shipping and describing it afterwards as a result. The tense change is the whole lie and it costs one word. “We expect this to save” becomes “this saved.” Reread your own case file for past-tense verbs attached to numbers you never went back and checked.
The falsifier
A measurement without a stated threshold is not a commitment, because any result can be narrated as a partial success after the fact. The falsifier fixes the narration in advance: a number, a direction and a date, written before you have the data.
Two properties make one real. It has to be a result you consider genuinely possible — a falsifier you are certain will not happen is decoration. And it has to be one you would actually act on, which usually means naming the action: stop the work, hand it back, take the claim out of the case.
If, by the end of Q2, the share of gated documents receiving at least one correction is under five percent, then operators are approving almost everything the model produces, the gate is friction rather than interception, and I will withdraw the cost-to-serve claim.
Notice what that sentence does not require. No holdout, no experiment, no data team. It is a measurement from rung two of the ladder with a threshold and a date attached, and it is enormously more credible than a retention claim you cannot check, precisely because you can be shown to be wrong by the end of the quarter.
Check your recall
Answer from memory — no scrolling back.
Retrieval check
The advisory recommendation surface recommends but never executes. Somebody asks what you will measure. What is the honest answer, and what does it cost you to give it?
Check your answer
The honest answer is that you can measure engagement with the recommendations — how many are viewed, how many are dismissed, how many are followed as far as your own product can observe — and that none of those is a financial outcome, because the product hands off before anything happens. Whatever the user does next occurs somewhere you cannot see.
So the measurement you can actually get is a mechanism measurement with a gap at the end of it, and the case should say so: here is where my observation stops, and here is what somebody would have to instrument — a confirmation step, a partner integration, a survey — to close it.
What it costs is the good version of the story. You do not get to say the surface drove revenue. What you get instead is a person who now believes your other two cases, because they have watched you decline a claim about work you are proud of. That trade is the entire subject of the last lesson in this course, and it is a better deal than it feels like in the moment.
Hands on
Fill the measurement and falsifier lines
Done when: Every case in VALUE-CASES.md names a measurement, states which rung of the ladder it sits on, records the baseline value or explicitly notes that no baseline exists, and carries a falsifier with a number, a date and a named consequence.
- For each case, list three candidate measurements and mark each with its rung: already reported, queryable, ask a person, needs instrumentation, needs an experiment. Three, not one — the comparison is what makes the choice defensible.
- Pick the highest-value candidate that sits on rungs one to three, and write in one line why you did not pick the better one above it. “A holdout would settle this and nobody will approve one” is a perfectly good line and tells a reader you understood the trade.
- Write the baseline: the current value, with the date you obtained it and who gave it to you. If the change already shipped and no baseline exists, write “no baseline — shipped before this was measured” and leave it visible. That admission is doing work.
- Write the falsifier as a sentence with a number, a date and the action you will take. Then apply the test: is this a result you genuinely think could happen? If not, the threshold is set where it cannot embarrass you, and it needs moving.
- Reread all three cases for past-tense verbs attached to unmeasured numbers. Change every one to the conditional. This takes two minutes and is the highest-yield edit in the file.
- Bring them into the chat. I will go after the falsifiers first, on one question: what would have to be true for this threshold to be crossed, and do you actually believe it could be?
What this does not cover
This lesson took cost as given, and the cost field on at least one of your cases still has a blank displacement line. Nothing here addressed the question that gets asked out loud in the room: if this is worth funding, what comes off the plan to fund it? That is the next lesson.
And nothing here has yet confronted the thing that makes measurement hard rather than tedious. Even a clean number, obtained from a real baseline on a real schedule, does not tell you that your change caused it. The attribution lesson closes the course on exactly that gap.
Read this next — primary source
Three Myths About Calculating the ROI of UXKate Moran, Nielsen Norman Group, 6 September 2020 — free; NN/g sells UX training, certification and consulting, so it is an interested party in the question of whether UX investment can be justified
Worth reading because it argues the opposite of what you might expect from a firm that sells UX services: that ROI calculations for design are estimates rather than forecasts, that they do not have to be monetary, and that a rough one is usually enough. This lesson agrees with the first two thirds of that and parts company on the last stretch, for reasons the prose sets out. Read it for the framing, then read the pushback here, and notice that both positions can be held honestly at once — an estimate is legitimate, and presenting an estimate as a measurement is not.
Stuck, curious, or think this lesson is wrong? Ask your teaching agent. The lessons are the scaffold; the conversation is where the learning gets unstuck.