The attribution trap
Claiming a retention number for a UI change that plausibly did nothing buys one good quarter and spends the credibility of every case you write afterwards — the failure mode this whole course exists to prevent.
Retention went up this quarter. Your change shipped in the same window. Somebody in the room, not hostile, genuinely curious, asks whether that was you. The sentence is already forming, and it is a good sentence, and everything in your professional situation rewards saying it.
This is the lesson the course was built to reach. Everything before it — the fund, the P&L lines, the metric definitions, the four-field page — exists so that you can decline that sentence for a reason you can state out loud rather than from vague caution.
What attribution actually requires
A causal claim is a claim about a world that did not happen: what the number would have been if you had not shipped. That world is not available for inspection, so every honest attribution is an argument about how well you have approximated it. There are three ways, and only three.
- Randomisation. Assign the change to some and not others, by chance. The untreated group is your counterfactual. This is the only method that works without you having to be right about anything else, and it is the one that requires somebody senior to agree that a set of customers gets the worse version on purpose.
- A comparable untreated group. Not random, but similar: a segment, a region, a product line that did not get the change. Weaker, because the groups may differ in ways that also affect the outcome, and much more often obtainable.
- A mechanism tight enough to leave no room. The chain from change to number is short, every link is separately observed, and no other cause plausibly produces the same signature. This is how you attribute a drop in a specific error class to the gate that intercepts it. It does not scale up to retention, because retention has too many doors into it.
Notice which method is not on that list. “The number moved after we shipped” is not an approximation of the counterfactual. It is the absence of one.
Everything else that happened in your window
Before claiming a quarter, write down what else was true of it. Not as a rhetorical exercise — as a list, because the list is what you will be asked for and it is much less frightening on paper than in your head.
- Other releases. Including ones from your own team. Two changes shipped in the same quarter cannot be separately attributed by looking at a quarterly number, and the other one may have a better claim than yours.
- Commercial motions. A pricing change, a packaging change, a renewal push, a new customer success playbook, a compensation change for account managers. Any of these can move retention on their own and most of them are designed to.
- Cohort mix. The customers up for renewal this quarter are not the same customers as last quarter. If the sales team changed who it sold to eighteen months ago, that arrives now, wearing your change’s clothes.
- Seasonality and renewal timing. Contracts cluster. A quarter heavy with multi-year renewals looks different from one heavy with monthly accounts, and the difference has nothing to do with the product.
- Survivorship. The accounts still present to be measured are the ones that did not leave earlier. Metrics computed on survivors drift upward for reasons that are not improvements.
- Outside events. A competitor’s outage, a competitor’s price rise, an acquisition, a regulatory deadline that made switching impossible this quarter.
You are not expected to rule all of these out. You are expected to know they exist, and to have looked. The difference between a case that survives a data team and one that does not is rarely the strength of the evidence. It is whether the author had already found the objection.
Divide before you claim
Here is a habit worth more than any framework in this course, and it takes about four minutes with a calculator.
Do not defend your claim. Derive what your claim implies about a part of the business you can look at, and then go look. If the implication is false, you are finished arguing and you have found out cheaply, in private, before a room full of people found out for you.
Where people get burned
Every number in the next four paragraphs is invented by this course to make the arithmetic legible. It is not a benchmark, it is not any company’s reported figure, and nothing in it should be repeated as a fact about anything. The assumptions are stated so that you can rebuild the same calculation with your own real numbers, which is the only version of it that means anything.
Stated assumptions. A business with 400 accounts at the start of the year. Gross retention measured by account count rather than by revenue, purely because counting accounts is easier to follow — a revenue-weighted version would need each account’s ARR, and the metric-set module has already warned you that these two are different calculations with the same name. Gross retention last year: 91%. This year: 93%. Your change shipped to the 120 accounts using the extraction product, and to nobody else.
ILLUSTRATION — invented figures, stated assumptions
Accounts at start of year 400
Accounts using the changed product 120 (30%)
Gross retention, prior year 91% -> 36 accounts lost
Gross retention, this year 93% -> 28 accounts lost
--------------
Difference 8 accounts
If ALL of that came from the treated 120:
Treated losses, prior year 120 x 9% = 10.8 accounts
Treated losses, this year 10.8 - 8 = 2.8 accounts
Implied treated retention 1 - 2.8/120 = 97.7%
So the claim implies: retention inside the treated 120 rose
from about 91% to about 97.7% — nearly 7 points — while the
other 280 accounts stayed exactly flat.That is the whole technique. A two-point move across the book, if it came entirely from the thirty percent of accounts you touched, is a nearly seven-point move inside that segment. Now you have a question you can actually answer: did retention inside those 120 accounts rise by seven points? Somebody can pull that. If it rose by two points, the same as everywhere else, then whatever moved retention moved it everywhere and it was not your change.
Be precise about what the calculation can do. It can only ever fail your claim, never confirm it. If the segment did move seven points, you have removed one objection and every confounder in the list above is still standing — the treated segment might be the segment the new customer success playbook was piloted on. The arithmetic is a filter, not a proof, and it is valuable because it is a filter you can run on yourself before anyone else runs it on you.
There is one more thing to read off it. The entire two-point move is eight accounts. Eight. One large customer’s reorganisation, one account manager’s territory, one delayed renewal that slipped into the next quarter — any of those is the size of the whole effect you are proposing to take credit for. When the effect is that small relative to ordinary variation, the honest word is not “caused.” It is “consistent with,” and it is weaker on purpose.
The statistic you will be handed, and why it is not here
At some point in this argument somebody will offer you a rescue: a large headline figure showing that design pays. Two families of these dominate. This course went looking for both and prints neither number, and the reason is worth more to you than the figures would have been.
The first is a widely-repeated ratio for the return on money spent on user experience, attributed in every secondary repetition found to a large analyst firm. Searching for the underlying report by title, author, year and sample turned up no primary publication — only further blog posts and vendor pages repeating the same sentence, none of them citing a report you could open. A critique of the figure exists, published by Chuk Moran under the title “ROI from UX: Fake Facts”; it returned an HTTP 403 when this course tried to read it, so nothing is quoted from it and no argument is attributed to it beyond its title. The figure itself is not printed, because a number whose source cannot be found is not a number.
The second is an index claiming that design-driven companies outperform a market benchmark over ten years. Two things about it. It is published by the Design Management Institute as its Design Value Index, and it is routinely misattributed in secondary write-ups to McKinsey — a different organisation, a different report, and the one the P&L lesson already declined to quote. Separately, DMI’s own page describing the index’s methodology returned an HTTP 403 when this course attempted it, so its sample, its selection criteria for “design-driven,” and its index construction could not be checked. The percentage is therefore absent here for the same reason the McKinsey percentages are absent from the P&L lesson.
Underneath both sits the objection Mauro and Thurman make about the McKinsey report: that findings of this shape are “both inaccurate and unsupportable” as research, and more generally that an index built on correlation between design practice and financial performance cannot carry a causal claim. Their firm sells usability research, so they are an interested party arguing about whose evidence should count — and the structural point survives the conflict anyway, because it is the same point you just applied to your own quarter. Companies that invest in design are different from companies that do not in a hundred ways. The index has no counterfactual either.
What the second quarter costs
Suppose you make the claim. The quarter goes well. Here is the part nobody describes, because it is invisible from where you are standing.
Somebody eventually splits the data by segment — not to catch you, but because a finance or data team splits everything by segment eventually. The claim does not replicate. Nothing dramatic follows. No meeting is called. What happens is that the next case you write gets read differently: the reader now applies a discount before they reach your mechanism, and the discount is permanent because they have no way to know when you have stopped doing it.
And you never find out. That is the asymmetry that makes this trap worth a whole lesson. Overclaiming has an immediate, visible reward and a delayed, invisible cost, which is the exact structure of every habit that is hard to break. The only defence available is a rule you follow before you know whether it would have mattered.
The rule this course proposes, in the course’s own words: make the smallest claim your evidence actually supports, name what you are not claiming, and say what would show you were wrong. Then, when a claim of yours does hold up, it arrives at full weight.
Check your recall
Answer from memory — no scrolling back.
Retrieval check
Rewrite “the review gate improved retention” into the strongest claim the evidence in your case file actually supports — and say what you gave up.
Check your answer
Something close to: the gate intercepted this many low-confidence extractions last quarter and operators corrected this share of them, so that many errors did not reach a downstream system. Each of those was previously capable of becoming a support contact. I have not measured how many became one, and I am not claiming a retention effect: I have no segment split for the treated accounts, so nothing available to me separates my change from everything else that happened this quarter.
What you gave up is the headline. What you got is a sentence in which every clause is countable, the inference is visible at the end where the listener can price it themselves, and the disclaimer is doing more work than the claim — because it tells the room you went looking for the objection before they did.
If it helps, notice that the abandoned version was never actually available. You could have said it, but you could not have defended it, and the version of you that says it is the version whose next three cases get discounted.
Hands on
Audit all three cases, and strike one claim
Done when: Every case in VALUE-CASES.md carries a confounder list for its measurement window, has had its strongest claim reduced to what the evidence supports, and at least one claim you wanted to make has been struck out and replaced with what you would need in order to make it.
- For each case, write the confounder list: what else was true of the measurement window. Aim for five. If you cannot reach five, you have not asked anyone outside your own team what happened that quarter.
- Run the division. For the case with the strongest outcome claim, work out what that claim implies about the segment your change actually reached, and write the implied number down. If you cannot get the segment split, write who you would ask for it.
- Find the claim you most want to be true and strike it out visibly — leave it on the page, crossed through. Underneath it write the two or three things you would need in order to make it honestly. That crossed-out line is the most valuable thing in the file and it is the one you will be tempted to delete.
- Rewrite each hypothesis to the smallest claim the evidence supports, and check that the Not claiming line still matches — it usually needs to grow after this exercise, not shrink.
- Say the hardest one out loud: a case about work you are proud of, concluding that you have no financial claim yet, followed by what you would need to build one. That is the outcome this course exists to make sayable, and saying it in an empty room first is how it becomes available in a full one.
- Bring all three back. I will take the position of a data team that does not know you and has no reason to be generous, and I will start with whichever claim you kept.
What this does not cover
This lesson has been about the causal argument and not about the arithmetic of measuring properly — sample sizes, confidence intervals, how long a test has to run, what a difference-in-differences design does. That is a statistics course, and it is deliberately outside this one: the failure this course was built to prevent almost never happens because somebody mis-specified a model. It happens because nobody asked what else could have caused it.
The metric definitions behind every number you might claim are in the metric-set module, and the four P&L routes a change can travel are in the earlier lesson on which line you move. From here the thing to keep open is the translations reference — design-language claims paired with their operating-language versions and the mechanism that makes each pairing honest. It is written to be reopened before a conversation, not read once.
And the exercise does not end with the file. The last row of that reference exists for the case you cannot make yet. Bring that one back first.
Read this next — primary source
Should McKinsey Design Retract Its “Business Value of Design” Report?Charles L. Mauro and Paul W. Thurman, Mauro Usability Science — free; the firm sells human factors engineering and usability research, so it is an interested party in the question of whose evidence for design’s value counts
You met this in the P&L lesson as a caution about a single famous report. Read it again here for the general argument underneath the specific complaint: that an index built on correlation between design practices and financial performance cannot carry a causal claim, however large the correlation and however respectable the publisher. That is the same objection a data team will raise about your case, applied at industry scale instead of to one product. Read the critique and the conflict of interest together — the authors sell competing evidence — because holding both at once is the exact skill this lesson is asking for.
Stuck, curious, or think this lesson is wrong? Ask your teaching agent. The lessons are the scaffold; the conversation is where the learning gets unstuck.