The red/amber/green test
A colour is only worth printing if two people applying your rubric to the same evidence would land on the same colour — which means writing the threshold down before you score, and refusing the numeric precision you have not earned.
The rubric now has thresholds under every question. Green means no remediation is implied. Amber means remediation is implied and bounded. Red means it is large, or its size is genuinely unknown. You wrote them, you understand them, and applying them feels obvious.
That feeling is worth nothing, and it is the specific thing this lesson is trying to take away from you. A threshold that only its author can apply is a preference with a colour swatch beside it. The instrument you are building is meant to survive being handed to somebody else, which means the only question that matters about a colour is whether a second reader, given the same evidence, would print the same one.
The thing you are claiming has a name
“Two people applying this to the same evidence would agree” is not a new idea and not the course’s idea. It is inter-rater reliability, and there is a real measurement literature behind it, none of which comes from anyone selling due diligence.
The concept transfers cleanly. The methods do not, and knowing why is the useful part. In that literature, reliability is quantified — Cohen’s kappa for categorical marks like yours, other statistics for scales — from many ratings made by trained raters on sampled material. The central caution in McHugh’s methods paper on the kappa statistic is that raw percent agreement overstates reliability, because some agreement happens by chance, and how much depends on how the categories are actually distributed.
Sit with that against your four marks. Suppose two people threw darts at green, amber, red and not assessed, all equally likely. They would land on the same mark a quarter of the time, having read nothing. Your marks are nowhere near equally likely — amber is where uncertain evidence goes, so amber will be over-represented — which makes the real chance-agreement rate different from a quarter and unknown to you, because you have not measured it and are not going to.
Write the threshold before you look
The operational rule falls straight out of the problem. A threshold written after you have seen the evidence is a description of what you already decided, and it will fit that evidence perfectly and nothing else.
So the sequence is fixed: write the threshold, freeze it, then go and get the evidence. When the evidence turns out to sit awkwardly across the line you drew, that is not a failure. That is the instrument telling you something, and the honest move is to score against the line as written and note the awkwardness — not to redraw the line so the answer comes out where your gut wanted it.
A threshold that a second reader can apply has three properties.
- It names evidence, not quality. “A versioned package other teams install, with a changelog and more than one consuming application” is checkable. “A mature design system” is a mood.
- It says what the evidence has to show, not how much. Amounts invite the precision you have not earned. Presence, absence and count-of-distinct-things survive being challenged.
- It has a stated behaviour when evidence is missing. Otherwise every reader improvises one, and they will improvise different ones.
Not assessed is the third state, and it is ours
The fourth mark on the rubric is a deliberate design choice by this course, and it is worth separating from everything above. The reliability literature is about agreement between raters who are actually rating. It has nothing to say about when a rater should decline to rate. That rule is the course’s, invented for a specific failure that the constraints lesson set up.
The failure is that amber absorbs everything. It is where a genuine bounded problem goes, and it is also where “I could not find out” goes if you let it, and those are entirely different messages to a buyer. One says fund this work. The other says your deal team negotiated access that did not cover this. Collapsing them into one colour destroys the second message completely, and the second message is sometimes the more valuable of the two.
There is a reliability argument for keeping them apart, too. Every mark you round into amber makes amber more dominant, and the more dominant a category is, the more of your apparent agreement with a second reader is just both of you defaulting to the same bucket. A page where most questions came out amber is not a reliable page. It is an uninformative one wearing the same colours.
Why there are no scores out of ten
The pressure to make the page more precise will come from the reader, not from you. Numbers look rigorous, they sort, and they can be put in a deck next to other numbers. Refuse anyway, and have the reason ready.
Numbers invite averaging, and averaging is where a real finding goes to die. Score four axes out of ten, get eight, eight, eight and a two, and the arithmetic gives you twenty-six over four, which is six and a half. A middling overall six and a half, and the two — the one axis that was going to cost the buyer something — has been dissolved into it. Four separate colours cannot be averaged, which is not a limitation of the scale. It is the feature.
The second reason is the one this whole lesson has been building to. A ten-point scale asks a second reader to reproduce your judgement to within a tenth of the range. Four marks with written thresholds ask them to reproduce a category. You have some chance at the second. You have none at the first, and claiming otherwise is a precision claim you cannot support from the evidence a diligence actually gets.
The evidence has to be frozen too
A second-reader test where each reader gathers their own evidence tests nothing useful. If you disagree, you will not know whether the threshold was ambiguous or whether you simply opened different screens, and you will guess wrong about which, because blaming the evidence is more comfortable.
Assemble the evidence once — screenshots, the pages you reached, the documentation URLs, the changelog entries, the job postings — and have both readers score from that identical pack. Then a disagreement is unambiguously about the threshold, which is the thing you are trying to fix.
Check your recall
Answer from memory — no scrolling back.
Hands on
Make a threshold survive a second reader
Done when: Every question in RUBRIC.md has a written threshold naming observable evidence, at least three questions have been rewritten after a second reader disagreed, and the evidence pack both readers used is recorded in the file.
- Go through
learning/ux-diligence/RUBRIC.mdand fill the Threshold column for every question, before you score anything. Each one must name evidence a stranger could go and look at. If you cannot write it without the words “good”, “mature” or “reasonable”, the question is not ready. - Write the missing-evidence behaviour into each threshold explicitly. One sentence: what happens to this question when the evidence is not there. Most will be not assessed; some will be red, because absence is itself the finding. Decide which now, not on the day.
- Pick a product and assemble one evidence pack: the screens you reached, the public documentation, the changelog, the job postings. Save it as a fixed set. Record in the file what is in it and, just as importantly, what you tried to get and could not.
- Score the rubric from that pack. Then hand the pack and the rubric to a second person and have them score independently, without seeing your marks.
- Compare. Count how many of your agreements were both of you writing amber — that number tells you how much of the agreement was real. Then take every disagreement and fix the threshold, not the rater.
- Rewrite at least three thresholds. Note next to each what the second reader read it to mean, because that reading is the ambiguity, and your own version of the sentence will never show it to you.
- Bring the before-and-after thresholds into the chat, along with the amber count. I will push on any threshold where a determined reader could still land on either side.
What this does not cover
Reliable colours still do not tell a buyer what anything costs. An amber says remediation is implied and bounded, and the whole value of that word “bounded” is the number that goes next to it. The remediation-in-weeks lesson is about that number: why it is a claim about an assumed team, an assumed scope and an assumed sequence rather than a measurement, and what has to be printed beside it so that somebody who knows more than you can correct it instead of just disbelieving it.
After that comes the finding management will dispute, which is where a written threshold stops being an internal discipline and becomes the thing you stand on in the room. The last lesson runs the whole instrument against a real product from outside.
Read this next — primary source
Interrater reliability: the kappa statisticMary L. McHugh, Biochemia Medica, 2012 — free full text on PubMed Central; a peer-reviewed methods paper, selling nothing
Read it for one idea carried all the way through: two people agreeing is not evidence that an instrument works, because some agreement happens by chance and the amount depends on how the categories are distributed. That is the argument against the reassurance you will feel the first time a colleague scores your rubric and mostly matches you. The paper is written for clinical measurement, where raters are trained and ratings are sampled in bulk. You will have neither. Read it to understand what a reliability claim requires, then notice how little of that you can honestly claim — which is the point.
Stuck, curious, or think this lesson is wrong? Ask your teaching agent. The lessons are the scaffold; the conversation is where the learning gets unstuck.