Running it on a real product
Score a product you can only see from the outside, in one sitting, and find out which questions your rubric cannot answer without access — the broken questions are the deliverable.
The instrument is finished on paper. A named reader, a bounded page, nineteen questions with written thresholds, an estimate with its assumptions showing, and a disputed finding with a rebuttal beside it. None of that has met a real product.
This lesson is the meeting. You pick a target, freeze an evidence pack, score the whole rubric in one sitting, and write down every place it broke. The scored page is not the deliverable and never was. The list of questions that broke is the deliverable, because it is the only part you cannot get by designing more carefully at a desk.
This course does not pick your target
There is no worked example here and no named company scored for you, and the omission is deliberate rather than unfinished.
Two reasons, both this course’s own. The first is pedagogical: a target chosen in advance quietly reshapes the rubric around whatever that target happens to expose, and you would end up with an instrument tuned to one company’s public surface. The rubric decides what evidence it needs first. The target comes second.
The second is the same rule the course has been applying to itself for twelve lessons. Publishing a colour-coded verdict on a named, real company — assembled from outside-only evidence, with no access, no interviews and no right of reply — would be precisely the unevidenced public claim this whole instrument is built to avoid making. Score one privately. Do not publish the page.
Choosing a target that can actually be scored
The selection rules below are the course’s, not anyone’s published methodology. They exist so that a failed question means your rubric broke, rather than meaning you picked a product with nothing to read.
- A public web app you can reach without a sales call. A free tier, a trial, or a product surface that renders before authentication. Without this, most of the front-end debt and graftability questions have nothing to resolve against.
- Public documentation. This is where the API surface question on the agentic-readiness axis lives, and where accessibility claims and version statements tend to be made in checkable form.
- A public changelog or release notes. Cadence is almost impossible to read from a static snapshot and almost trivial to read from a dated list. It is also the cheapest evidence of whether the product is actively maintained.
- Not your employer, and not something you use daily. If you know the product well you will score from memory, and the scores will be better than your rubric deserves. The instrument is being tested here, not your knowledge.
Your uncurated sample is still a biased one
The constraints lesson made the case that a front-end diligence survives losing repository access better than any other workstream, because the product is running on the public internet and nobody can cherry-pick which screens you open. That is true and it is the reason this axis is worth having.
It is not the whole picture, and the missing half matters today. The publicly reachable surface is the marketing site, the sign-up flow and the free tier — which is to say, the part of the product the company has the most reason to keep polished. Your sample is not curated for your visit, but it is systematically drawn from the best-maintained screens. Score what you can see, and write that bias into the “what I could not see” row rather than pretending the outside view is neutral.
Freeze the evidence, then score
The order matters and it is the same discipline the red/amber/green test lesson established. Gather first, score second, and do not go back for more evidence once scoring has started.
The reason is that returning to the evidence mid-scoring is never neutral. You will go back for the questions whose answers you dislike, and not for the ones that came out how you expected, which quietly converts your rubric into a machine for confirming first impressions. Freezing the pack also makes the second-reader test possible later, since somebody else can score the identical material.
The pack is: the screens you reached, saved rather than remembered; the documentation URLs; the changelog entries with their dates; the job postings, which are the cheapest public evidence about a stack and a team; and whatever the rendered front end tells you from outside. Write down what you tried to get and could not. That list is not an apology, it is the row that tells the reader how much weight the rest of the page carries.
One sitting, and what that constrains
The time box is fixed at one sitting because the environment the instrument runs in is time-boxed by somebody else, and a rubric that only works when you can keep going is a rubric that has never met the constraint it was designed for.
There is no published figure for how long a front-end diligence should take, and this course is not going to invent one. Every duration in this literature is self-reported by a firm quoting for the work. So set your own box, write down what it was before you start, and treat overrunning it as data about the instrument rather than as permission to keep working. A rubric that needs three sittings is telling you something specific: too many questions need evidence that takes a session to gather.
Retrieval check
Halfway through the run you realise question 1.3 — design system adoption — cannot be answered from anything you can reach. What do you write, and what do you do next?
Check your answer
Write not assessed, and resist every instinct to round it into amber. Amber tells the reader you looked and found something middling. Not assessed tells them a true and separate thing: this question needs access nobody granted.
Then keep going. Do not stop to redesign the question mid-run, because a rubric edited while it is being scored is no longer the rubric you were testing, and you lose the clean result. Log it and carry on.
Afterwards, ask which of three things happened, because the fix is different for each. It may be genuinely unanswerable without the repository, in which case the question is fine and belongs below the cut line as a second-stage question. It may be that you could not tell which side of the threshold you were on, in which case the threshold is ambiguous and needs rewriting. Or it may be that you got a clean answer and it changed nothing — which is the one people never catch, and it means the question should be deleted rather than fixed.
Three ways a question breaks
That taxonomy is the course’s own and it is worth having in front of you during the run, because in the moment all three feel identical and they have opposite fixes.
| Break | What it looks like | What to do |
|---|---|---|
| Needs access | The evidence exists but is inside the company. Nothing you can reach resolves it. | Keep the question, mark it not assessed, and move it below the cut line as a second-stage question. |
| Ambiguous threshold | You have the evidence and still cannot tell which side of the line it falls on. | Rewrite the threshold using this exact case as the worked example. The question is fine; the sentence is not. |
| Answerable and useless | You got a clean answer quickly and it would not change a price, a closing condition or a 100-day item. | Delete the question. It is passing the evidence test and failing the decision test. |
Expect the third to be the rarest thing you find and the most valuable. Questions that are easy to answer feel productive, which is exactly why they survive on a rubric long after they have stopped earning their place on a page you already agreed to bound.
What the run is actually evidence of
Be careful with the claim you make about this afterwards, because overclaiming here is expensive in a room of people who do this professionally, and the honest version is strong enough.
What you will have done: scored one product, from outside only, with no access, no interviews, no data room, and no counterparty. What you will not have done: run a diligence. There was no deal, no reader with money at stake, no management team with an interest in your answer, and none of the three constraints that make the real thing hard were actually applied to you — you simulated one of them and chose your own clock.
The defensible sentence is about the instrument, not the engagement: I built a front-end and UX diligence rubric, ran it end to end against a live product from outside, and here are the questions it broke on and what I changed. That is checkable, it is unusual, and it is true.
Hands on
Run the rubric end to end
Done when: RUBRIC.md is scored in full against one real product from a frozen evidence pack, every question carries a mark including not assessed, and a Broken questions section records each break classified as needs-access, ambiguous-threshold or answerable-and-useless with the change you made.
- Pick a target against the four selection rules: a reachable public web app, public documentation, a public changelog, and not something you know from the inside. Write the target and the reason you chose it into the cover block of
learning/ux-diligence/RUBRIC.md. - Set your time box and write it down before you start. Also write down the time now. This is the only way you will find out honestly whether the instrument fits the constraint it was designed for.
- Gather the evidence pack and freeze it: saved screens, documentation URLs, dated changelog entries, job postings. Record what you tried to get and could not. Do not score anything yet.
- Score all nineteen questions in order, against the thresholds exactly as written. No going back for more evidence. Where a question breaks, log it in one line and keep moving — the run has to finish for the result to mean anything.
- Fill the cover block honestly, especially the what I could not see row, and include the sampling bias: you scored the publicly reachable surface, which is the part the company has the most reason to keep polished.
- Assemble the remediation estimate from the ambers and reds only, grouped into workstreams, with the assumed team and sequencing printed beneath. Then write one disputed finding with its rebuttal and its falsification line.
- Now add a Broken questions section. Every break gets a classification — needs access, ambiguous threshold, or answerable and useless — and the change you made in response. Deleting a question is a legitimate change and should appear at least once if the run was honest.
- Check the page still fits on one page after everything you learned. It probably does not. Cut it back, and add what you cut to the cut list with a reason.
- Bring the broken-questions list into the chat, not the scores. I will ask which classification you got wrong, because at least one break you filed as needing access is almost certainly an ambiguous threshold you could have resolved.
What this does not cover
This is the last lesson, and the thing that comes after it is not another lesson. It is the transfer test, and it is the only one that settles whether you built an instrument or a memo: hand the page to somebody else, have them score a second product with it, and see whether the colours broadly agree. Nothing in this course can do that for you, because the whole point is that it happens without you in the room.
The rubric template, with the thresholds written out per question and the assembly order for the estimate, lives at /ux-diligence/reference/rubric and is meant to be reopened rather than read once. Your filled-in copy is learning/ux-diligence/RUBRIC.md.
And the boundary stays where it has been since the first lesson. This covers the front-end and UX half of one technical workstream. The commercial, financial, legal, tax and security diligences are real, larger, and staffed by other people. Say so on the cover of anything you produce from this, because a scope statement is what stops four axes being read as a verdict on a whole company — and because the fastest way to lose a room of people who run deals is to imply you were running one.
Read this next — primary source
Technical Due Diligence Best PracticesMarty Abbott, AKF Partners, January 23, 2018 — a firm that sells technical due diligence, arguing against its own interest about what its own product cannot see
Read it as the honest account of sampling in this work. It concedes that the material a reviewer gets is chosen by the company, and that a chosen sample is not a representative one. You are about to run into the mirror image of that problem: nobody curates your evidence for you, and the public surface you can reach is still a biased sample, just biased in a different direction. Read it whole so you can name which bias you are working under on the day, rather than discovering it in the room.
Stuck, curious, or think this lesson is wrong? Ask your teaching agent. The lessons are the scaffold; the conversation is where the learning gets unstuck.