UX maturity, read from the outside
Research cadence, a design system and its real adoption, consistency debt and accessibility posture are the four proxies this course reads from the outside — an operationalization of UX maturity, not the four factors any published model names, because the model’s own method (interviews, org-wide surveys) needs access a diligence never gets.
You can open the target’s product this afternoon. You cannot interview their designers, you cannot survey their engineers, and the one UX person you will be introduced to has been chosen by the person selling the company. That is the whole problem with this axis, and it is why it comes first: it is the axis where the gap between the published method and the access you have is widest.
The temptation is to reach for a named model, put its label on your page, and let the name do the work. Do not. There is a published model, it is good, and you cannot run it. What you can do is say exactly what you are substituting for it and why — which is a stronger position than borrowed authority, because borrowed authority collapses the moment someone in the room has read the source.
What the published model actually says
The Nielsen Norman Group’s UX Maturity Model defines six stages, in this order: Absent, Limited, Emergent, Structured, Integrated, User-Driven. Those are the names. Learn them in order, because there is a trap waiting.
NN/g evaluates those stages across four cross-cutting factors: strategy, culture, process and outcomes. Note what is not on that list. Research cadence is not a factor. Design systems are not a factor. Accessibility is not a factor. The model is about how an organisation is built, not about which artifacts it has shipped.
The version trap, and it is checkable
NN/g still hosts a superseded eight-stage model from 2006 with completely different stage names, running from “Hostility Toward Usability” up to “User-Driven Corporation.” Both pages are live on the same site. Citing the 2006 stage names as the current model, or blending the two lists, is the kind of error someone verifies in thirty seconds on their phone while you are still talking.
The other bias worth stating out loud before you lean on any of it: NN/g sells a paid UX Maturity Assessment. The free model and the free quiz are the top of that funnel. That does not make the model wrong — it is the best-documented one available — but a firm publishing a maturity model it also sells assessments against has an interest in maturity feeling assessable.
Why you cannot run it
NN/g is explicit about what a real evaluation takes. A complete assessment, they write, should be based on diverse assessment methods — observation of and interviews about work practices, analysis of processes, people and tools, assessment of deliverables, and surveys of people from across the organisation.
Read that list against the three constraints. Interviews about work practices: you get one scheduled call with someone management picked. Analysis of processes and tools: that is a data-room request you will not win before exclusivity. Surveys across the organisation: nobody is letting a prospective buyer survey their staff during a live deal. Three of the four methods the model calls for are unavailable to you by construction.
NN/g’s companion piece on assessing UX maturity informally is closer to your situation, and it carries a warning worth turning directly on the thing you are building: overassessment risks maturity becoming “a scoreboard rather than a healthy system”. A diligence scorecard is a scoreboard. That is its job, and the criticism still lands: the colour you print is a decision input for one transaction, not a diagnosis of the company’s health, and writing it as though it were the latter is how a diligence finding becomes an insult.
What this course substitutes, and whose idea it is
So the rubric does not score NN/g’s four factors. It scores four proxies chosen for one property the factors do not have: each one is visible from outside the company.
- Research cadence. Is there any, and how often?
- A design system, and its real adoption. Two questions, not one: does it exist as a shipped artifact rather than a Figma file, and does the product actually use it?
- Consistency debt. How many ways does this product do the same thing?
- Accessibility posture. Claimed, tested, or neither?
The evidence, proxy by proxy
Research cadence
Uncurated evidence first. Job postings for researchers, and how long they have been open. A public research or design blog, and the date of the last post. A beta or community programme that recruits participants. Changelog entries that describe a change in terms of what users were doing rather than what shipped. None of that was written for your visit, and the absence of all four across a large product is itself a finding.
The design system, and its adoption
Existence is cheap to check: a public documentation site, a published package, a Storybook. Adoption is the hard half, and it is where an interested party’s estimate will be offered to you.
The one published measurement method worth borrowing comes from Pinterest’s design systems team, writing on Figma’s blog — a vendor whose paid analytics product the method happens to need. The formula is design-system component layers divided by total layers, scoped to recently edited files marked ready for handoff. Their worked example is seven system layers out of fifteen total, which they report as 46.6%.
There is no adoption benchmark. Do not invent one.
The Pinterest post declines to name a target, in its own words: it is hard to say there is an ideal number for adoption. When it mentions one team at 50% against another at 2%, it is illustrating a spread hypothetically to argue against absolute targets, not reporting its own measured results for named teams. Quoting those figures as Pinterest’s data is a misread of the source.
Nothing else in this research found a published adoption benchmark either. So the qualitative point is the finding: the spread between teams is the signal. One product area on the system and another ignoring it tells you about governance, which is what you actually wanted to know. A single percentage tells you nothing without a target nobody has published.
Consistency debt
Brad Frost’s interface inventory is the method: sweep the product and collect every instance of each category — buttons, form fields, navigation, icons, and a dozen more. Frost sells consulting and books; the method itself is free and he publishes no counts and no thresholds, so there is no statistic to attribute to him. What you get is a picture, and a picture of nineteen buttons is persuasive without a number.
For something closer to objective, their stylesheet is downloadable without asking anyone. css-analyzer (MIT, open source) reports over 200 metrics with a total and a totalUnique per category. The ratio between them is the closest thing to hard consistency-debt evidence you can obtain from outside: fifty-one unique font sizes in a production stylesheet is a fact, not an opinion, and nobody chose to show it to you.
Accessibility posture
Three states, and the distinction is the whole question. Claimed is an accessibility statement on the site. Tested is a claim that names a version and a level. Neither is silence.
WCAG 2.2 is a W3C Recommendation of 5 October 2023, revised 12 December 2024, with conformance levels A, AA and AAA. So “WCAG compliant” with no version and no level is not a checkable claim, and you can say so without having tested anything. WCAG 3.0 exists as an early draft and is not a Recommendation; a target citing it is citing a draft.
Retrieval check
The target’s marketing site says “our platform is fully accessible and WCAG compliant.” You have not run a single automated check. What can you already write in the Findings table?
Check your answer
That the claim is unfalsifiable as stated. WCAG has versions and three conformance levels; a claim naming neither cannot be checked, agreed with, or disagreed with. That is a finding about how the company makes claims, and it costs you nothing to write.
What you must not write is that they are non-compliant. You have not tested. The mark for question 1.5 is not assessed on conformance, with the unfalsifiable claim recorded separately as evidence about posture. Two different things, two different rows, and collapsing them into an amber is exactly the rounding the rubric forbids.
What would move the score
Every question on this axis needs an answer to “what evidence would change this colour,” because that sentence is what makes a disputed finding survivable. For adoption, it is a rendered screen using a component that is not the published one. For research cadence, it is a dated artifact — a post, a study, a recruitment page. For consistency debt, it is the uniqueness ratio out of their own stylesheet. For accessibility, it is a conformance claim carrying a version and a level.
And be honest about the ceiling. NN/g’s own state-of-maturity survey drew 5,371 respondents and NN/g attaches three caveats to it: the sample is self-selected and not representative, low-maturity organisations are likely under-captured, and an individual’s assessment is not the best way to determine organisation-wide maturity. That last caveat applies to you. You are one individual, with less access than an employee, scoring an organisation from its output.
Check your recall
Answer from memory — no scrolling back.
Hands on
Fill axis 1 with evidence you could actually get
Done when: All five questions under Axis 1 in RUBRIC.md have a named evidence source and a written threshold, and the axis carries one line stating that the four proxies are yours rather than NN/g’s.
- Open
learning/ux-diligence/RUBRIC.mdat Axis 1. Under the axis heading, write the one-sentence attribution: the anchor model, its six stage names, and the fact that these five questions are your own outside-observable substitute for its four factors. - For each of 1.1 to 1.5, fill the Evidence I would need column with a specific artifact you could open today — not a category. “Their published component documentation site” rather than “design system evidence.”
- Write the Threshold column for 1.5 first, because it is the easiest: green needs a conformance claim naming a WCAG version and level, amber is a claim without them, red is a product with visible barriers and no claim at all. Notice how much of the work is just refusing to accept an uncheckable claim.
- Now write the threshold for 1.3, adoption. This one will fight you, because there is no benchmark to anchor to. Write a threshold in terms of spread and governance rather than a percentage, and if you cannot make it reproducible, write not assessed into the threshold cell as the honest default.
- Pick a real product you use daily. Download its main stylesheet, run
css-analyzerover it, and record the unique-to-total ratio for colours and font sizes. Fifteen minutes, and you now have a worked number and know how long the method takes under time pressure. - Bring the filled axis into the chat. I will pick the weakest threshold and ask you to score two different products with it, which is the only test that matters.
What this does not cover
This axis asks whether anyone here knows what their users do. It says nothing about what it would cost to change the interface once you own it, which is a question about frameworks, language settings, test coverage and dependencies rather than about design practice. That is the front-end debt lesson, and it is the easiest of the four to score defensibly, because the facts behind it are published with dates on them.
The thresholds you started writing here also do not yet meet the test they will eventually have to pass, which is that a second reader applies them and lands on the same colour. That test gets its own treatment in the red/amber/green lesson, in the module on the scorecard as a product.
Read this next — primary source
The 6 Levels of UX MaturityKara Pernice, Sarah Gibbons, Kate Moran & Kathryn Whitenton, Nielsen Norman Group — published June 13, 2021, updated January 24, 2024; free, and the top of the funnel for NN/g’s paid UX Maturity Assessment engagement
Read it whole for two things this lesson only samples. First, the stage names in order and what behaviour NN/g attaches to each, because naming the model you are anchoring to and getting its stages right is the difference between a defensible score and a vibe. Second, the methodology paragraph, which tells you plainly that a complete evaluation needs interviews, process analysis and organisation-wide surveys. That paragraph is the reason this lesson exists in the shape it does: you are not going to run their model, so you need to be explicit about what you are running instead.
Stuck, curious, or think this lesson is wrong? Ask your teaching agent. The lessons are the scaffold; the conversation is where the learning gets unstuck.