The platform corner and the research corner
Apple’s guidance is platform rules with App Review behind them and NN/g’s is small-sample research sold alongside training — two different kinds of claim that get quoted as if they were the same kind.
In the same meeting, two people cite two sources.
The first says Apple’s guidelines advise against showing confidence values unless you know they correlate with quality. The second says NN/g research found users abandon agents that hide pricing. Both are accurate. Both get written into the notes as findings. One is a platform rule with an enforcement mechanism behind it and no stated evidence. The other is an observation from six people.
Neither is worthless and they are not the same kind of claim. Reading each one correctly is the difference between a designer who has read the canon and one who can use it.
Apple: instructions with no visible evidence
There is no “machine intelligence” section in the Human Interface Guidelines. The AI pages are scattered: Generative AI, Machine learning and Siri sit under Technologies, App Shortcuts under Components. The Machine learning page is the prediction-and-recommendation-era document, still live, pointing at the Generative AI page for generative work. It carries nine patterns: Explicit feedback, Implicit feedback, Calibration, Mistakes, Corrections, Multiple options, Confidence, Attribution and Limitations.
The Confidence pattern contains the single sharpest sentence any publisher in this canon has written:
If you’re not sure how your confidence values correlate with the quality of your results, it’s not a good idea to convey confidence to people.
The neighbouring patterns are equally direct. “Be especially careful to avoid mistakes in proactive features.” “Never rely on corrections to make up for low-quality results.” Never. Not a good idea. Be especially careful. This is the grammar of a rule.
Nowhere on the page is there a study, a sample, a methodology, or a citation of any kind. Not a footnote, not a research link, not a “we tested this with.” Apple asserts. That is consistent with what the document is: platform guidance from the company that runs App Review, where much of the HIG has enforcement behind it. A rule does not need to show its working, because compliance is the point.
Get the criticism narrow or lose it
Do not say Apple’s guidance is unresearched or made up. You do not know that, and Apple almost certainly runs research it does not publish. The claim that holds is about the document rather than the company: it carries no visible methodology and no sample, so there is no way to audit how it was reached. Apple publishes no methodology page and no rationale beyond one-line change-log notes — the Machine learning page was consolidated on 2 May 2023 and last touched with minor updates on 8 June 2026.
That narrower version is also the more useful one. It tells you what to do with the Confidence sentence, which is to treat it as a strong prior from an experienced platform team, weigh it against your own measurements, and not mistake its confident tone for a result.
NN/g: evidence with a number attached, and the number is small
The Nielsen Norman Group publishes free research and sells training, consulting and paid reports, advertised on the same pages as the research. That is a real interest and naming it is not a dismissal; it is the same courtesy the course extends to Microsoft, Google and Apple.
The most agent-relevant empirical piece on the site is “Designing AI Agents: 4 Lessons from China’s Qwen Agent” (Feifei Liu and Maria Rosala, 8 May 2026). Its sample is six participants. NN/g says so, in the article, which is the thing to notice: the sample is stated rather than hidden, and stating it is what makes the piece usable.
Its own scope is tighter than an agent designer might hope. The findings are about discoverability, familiar patterns, data-handling transparency and pricing transparency, and the article records that the agent it studied never took an unattended action: “Users still retained control; they needed to authenticate with AliPay before allowing payments.” So even NN/g’s most agent-relevant study is a study of a system with a human gate in front of every commitment. It does not reach autonomous execution or mid-run failure either.
And the pattern worth carrying: the one large sample on the site is not NN/g’s data. The n = 410 in Evan Sunwall’s piece (19 September 2025) belongs to Colombatto, Birch and Fleming, whose study NN/g is summarising. NN/g is not uniformly small-n. It is small-n when the data is its own, and large-n when it is reporting somebody else’s.
Two kinds of claim, and what each can bear
| Apple HIG | NN/g research | |
|---|---|---|
| What it is | Platform guidance, much of it enforced via App Review | Published qualitative studies plus summaries of others’ work |
| Evidence shown | None on the page: no sample, no method, no citations | Sample size stated, method described, own data distinguished |
| Why it is binding | Because Apple can reject your app | It is not binding on anyone |
| Correct use | A strong prior and a shipping constraint on Apple platforms | An observation about a handful of users of one product |
| Misuse to avoid | Quoting it as a research finding | Quoting the n out, or attributing a borrowed n to NN/g |
The asymmetry that makes this worth a lesson: the source with no stated evidence is the one you may be contractually obliged to follow, and the source that shows its work has no power over you at all. People collapse that into “Apple is more authoritative,” which is true in exactly one sense and false in every other.
How this was checked
Both characterisations above are claims about what a page does not contain, so here is what would have falsified them:
- Apple’s Machine learning page was read in full through Apple’s own DocC JSON route, not a search snippet, because
developer.apple.com/design/…is client-rendered and returns only a title to a plain fetch. A hit would have been any citation, sample, methodology note or research link on the page. There are none, and the change log carries dates but no rationale. - The Qwen agent article was fetched for its stated sample. A hit would have been a sample larger than a handful, which would have broken the small-n characterisation entirely. It is six.
- The Sunwall piece was re-read to confirm the n = 410 belongs to the academic study it summarises rather than to NN/g. It does.
Retrieval check
A designer closes an argument with "Apple says not to show confidence values unless you know they correlate with quality, so we are removing our per-field confidence." You disagree. What do you say?
Check your answer
That is Apple’s Confidence pattern and it is quoted correctly. Two things about it. It is platform guidance with no stated sample or methodology — it is a rule, not a finding, and the rule is binding on the App Store rather than on us. And the condition it sets is the operative bit: it says unless you know how your values correlate with quality. So the question for us is whether we know that, and if we do, we are inside the exception rather than arguing with the rule.
Notice what that does not do. It does not call Apple wrong, does not call the guidance unresearched, and does not treat the sentence as optional. It reads the conditional clause that the person quoting it skipped, which is usually where the argument actually is. It also does not claim a measurement you have not made — if nobody has checked your confidence values against corrected ones, the honest answer is that Apple’s condition is unmet and the work is to meet it.
Two kinds of claim
Answer from memory — no scrolling back.
Hands on
Label every source in the position document by kind
Done when: Every framework row in POSITIONS.md has a filled Publisher’s interest cell that states what kind of claim the source makes, what evidence it shows, and what leverage the publisher has over you — each traceable to a page you opened.
- For each of the five rows in
learning/agentic-ux-canon/POSITIONS.md, write the publisher’s interest in one line. Not “big tech company” — the specific interest. Apple sells the platform and enforces this guidance through App Review. NN/g sells training and consulting alongside the research. - Add a second line to each: evidence shown. A sample size if there is one, “none stated” if there is not. “None stated” is a finding and belongs in the cell in those words.
- For the Apple row, fill the disagreement cell with the Confidence pattern verbatim, including the conditional clause. The conditional is the part people drop when they quote it, and having it in your own document means you never will.
- For the NN/g row, pick one finding and write it with the author, the date and the n inline, as a single sentence you could say out loud. If the sentence sounds weaker with the n in it, that is the sentence being honest and you should keep it.
- Bring the Apple row into the chat. The check is whether the criticism is the narrow one — no visible methodology — or whether it has drifted into “unresearched”, which is the version that loses the argument.
What this does not cover
This lesson separates two kinds of published claim. It leaves out a third kind entirely: the pattern catalogue, which is neither a rule nor a study but a record of what somebody observed shipping. Reading one of those correctly — taking the vocabulary and refusing the authority — opens the next module, in the catalogues-are-vocabulary lesson. The module you are finishing named three gaps in the canon; the one ahead is about turning that into a position that survives being questioned, starting from the words a room already shares.
Read this next — primary source
Human Interface Guidelines — Machine learning, the Confidence patternApple — fetched 2026-09-05 via Apple’s own DocC JSON, since developer.apple.com returns only a page title to a plain fetch. Platform vendor publishing guidance it enforces through App Review; page consolidated 2 May 2023, minor update 8 June 2026.
Read this page for its register rather than its content. Nine patterns, every one of them stated as an instruction, and not a single citation, sample size or methodology note anywhere on it. Then notice that its Confidence pattern is the sharpest sentence any publisher has written on the subject and it cuts directly against a per-field confidence UI like yours. Holding both of those at once — this is the best line in the canon, and it is an assertion — is the reading skill this lesson is for.
Stuck, curious, or think this lesson is wrong? Ask your teaching agent. The lessons are the scaffold; the conversation is where the learning gets unstuck.