Every factual claim in this course traces to something here, grouped by the module it serves. The bar is higher on this course than on the rest of the shelf, because this course does nothing but characterise other people’s published work: a claim that could not be traced to a page someone actually opened was cut, not softened. If you find a source that contradicts a lesson, or a lesson that characterises a source in a way its own page does not support, that is worth raising with your teaching agent immediately — a mischaracterised framework makes every disagreement built on it worthless.
The frameworks themselves, plus the dates that make the provenance argument checkable.
WhyThe primary source for the whole first module, and the only place the eighteen guidelines appear with their phase groupings printed together (Table 1). Table 2 lists the twenty products across ten categories the guidelines were validated against — recommenders, autocomplete, feed filtering, web search — which is the entire provenance argument in one table. Methodology: 168 candidate recommendations to 20, cut to 18 by eleven researchers over thirteen products, then validated by 49 practitioners. Section 7 carries the authors’ own scope limit. Two numbers are easy to garble: the contributions bullet says “over 150” recommendations where the methodology says 168, and validation used twenty products, not eighteen. Read a second time in the second module: Table 1’s example-application column for G7–G11 is the primary source for the suggest-versus-act lesson, and Section 7’s exclusion note — the authors deliberately left out Horvitz’s principle of inferring ideal action from costs, benefits and uncertainties because it belongs at the modeling layer — is what the over-hedging lesson turns on.
WhyFour components, in Microsoft’s own words: the Guidelines for Human-AI Interaction, the HAX Design Library (patterns and examples), the HAX Workbook (a team exercise for prioritising which guidelines to implement) and the HAX Playbook (common failures for NLP applications). Three of the four are about getting a team to apply the guidelines, which tells you Microsoft does not treat the eighteen as a checklist to be passed.
WhySource for the count (eighteen), the claim that they are “a synthesis of more than 20 years’ research, introduced in this award-winning 2019 CHI paper”, and the four phases named in running prose. Two facts about this page are used in the course as findings in their own right: it does not print the list of guidelines, and it carries no revision date.
WhyThe bridge document. It links the HAX Toolkit explicitly and reuses eight guidelines verbatim under the same four lifecycle headings, then adds three principles the eighteen do not contain — human in control, avoid anthropomorphizing copilot, consider direct and indirect stakeholders — plus the output-design heading “Add appropriate friction (it’s a good thing!)”, which inverts the reflex most designers were trained on. Also the source of a small, instructive wording drift: it renders G11 as “Make it clear why the system did what it did” where the paper says “Make clear why the system did what it did.”
WhyThe human-authored companion to the Learn agent page, carrying the same three principles: built for intent, differentiated from humans, bias resistant. Useful when you want a dated, human-bylined Microsoft position on agent design rather than an AI-generated docs page.
WhySix chapters organised as a product lifecycle: User Needs + Defining Success, Data + Model Evolution, Mental Models + Expectations, Trust + Explanations, Feedback + Controls, Errors + Graceful Failures, plus patterns, workshop kits, case studies and a glossary. Three of the six have no counterpart anywhere in the eighteen. Cite this URL and never a /chapter/ path — see the note on the superseded edition below.
WhyThe citable source for the Guidebook’s edition history in Google’s own words: originally published 8 May 2019, current release the third edition updated April 2025, with the 2025 work explicitly being the generative-AI update. The Guidebook’s own footer says the same, but the site is client-rendered and awkward to quote.
WhyListed here as a trap, not as a source. The first edition is still hosted on the same domain, at this URL and at a set of /chapter/… paths that are the ones in pair.withgoogle.com’s own sitemap — so they are what search returns. Nothing on those pages says they have been replaced. The chapter names are the tell: “Explainability + Trust” and “Data Collection + Evaluation” are the 2019 names; the current ones are “Trust + Explanations” and “Data + Model Evolution.”
WhyA fair sample of the Guidebook’s workshop material, and a fetchable artifact of the current edition. The activity asks teams to find where users would under-trust the feature and where they would over-trust it. Calibrating trust rather than maximising it, published by a company that sells AI, is worth being able to quote.
WhyEvidence that a second edition existed, introduced at Google I/O with the first “released two years ago”. Worth noting that the codelab itself still teaches the first edition’s chapter names and contains no generative-AI material — another artifact Google left in place.
WhyThe date anchor for the provenance argument: the interleaved reason-and-act loop that essentially every tool-using agent now implements is posted three and a half years after the eighteen guidelines were presented.
WhyUsed for the date, and for Anthropic’s own description of the capability as “still experimental—at times cumbersome and error-prone”, which is a more honest sentence than most vendor launch copy and worth quoting as such.
WhyThe fourth date anchor. Five and a half years after CHI 2019.
WhyDeliberately the weakest citation in this course, and flagged as such in the provenance lesson. OpenAI’s own announcement page returned HTTP 403 to every fetch attempt behind these lessons, so the date for the API that put tool use in ordinary developers’ hands comes from a dated, contemporaneous third-party writeup rather than from the vendor. Do not upgrade it to a primary citation without fetching OpenAI’s page successfully.
The three gaps this module demonstrates, and the evidence on both sides of the one that is genuinely contested. Every publisher here has an interest: Microsoft, Google and Apple all sell the category of product they publish guidance about, and NN/g sells training and consulting alongside its research.
WhyThe only guidance in the canon that addresses a system taking an action: “Consider consequences and get permission before performing irreversible or potentially problematic tasks” — avoid automating destructive actions, ask for confirmation before a significant action on someone’s behalf. It is one bullet, framed entirely as permission before a task begins, with nothing about a run already underway. Its hallucination guidance — “Raise awareness about and minimize the chance of hallucinations” — is one of the two published pushes toward disclosure that the over-hedging lesson argues against. Cited in the suggest-versus-act and over-hedging lessons.
WhyNine patterns — Explicit feedback, Implicit feedback, Calibration, Mistakes, Corrections, Multiple options, Confidence, Attribution, Limitations — every one stated as an instruction, and not a citation, sample or methodology note anywhere on the page. Its Confidence pattern is the sharpest sentence any publisher has written on the subject and cuts directly against a per-field confidence UI: “If you’re not sure how your confidence values correlate with the quality of your results, it’s not a good idea to convey confidence to people.” The primary source for the platform-and-consultancy lesson, which uses it as the worked example of a rule with no visible evidence.
WhyMicrosoft’s newest, agent-labelled guidance, and it holds the suggestion frame in both places this module tests. Its “Built for intent” principle keeps the user as the one initiating action, illustrated as “Summarize with Copilot” rather than “Copilot, summarize” — a per-action, user-initiated model rather than autonomous multi-step execution. Its “When the system is wrong” section defines failure entirely as bad output: “Users should be able to regenerate responses, revise prompts, or manually edit outputs without friction.” A full-text read found no undo, rollback or compensation language anywhere on the page, which is the falsification criterion the halfway-failure lesson states. It also carries the HAX phase structure almost word for word — first-run experience, during interaction, when the system is wrong — without ever naming HAX or the eighteen guidelines, which the-canon-was-written-for-a-different-machine lesson in the first module reads as unacknowledged lineage. Cited in suggest-versus-act, failure-halfway-through and the-provenance-problem.
WhyThe primary source for the halfway-failure lesson, and the readable route into the current edition’s error chapter: the Guidebook is client-rendered and its chapter pages return only a title to a plain fetch. Its taxonomy is three types, verbatim — System limitation (“can’t provide the right answer, or any answer at all, due to inherent limitations”), Context (“working as intended,” but the user perceives an error), and Background (“the system isn’t working correctly, but neither the user nor the system register an error”). A run that committed two side effects and then stopped fits none of them; the only compounding language on the sheet is one checkbox, “Is your feature unusable as the result of multiple errors?”, which is about the feature degrading rather than about unwinding what already fired. Treat it as one artifact of the chapter, not the chapter’s prose. Its cover page reads “Errors + Grace Failure” where the running header says “graceful” — Google’s typo, quoted as found.
WhyThe reason this module’s absence claim is narrow rather than sweeping. Yocco proposes six named patterns — the Intent Preview, the Autonomy Dial, the Explainable Rationale, the Confidence Signal, the Action Audit & Undo, and the Escalation Pathway — several of which aim squarely at the suggest-versus-act and partial-failure gaps. The article states no methodology, no sample and no academic citations for the patterns, and it is one practitioner rather than an institution publishing validated guidance. So the accurate claim is that no institutional canon publisher addresses these cases, and one individual has started to. Named in the suggest-versus-act lesson for exactly that reason: naming what exists makes the absence more credible, not less.
WhyThe primary source for the over-hedging lesson and the only study found that manipulates hedging language directly and measures the result. First-person expressions of uncertainty such as “I’m not sure, but…” “decrease participants’ confidence in the system and tendency to agree with the system’s answers, while increasing participants’ accuracy,” which the authors attribute to “reduced (but not fully eliminated) overreliance on incorrect answers.” The scope matters as much as the finding: medical questions on a fictional LLM search engine, which is high-stakes and unlike an advisory recommendation product. Also cited in the what-would-change-your-mind lesson (Having a position) as the model of a well-constructed piece of evidence — pre-registered, n = 404 — though that lesson leans on it as a supporting example rather than as its primary source.
WhyListed here twice over as a caution. It is the one large sample on the NN/g site and it is not NN/g’s data, which is the pattern the platform-and-consultancy lesson teaches. And it is not evidence about hedging language: it measures perceived intelligence against perceived emotional capacity — “perceptions of intelligence were positively related to taking ChatGPT’s advice” while “perceptions of emotion were negatively related” — which is a different variable. The over-hedging lesson names it explicitly as the study that looks relevant here and is not; the platform-and-consultancy lesson uses the same n = 410 as its example of a large number NN/g reports but did not itself collect.
Most of this module is the course’s own methodology and is labelled as such in the lessons — these are the four outside sources it leans on, and each carries its own kind of interest to weigh: a catalogue that advertises its maintainer’s consulting, a consultancy that sells training alongside its research, a vendor publication summarising work it did not do, and a peer-reviewed study from two research institutions with no AI vendor to declare.
WhyThe primary source for the catalogues-are-vocabulary lesson, and the course’s worked example of a source that supplies names and no findings. Seven pattern types with verbatim definitions — Wayfinders, Prompt actions, Tuners, Governors, Trust builders, Identifiers, Dark matter. Governors, “Human-in-the-loop features to maintain user oversight and agency,” is the closest published name for the HouseWarm review gate. Three absences are used as findings in the lesson: no stated methodology, no citations to research anywhere, and no date of any kind — no launch date, no last-updated, no changelog, no lastmod in the sitemap. Do not quote a pattern count: two live fetches of the same page on the same day produced different implied totals.
WhyListed separately because it is the corroborating route, not a second reading. Two summarised fetches of the Shape of AI homepage both dropped this category entirely; it was confirmed through the site’s own sitemap.xml and then by fetching this page, which returned the definition — “Potentially nefarious, but certainly ambiguous patterns that impact user trust with questionable user value” — word for word. Cited in the catalogues-are-vocabulary lesson as the reason a count that cannot be read twice is not a fact you own.
WhyThe primary source for the grounding-in-what-you-shipped lesson, read for its shape rather than its conclusions: the most directly agent-relevant empirical piece on NN/g’s site, and it is six sessions and says so. Its own scoping line — “Users still retained control; they needed to authenticate with AliPay before allowing payments” — is a limit statement, and it tells you the study never observed a committed autonomous action. The lesson uses it as the published model for naming your own narrowness. Also the platform-and-consultancy lesson’s counterweight to Apple in the second module: a source that shows its evidence, has no power to compel anyone, and is honest about its own size.
WhyThe primary source for the surviving-the-third-why lesson, used as the worked example of what a third-level answer looks like: a within-subjects study on a logic-puzzle decision-support task finding that “well-calibrated confidence scores significantly improved decision accuracy (+20%…), whereas miscalibrated scores yielded minimal accuracy gains (+2%…) and increased vulnerability to automation bias and conservatism bias.” Read it as a third position rather than as ammunition — it says the operative variable is calibration, so neither side of the confident-versus-hedged binary owns it. The task is logic puzzles, not advisory finance; the mechanism may transfer, the effect size should not be quoted as if it did. Also the source the over-hedging lesson (Where the canon is thin) uses to dissolve its own hedged-versus-confident binary — the same scope caution applies there.
WhyThe primary source for the what-would-change-your-mind lesson, and deliberately a weak one — reading it is the exercise. It reports work by Valerio Capraro and colleagues as making participants “less accurate, more confident and far less likely to say ‘I don’t know,’” with accuracy declining “by a factor of three, while confidence rose by a factor of 2.5.” The underlying paper is an arXiv preprint that has not been peer reviewed, and this course knows it only through this secondary writeup. Cite it as a preprint reached through a summary, or not at all — and note that it supports the course’s position rather than threatening it, which is why it does not belong in a change-my-mind cell. Also cited, with the same qualifiers, in the over-hedging lesson (Where the canon is thin).
Every framework on these pages links to the page its publisher actually wrote. If a lesson characterises a guideline in a way its own source does not support, check the resource list and tell your teaching agent — a mischaracterised framework makes every disagreement built on it worthless.