Every factual claim in this course traces to something here, grouped by the module it serves. Each of these was opened and read; where a page could not be reached, the claim that would have rested on it was cut rather than sourced to a summary. If you find a source that contradicts a lesson, that is worth raising with your teaching agent — the lesson should change, or it should say why it disagrees.
What question each number answers, how to define it so a stranger computes the same thing, and why every one of them has a second reading.
WhySource of the HEART categories (Happiness, Engagement, Adoption, Retention, Task success) and the Goals-Signals-Metrics ladder, plus the PULSE critique: page views may rise “because the feature is genuinely popular, or because a confusing interface leads users to get lost in it.” Written by three Google employees, applied to more than twenty Google products, with every worked example a Google product — a company documenting its own internal practice, not an independent evaluation of whether the method produces better decisions. Its own caveats are worth keeping: not every category belongs on every product, and these metrics evaluate launched products rather than replacing formative research. Primary source for “the users like it is not a finding.”
WhyPage 308 carries the sentence everyone quotes: “When a measure becomes a target, it ceases to be a good measure.” Strathern does not claim it as her own — she attributes the naming to Hoskin, who named it after Charles Goodhart’s observation on monetary control. Goodhart’s 1975 original could not be reached for this course, so the widely circulated “any observed statistical regularity will tend to collapse…” wording is deliberately not quoted anywhere in these lessons. The most-repeated line in metrics criticism is itself a claim nobody checks the source of.
WhyThe rare standards document that is candid about the limits of its own subject: it names a “current lack of consensus on robust and verifiable measurement methods for risk and trustworthiness,” and warns that measurement approaches “can be oversimplified, gamed, lack critical nuance, become relied upon in unexpected ways, or fail to account for differences in affected groups and contexts.” MEASURE 1.1 requires documenting the characteristics that will not or cannot be measured — the requirement this course steals for its “will not measure” list — and it states plainly that an inability to measure a risk does not imply the risk is low. Its human-AI configuration guidance suggests collecting data on the frequency and rationale with which humans overrule AI output, which is the escalation metric in another vocabulary. Also cited in the instrumenting module: “The schema comes before the component” borrows MEASURE 1.1’s requirement to document what will not be measured, while explicitly declining to credit the framework for the schema-before-component argument itself.
WhyDefines trust as “the attitude that an agent will help achieve an individual’s goals in a situation characterized by uncertainty and vulnerability,” and gives the vocabulary this course uses for what “appropriate” means: calibration (does trust match capability), resolution (how finely trust tracks capability), specificity (does trust attach to a component or spill across the system). Source of the design recommendation “design for appropriate trust, not greater trust,” and of the misuse/disuse framing in which over-reliance and under-reliance are both failures.
WhyThe basis for this course’s honesty about “time-to-trust”: the review catalogues thirty measures of trust in automation — sixteen self-report, nine behavioral, four physiological — and none of them is a time-to-trust. The nearest thing in the literature is a methodological proposal (attributed there to Yang et al., 2017) to quantify trust as the “area under the trust curve,” which requires sampling trust dozens of times within an experimental block: a research protocol, not a product metric. Also the source for the validated instruments named in the lessons, Jian, Bisantz & Drury’s twelve-item checklist and Madsen & Gregor’s twenty-five-item questionnaire. Primary source for the definitions lesson.
Why1,626 crowd workers across three tasks, with AI accuracy deliberately pinned at 84% — roughly human parity — so that deferring to the AI could not masquerade as improving. Headline: “explanations increased the chance that humans will accept the AI’s recommendation, regardless of its correctness.” A bare confidence score matched explanations on accuracy (beer task: 0.89 ± 0.05 versus 0.88 ± 0.06, not significant), and neither adaptive nor expert-written explanations changed that. The mechanism matters more than the headline: explanations raised accuracy when the AI was right and lowered it when the AI erred, cancelling to a flat aggregate. Primary source for the interpretation lesson.
Why199 analysed participants against a simulated AI held at 75% accuracy, comparing ordinary explainable-AI interfaces with three cognitive forcing functions (reveal on demand; commit unaided first, then reveal; wait while the AI “processes”). Overreliance on incorrect predictions fell 0.64 → 0.48 on the sub-decision analysed most closely. The finding this course actually uses is the correlational one: self-reported trust was positively correlated with overreliance, while mental demand was positively correlated with performance on exactly the cases where the model was wrong. Two caveats the lessons repeat rather than bury — participants with no AI at all beat every AI condition on incorrect predictions, and the forcing functions helped high-need-for-cognition participants disproportionately, which the authors frame as an intervention-generated inequality. Also the primary source for the instrumenting module’s “What only the UI can see,” which takes the correlation between self-reported trust and overreliance as license to treat hesitation, edits and reverts as measurable proxies — without claiming the paper measured dwell time itself.
Why13,821 papers screened to 74 studies, with a meta-analytic risk ratio of 1.26 (95% CI 1.11 to 1.44) for following erroneous advice: decision support raised the risk of an incorrect decision by 26%. Source of the definition this course uses — users over-accept computer output “as a heuristic replacement of vigilant information seeking and processing” — and of the split between errors of commission (following incorrect advice) and errors of omission (failing to act because nothing prompted you to). The omission half is the one a review gate structurally cannot observe.
WhyA survey of 41 policy documents mandating human oversight of algorithms, finding that none of the seven requiring “meaningful” human input proposes a definition of what meaningful means — the parallel this course uses for why a metric name without a definition does no work. Also the source for two hard findings quoted in the interpretation lesson: London police “overwhelmingly overestimated the credibility” of a live facial recognition system, judging matches correct at three times the actual rate of accuracy; and automation bias persists after training and explicit instructions to verify. Green also reports the honest counterexample — an evaluation of the Allegheny Family Screening Tool where staff were able to override many algorithmic errors.
WhyNames automation bias inside the statute: an overseer must be enabled to “remain aware of the possible tendency of automatically relying or over-relying on the output produced by a high-risk AI system (automation bias).” Article 14(4) also requires that they be able to decide not to use the system, to disregard, override or reverse its output, and to interrupt it through a stop button or similar procedure — a list of affordances, each of which is an emittable event and therefore a metric you cannot otherwise have. This is a consolidated reading copy, not the Official Journal; EUR-Lex would not load for this course, so check the OJ text before relying on any of it for legal purposes.
WhyA synthesis of roughly sixty papers, defining overreliance as users accepting incorrect AI recommendations — errors of commission — and automation bias as favouring recommendations from automated systems while disregarding non-automated sources. Used in the interpretation lesson for one asymmetry it collects: trust drops by a relatively large amount when system capability decreases, and returns by a much smaller amount when capability recovers. Vendor-affiliated: this is Microsoft’s own AI ethics committee reviewing a problem in products Microsoft ships, and it is not peer-reviewed — treat it as a well-organised secondary index and follow it to the primary sources.
What must never leave a review surface, and the emerging convention for the trace side of the routing decision — which is not yet a settled standard, and is described here as what it is. The argument that the schema comes before the component is this course’s own; no source is cited for it, because every source making it sells the tooling that implements it.
WhyThe authority behind the never-logged table. Section 3.5 states that only personal data “adequate, relevant and limited to what is necessary for the purpose” may be processed, and that controllers should verify whether the purpose can be achieved “by processing less personal data, or having less detailed or aggregated personal data or without having to process personal data at all.” Its listed design elements — data avoidance, limitation, relevance, necessity, aggregation, pseudonymisation, anonymisation and deletion — map almost directly onto analytics payload design, and its worked example of a transport operator that deliberately does not store the ticket identifier is the exact shape of decision an event schema has to make. Primary source for “what must never be logged.” The ICO’s equivalent pages returned HTTP 403 for this course and are therefore not cited anywhere in it; the NIST Privacy Framework downloaded but its text could not be extracted, so this course has no US-side authority for a never-logged table and says so in the lesson.
WhyPrimary source for “analytics, or trace,” cited for the shape of the decision rather than for the names. It defines gen_ai.evaluation.result (Recommended, Development) to capture the result of evaluating generative-AI output for quality, accuracy or other characteristics, carrying an evaluation name, a score value or label and an explanation — which is where a judge score belongs, next to the run it scored rather than averaged in product analytics. It also defines gen_ai.client.inference.operation.details (Opt-In, Development), the event that carries input messages, output messages, system instructions and tool definitions, kept deliberately separate from the spans. Read that split as the routing argument this course makes, arrived at independently one layer down. What it is not: a stable standard. As of the v1.42.0 release on 12 June 2026 the GenAI conventions were moved out of the main semantic-conventions repository into this dedicated one, and names in the namespace have already changed at least once — gen_ai.system became gen_ai.provider.name, gen_ai.usage.prompt_tokens became gen_ai.usage.input_tokens. A schema that hard-codes these attribute names will need renaming.
WhyTwo things the lessons use. The content-capture default, stated plainly: “By default, no prompt content or tool arguments are captured with GenAI telemetry, as these can contain sensitive data.” That is the same call the never-logged lesson makes about a review gate, taken one layer down by people with the same problem, and it is why content-off is treated here as the starting position rather than as a restriction to argue against. And the span hierarchy — an invoke_agent span wrapping the chat and execute_tool spans beneath it — which is the structure a review-gate span has to attach to. Cited in “what must never be logged” and “analytics, or trace.”
WhyThe plainest available statement of the maturity position that “analytics, or trace” is built around: every gen_ai.* attribute, span, metric and event in the registry carries the “Development” badge and none is marked Stable, contrary to a run of posts claiming the namespace went stable at OTel 1.30. It records the 12 June 2026 move to a dedicated repository with no tagged release, and it states the consequence in one line — Development status explicitly means names can still change — with the practical recommendation to treat the wire protocol as the integration point rather than betting a schema on particular attribute names. A community blog post, not a peer-reviewed or official source; it is used for the currency check on a claim the project’s own docs support in substance and understate in tone.
What an eval scores, where a model judging a model breaks, and which tool answers which question.
WhyThe measured case both for and against LLM judges, which is why it is cited rather than any vendor’s claim about its own scorers. On MT-Bench first turn, GPT-4 agreed with human preferences 85% of the time against 81% human-human agreement (non-tie votes). The same paper measures the failure modes: position bias, where consistency under swapping the two answers was 65.0% for GPT-4, 46.2% for GPT-3.5 and 23.8% for Claude-v1; verbosity bias, where a padded answer with no new information fooled Claude-v1 and GPT-3.5 91.3% of the time and GPT-4 8.7%; a possible self-enhancement bias the authors explicitly say their data cannot confirm; and weak math grading, with failures dropping from 14/20 to 3/20 only under reference-guided prompting.
WhyThe scale check on the position-bias finding above: 15 LLM judges across 22 tasks and roughly 40 solution-generating models, over 150,000 evaluation instances, with three defined metrics — repetition stability, position consistency, preference fairness. Reports that position bias is not random, varies significantly across judges and tasks, and is strongly affected by the quality gap between the two solutions. Only the abstract page was opened for this course, so no per-judge numbers from it appear in any lesson.
WhyThe vendor’s own account of its evaluation primitives: datasets of examples, workspace-level evaluators (including LLM-as-judge and pairwise), the offline/online split between pre-deployment benchmarking and production monitoring, and annotation queues for structured human feedback. Tracing is the base record from which datasets get built. Its published pricing has a $0 developer tier and a $39/seat/month Plus tier. Every capability claim here is the vendor’s, unverified by anyone independent.
WhyA config-and-CLI offline eval tool: test cases, deterministic assertions and graders, model-graded (LLM-as-judge) checks, red teaming, and CI integration. No annotation-queue or online production-eval primitive comparable to the other two. MIT licensed, with the source at github.com/promptfoo/promptfoo. Note the ownership change: promptfoo announced on 9 March 2026 that it is joining OpenAI, stating it “will remain open source.” That is the vendor’s own commitment about its own future, published on the day of the announcement — exactly the kind of claim to re-check rather than plan around, and it is not mentioned on the marketing or pricing pages.
WhyThe vendor’s decomposition of an eval into data (test cases with inputs, optional expected outputs, metadata), task (the function under evaluation) and scorers (built-in autoevals, LLM-as-judge, or custom code). Its online scoring evaluates production traces as they are logged, using LLM-as-judge because live requests have no ground truth — which is where the judge-reliability numbers above become load-bearing rather than academic. Published pricing runs from a $0 starter tier to $249/month Pro. Vendor self-description throughout.
WhyThe source for self-preference bias, and the reason the judge lesson does not attribute that bias to Zheng et al. (2023) the way most write-ups do. Zheng and colleagues studied self-enhancement bias and reported no significant evidence that GPT-4 or Claude favour their own responses — an explicitly unconfirmed finding. This paper measures it directly, reporting that LLM evaluators including GPT-4 and Llama 2 have non-trivial accuracy at distinguishing their own outputs from those of other models and of humans, and that self-recognition capability correlates linearly with the strength of self-preference bias in controlled experiments that rule out straightforward confounders. Only the proceedings abstract was opened for this course, so no accuracy or correlation figure from the full paper is quoted in any lesson. Cited in “LLM-as-judge, and where it fails.”
WhyUsed for one thing: evidence that the correction-to-dataset loop is a recognised practice rather than an invention of this course. Production usage surfaces failures, humans annotate and fix them, the fixes are stored, and the stored fixes come back as material for evaluation and improvement. Her emphasis falls as much on retrieving corrections as few-shot context as on building an eval set, which is a different endpoint from the review-gate framing in “Corrections become the eval set” — that framing, and the scoped claim that a discarded correction is a label the interaction will never produce again, are the course’s own and are not attributed to her. Chosen over any vendor’s flywheel page precisely because she is not selling the practice she describes. Primary source for “Corrections become the eval set.”
WhyThe published figures the tooling lesson quotes, with the date attached because they are exactly the kind of fact that moves: a $0 Developer tier including 5,000 base traces per month, a Plus tier at $39 per seat per month including 10,000 base traces, and Enterprise quoted individually. Re-check before repeating either number.
WhyThe published figures the tooling lesson quotes: a $0 Starter tier and a Pro tier at $249 per month described as including roughly 5GB of data and 50,000 scores per month with per-unit overage beyond that, plus Enterprise quoted individually. Date-stamped for the same reason as the LangSmith entry.
Every claim on these pages links to its source. If a source looks wrong or out of date, check the resource list and tell your teaching agent — the course is meant to be corrected.