You can build the review gate. You have built one. What you cannot yet do is stand in front of a portfolio company and say, with evidence, that it changed anything — and “we standardised the review gate across ninety products” is worth nothing without a number attached to it. This course is about the number: which ones a UI can actually produce, what each one is silent about, and what the components have to emit before any of them exist.
Start lesson one: “The users like it” is not a finding →
The scope, stated plainly
This is not an ML evaluation course and it does not try to make you an evals owner. The discipline being taught is instrumentation design: knowing what to make the interface observe so that somebody else can measure it. The evals module exists so you can argue with that person, not replace them.
This course is being written as you work through it. The first module is ready; the remaining two are registered so the numbering stays stable, and ship once the metric set has been argued with rather than read.
The five numbers a review gate can produce, defined precisely enough to survive an argument.
Designing the events before the component, so the numbers exist rather than being reconstructed later.
Not owning the eval stack — knowing enough about it to argue with the person who does.
Every claim on these pages links to its source. If a source looks wrong or out of date, check the resource list and tell your teaching agent — the course is meant to be corrected.