The three-hour rule and what to cut first
Three hours is a constraint, not an estimate — which means the skill being practised is deciding what to drop in hour one, and having that cut list written down before the clock starts is what separates a finished prototype from an abandoned one.
Somebody from a Vista portfolio company describes a problem at eleven in the morning and wants something to react to before the day ends. That is the shape of the job, and it is why every practice run in this course is three hours long. Not because three hours is optimal. Because three hours is roughly what an afternoon is, and the claim you are training to make is about afternoons.
Which makes the number a constraint, not an estimate. An estimate is a prediction you can turn out to be wrong about. A constraint is a wall you do not get to move, and the only variable left is what you decide not to build. So the skill being practised is not speed. It is triage under a clock that will not negotiate, and the difference between a finished prototype and an abandoned one is almost always a cut made in hour one.
What the evidence for timeboxing actually is
Before the rule, the thing the rule is not. You will meet the claim that deadlines improve output, usually attached to one famous experiment. This course does not make that claim, and the reason got considerably stronger in the week this lesson was written.
The experiment is Ariely and Wertenbroch’s 2002 paper on procrastination and self-imposed deadlines, published in Psychological Science. Two separate things have happened to it, and they are worth keeping apart, because they are different kinds of finding.
First, it did not replicate. Kyle Hyndman and Alberto Bisin ran a preregistered replication, published in Psychological Science 37(8) in August 2026 — the same journal that published the original. Their conclusion: “the results of the paper do not replicate… changes in the deadlines have a negligible effect on the three performance metrics”.
Second, and separately, the original data were fabricated. On 31 August 2026, Uri Simonsohn, Leif Nelson and Joe Simmons published a forensic analysis of the original dataset. Their evidence on Study 2: an effect size of roughly 2.5 standard deviations, which they call implausibly large for the design; 18 of the 20 participants in one condition having what they name a “Corrections Twin,” a matching participant exactly ten ID positions away; correlations that are strongly positive in the replication data and near zero in the original, for variables that have to move together; and a rounding pattern in which only 11.7% of estimated minutes were round numbers, against 85% in the honest replication data. Their conclusion is that the data “were severely tampered with or fabricated”. Coauthor Klaus Wertenbroch requested retraction on 23 July 2026, and the journal retracted the paper in full — both studies — on 2 September 2026, with Wertenbroch quoted saying that much or all of the data, and therefore the results, are false. Retraction Watch, the Duke Chronicle and the Chronicle of Higher Education reported it independently.
Do not merge those two events. The replication found no effect. The fraud finding is a separate forensic analysis of the original data, done by different people, using different methods, and published a year and a half after the replication preprint. Getting that wrong in a room is the exact kind of sloppiness this course is trying to train out of you.
A retracted paper is not weak evidence
It is tempting to keep a famous result around as a hedged mention — “influential but contested,” “the classic study, since questioned.” Do not. A paper retracted for fabricated data is not a weak data point, a dated data point or a disputed one. It is not a data point. This course cites it nowhere as support for anything, and neither should you.
Nor should you overcorrect into the opposite claim. Nothing here shows that timeboxing hurts. The honest position is that there is no rigorous surviving evidence that timeboxing improves output — although the replication authors note that their own earlier work and a 2011 study of intermediate targets both found performance worse under deadlines in their particular settings. That is a reason not to sell three hours as a productivity technique. It is not a reason to stop practising against a clock.
You will also meet Parkinson’s law, usually as “work expands to fill the time available.” That line comes from a 1955 Economist essay about the growth of the British civil service, and it is satire, not a study. The original text is behind a paywall that returned a payment-required response on both hosts tried during research for this course, so it is not cited here as something anyone read. Its popularity is not evidence either.
So why three hours
Because this course chose it, for three stated reasons. That is the whole justification, and dressing it up in citations would be a worse version of the thing the previous section just took apart.
- It matches the promise. The role expects an agentic surface the same afternoon. Practising against six hours would measure a claim nobody is asking you to make.
- It is short enough to force a cut. A constraint that never makes you drop anything trains nothing. Three hours against a real portfolio-company target reliably forces at least one real cut, which is the repetition the practice is for.
- It is long enough to reach something demoable. With a kit and a mock backend already in place, three hours clears the scaffold, the surface and one approval loop. A constraint that guarantees failure measures nothing either.
None of that is a research finding, and this course does not present it as one. It is a design decision with reasons attached, which is a different and more honest kind of claim.
The cut list, and why it is written first
Rule 2 of the practice log: the cut list is written before the clock starts. The reason is not discipline for its own sake. A cut decided at the hundred-minute mark always looks reasonable, because you make it with full knowledge of which part is currently painful. That is precisely the input you want excluded. Writing the list cold, before you know which phase will hurt, is what makes the run a measurement of your judgment rather than a record of your mood.
A cut list that works has three properties:
- It is ordered. First, second, third. “Some of the polish” is not a cut list; it is a feeling with bullet points.
- Each item is a thing you can stop doing. “Support one document type instead of three” is executable. “Move faster on the surface” is not.
- It names what is not on it. The protected set matters more than the cut set, because the demo’s claim lives there.
Cut breadth before depth
This is this course’s own ordering principle, not a sourced one, and it is worth stating plainly because it inverts the instinct. Under pressure most engineers cut quality: three entity types, all of them rough. The demo that survives contact with a room is the opposite — one entity type, complete, with the approval step working and the failure case visible. A viewer can extrapolate breadth from one working path. Nobody can extrapolate a working path from three broken ones.
A default order against the phase table you already have in the practice log, which you should adapt rather than adopt:
Cut 1 Scope of the domain surface — one entity type, not three
Cut 2 Mock breadth — one happy path and one failure, not a matrix
Cut 3 Polish — empty states, sub-tablet layout, transitions
Protected — never cut, or the demo stops making its claim
The approval gate
Streaming cadence and run state
The one path the walkthrough narrates end to endNotice what is absent from both lists. The trace view and the review gate are not decisions in a run any more, because the fake-backend module wired them into the kit once. That is the whole point of having built them: they left the cut list by leaving the run.
The checkpoints that make cuts happen on time
A cut list you never look at is a document, not a mechanism. Two glances at the clock convert it into one:
- At 60 minutes, the domain surface should be started. If it is not, execute cut 1 now. Not soon.
- At 120 minutes, the approval loop should run end to end, even ugly. If it does not, execute cuts 2 and 3 together and spend the remainder on that loop.
Cuts made at 150 minutes are not cuts. By then the alternatives have already been spent, and what feels like a decision is just an observation about what did not get finished.
Check your recall
Answer from memory — no scrolling back.
Retrieval check
Someone challenges the three-hour rule: “where is the evidence that timeboxing produces better work?” What is the honest answer?
Check your answer
That there is none, and that this course does not claim otherwise. The most-cited experimental support for deadlines improving performance was retracted in September 2026 for fabricated data, and the preregistered replication that preceded the retraction found deadline changes had a negligible effect. Nothing rigorous survives on that side of the question.
Three hours is not here to make the work better. It is here because the job is described in afternoons, and because a constraint that forces a cut is the only way to practise cutting. That is a design decision with reasons, stated as a design decision. It is a stronger position than a citation, because it cannot be retracted out from under you.
Hands on
Run 1, against a written cut list
Done when: PRACTICE-LOG.md contains a completed Run 1 with a cut list timestamped before the clock started, live-recorded phase times, and an explicit comparison of what you actually cut against what was on the list.
- Pick a target from a domain you do not know, with a step a human would have to approve before anything fires. Write the target sentence and the demo claim — the one sentence you would say while showing it — before anything else.
- Write the cut list. Three ordered items, each a thing you can stop doing, plus the protected set. Then write the time you finished writing it, so the log shows it preceded the run.
- Set alarms at 60 and 120 minutes. Not reminders you intend to notice — alarms.
- Run it. Record phase minutes as you leave each phase, including the yak-shaving row. Stop at three hours whether or not it works.
- Fill in What actually got cut against the list. If you cut something that was not on it, write why you did not foresee it. That mismatch is the interesting number, not a failure.
- Fill in What the kit had that got in the way. It is the column everyone leaves blank, and a kit that never loses an item becomes the framework this whole course exists to avoid.
- Bring the cut list and the actual cuts into the chat. I will push hardest on any cut item that is a mood rather than an action, and on a protected set that quietly contains everything.
What this does not cover
A cut list is only as good as the thing it is cutting from. Run 1 above told you to pick a target with an approval step and left the rest to instinct, which is enough for one run and not enough for six — a badly chosen target turns a practice run into an afternoon of styling, and you will not notice until the log has three rows that all measure the same thing. Choosing targets deliberately is the next lesson, Picking a target worth the clock.
This lesson also says nothing about what you build with. The runs are where your position on v0, Claude Code, Cursor and Figma Make comes from, and that position has its own requirements about what counts as evidence, in the lesson on an evidenced position on the tools.
Read this next — primary source
[138] Artificial Deadlines (Part 1): Evidence of Fraud in an Influential Study About ProcrastinationUri Simonsohn, Leif Nelson & Joe Simmons — Data Colada, 31 August 2026 — free
Read this for the method, not the scandal. The post walks through how you establish that a dataset was fabricated when you cannot see the fabrication directly: an effect size too large for the design, an identifier pattern that pairs participants who should be unrelated, correlations that appear in honest replication data and vanish in the original, and a rounding distribution no human self-report produces. It is a working demonstration of reading a claim rather than counting its citations. This lesson uses it for one purpose only, which is to establish that the most-cited experimental support for deadlines is retracted and therefore unusable — the three-hour rule here rests on nothing it says.
Stuck, curious, or think this lesson is wrong? Ask your teaching agent. The lessons are the scaffold; the conversation is where the learning gets unstuck.