Your own reference class
Seven finished-or-half-built projects are now a dataset about your own conversion rates and maintenance habits, not just a graveyard — how to read it before you start project number eight.
Every field the gates asked for was written to answer a question about the project it sits under. Is this one worth pursuing. What does it cost to keep alive. Who would hand over the money. When do you open it again. Each project got decided on its own merits, which is what the gates are for, and each verdict carries a date because the review is going to come back round.
Read the same file down its columns instead of across its rows and it stops being about any project at all. Seven state labels. Seven maintenance figures. Seven numbers in a table you have been forbidden from using since triage. Nothing in that reading is evidence about a recipe app or a guides site. It is evidence about the person who started them.
Seven attempts are a dataset, not a graveyard
The instinctive reading of a shelf of unfinished and unearning projects is that it is a record of failure, and that the correct emotional response to it is to look away and start something better. That reading throws away the only information on the shelf that a stranger’s framework could never have given you.
Bent Flyvbjerg’s paper is about forecasting the cost of very large infrastructure, which could hardly be further from your situation, and the definition at the top of it is the part that transfers without modification:
“The outside view on a given project is based on knowledge about actual performance in a reference class of comparable projects.” (Flyvbjerg, Getting Risks Right)
And what that view refuses to do is as important as what it does: “reference class forecasting does not try to forecast the specific uncertain events that will affect the particular project, but instead places the project in a statistical distribution of outcomes from the class of reference projects.” No scenario, no imagined obstacle, no judgement about how motivated you are this time. Only where the thing lands among the things like it.
The paper carries an example that is worth having in front of you before you open your own file, because the shape of the answer is exactly the shape of the answer you are about to produce. It is a story Flyvbjerg takes from Kahneman, about a team building a school curriculum who were each asked to write on a slip of paper how many months the work needed: “The estimates ranged from 18 to 30 months.” One member, an expert in curriculum development, was then asked to recall comparable projects and say how long those had taken from a comparable point:
“After a while he answered, with some discomfort, that not all the comparable teams he could think of ever did complete their task. About 40 percent of them eventually gave up.” (Flyvbjerg, Getting Risks Right)
Of the ones that did finish, he could not think of any that took less than seven years or more than ten. The team pressed on and, in the paper’s words, “finally completed the project eight years later, and their efforts went largely wasted”.
Notice what the expert already knew. The finishing rate was in his head the whole time, available to be stated the moment somebody asked the right question, and it never once reached the slips of paper. That is the failure this lesson is aimed at, and it is not a failure of data collection. Your ledger has been sitting in a repository for the length of this course.
What the hours-sunk column is finally for
The triage module recorded hours sunk and then fenced it off in the same breath, and the ledger prints the fence directly above the table: not an input to any decision. Every gate since has held that line. This is the lesson where the number gets its one job, and it is worth being exact about why the job is legitimate here and nowhere else.
Hours sunk on a specific project is a fact about a decision that has already been made. Aimed forward at that same project it can only distort, because it answers how much have I already paid when the question on the table is what will this return from here, and those two have no arithmetic relationship at all. That is why it was fenced.
Aim it backwards at the class and the question changes underneath it. Not should I keep going on this one, which the number cannot answer, but what does a project of this shape cost me before it stops — and that question cannot be answered without the number. There is nothing else in the file that measures the price of an attempt.
The fence was never around the number. It was around a direction. Which is the inside and outside view distinction arriving at your own ledger: the same figure is poison inside a project and evidence outside it.
The row discipline, borrowed from another course
There is a lesson in the planning course that teaches the mechanics of this and teaches them better than a section here could, because it is the whole of what that lesson is about. It is worth opening in another tab: the reference class you’re building for next time.
Its subject is the stretch immediately after a project closes, and its argument is that what you should write then is far smaller than the retrospective document you are imagining — which that lesson says gets written occasionally and read never, and is useless as a reference class because prose documents do not compare. What compares is rows. The range you estimated, what it actually took measured to the same boundary you estimated to, and one sentence naming what consumed the difference. Not a cause analysis, that lesson says; the dominant term.
Three things it establishes carry straight over, and it is worth taking its words rather than inventing your own, because each of the three names a failure that is invisible until you have already made it.
- The boundary. The start-and-end definition every row shares. Rows measured to different boundaries do not compare, and nothing in the file will tell you that they did not.
- The dominant term. One sentence per row, naming the thing that ate the gap. If the sentence needs a clause defending you, it is the wrong sentence.
- The second story. Borrowed there from John Allspaw’s blameless-postmortem argument, and stated as “not what someone should have done, but why the thing they did looked correct to them at the time, given what they knew”. Pointed at an estimate, it turns why were you wrong into what made that number look right at the start.
Point the second story at a project on your shelf and it stops being why did I abandon this, which produces a paragraph about character, and becomes what made finishing this look likely at the start — which produces answers like: nobody was charging for anything comparable and I did not check, or the maintenance was invisible until it arrived, or the distribution was a phase I had planned to start afterwards. Every one of those is structural, every one of them is a thing the gates now ask about, and none of them is about whether you are the sort of person who finishes things.
The warning that lesson turns on is the one to carry into your own file, and it is stated there as plainly as it can be: a padded reference class is worse than no reference class. With none you are obliged to say so and everyone discounts your number correctly. With a padded one you get the entire apparatus of the outside view wrapped around figures that were edited to be comfortable, and it is far more persuasive than it has earned.
Seven rows is not a distribution, and saying so is the exercise
Now the part where the borrowed method does not fit, which you need before you start rather than after. Flyvbjerg states the procedure as three steps, and the first two are entry requirements:
“(1) Identification of a relevant reference class of past, similar projects. The class must be broad enough to be statistically meaningful but narrow enough to be truly comparable with the specific project. (2) Establishing a probability distribution for the selected reference class. This requires access to credible, empirical data for a sufficient number of projects within the reference class to make statistically meaningful conclusions.” (Flyvbjerg, Getting Risks Right)
Your ledger fails both, and not narrowly. Seven rows written by one person about their own evenings is not a statistically meaningful class, there is no probability distribution to fit to it, and there is no percentile to place the next project in. Anything that comes out of this exercise wearing a decimal point is decoration.
Given both of those, call the output what it is. It is a prior, not a forecast. It is the number an honest person would start from before hearing anything specific about the next project, and its whole job is to be argued away from deliberately rather than skipped past silently.
The one condition your version has in its favour
The paper splits inaccurate forecasts into two causes and treats them completely differently, and the split is the reason this exercise is worth running at all. In the first, forecasters are honestly trying:
“when forecasters are honestly trying to gauge the future—the potential for using the outside view and reference class forecasting will be good. Forecasters will be welcoming the method and barriers will be low, because no one has reason to be against a methodology that will improve their forecasts.” (Flyvbjerg, Getting Risks Right)
In the second, the wrong number is doing a job for somebody. His example is cities competing for scarce national funds, and his sentence about them is flat: “There is no incentive for the individual city to debias its forecasts, but quite the opposite.” His verdict on that situation is blunter still: “In this type of situation the potential for reference class forecasting is low—the demand for accuracy is simply not there”.
You are unambiguously in the first situation. No committee reads your ledger. No funding depends on the numbers being attractive. Nobody is competing with you for the budget the file describes, because the budget is your own evenings. The padding pressure the planning lesson warns about comes from an audience that punishes a bad row, and you do not have one.
Which removes the excuse along with the pressure. If the rows come out flattering, there is nothing to blame it on.
What is actually countable in your file
Be precise about what the ledger can and cannot answer, because the temptation is to reach for the numbers a business would have and quietly invent them.
A finishing rate. Every project carries one of the five state labels the triage module defined, and three of those labels turn on the same question: whether a stranger can reach the thing. Built and unlaunched means no stranger ever could. In testing means only people you invited. Built and live means anybody can, right now, whether or not a single stranger has shown up. So how many of the seven got as far as a stranger being able to reach them is a countable fact, and it is the number the rest of this section leans on.
Where they stop. More useful than the rate, and only visible when you read the labels together. If more than one row carries the same label, that label is your dominant term — not the dominant term for any one project, but for the class. A person who stalls at the same point every time is looking at a process fact, and process facts are fixable in a way that motivation is not.
The price of an attempt. This is the hours-sunk table doing its job. Not an average — seven numbers do not have a meaningful average — but the smallest, the largest, and which state labels the largest ones are sitting next to. Hours parked beside a project that never became reachable are the hours that bought the least, and they are the rows worth reading hardest before starting another one.
How far your estimates sit from your measurements. The maintenance-load lesson asked you to measure exactly one project properly. The verdict lesson then asked you to estimate the rest, anchored to that one and labelled as estimates, because you were about to decide with them. That leaves you one reading and six guesses, and the comparison between them is the closest thing to an estimate-versus-actual row that this ledger holds. If the measured project came in above its own estimate, the six estimates around it are not neutral figures. They are figures of a kind you have now seen run low once.
And a conversion rate, which may well be zero. Nothing in this course asked you to launch anything, charge anybody, or record a month of income; every exercise was built to run on estimates and one measured number, deliberately, because a framework you cannot run until you are already earning is one you will never run. So the actuals fields are empty and the honest count of projects that have taken money may be none of them. Write that down as a fraction with its denominator rather than skipping the line. A base rate of zero is a real finding, and a blank is not.
Retrieval check
Say the actuals fields come out empty across the board, so on the record no project you have built has ever taken money. Is that a base rate, or is it a missing measurement?
Check your answer
It is a real base rate, for a narrower question than the one you want answered. It tells you how often a project you start ends up taking money, and on the record so far the answer is none of them. That is a fact about your process from idea to income, and it is the correct number to start the next project from.
It does not tell you how often a project you build and actually try to sell takes money. If nothing in the ledger has ever been put in front of a buyer at a price, the denominator for that second question is zero too, and there is no rate there at all — only an absence.
Naming which of the two you are extrapolating from is the entire skill. The first is evidence about you and is allowed to move your estimate for the next project. The second is a question you have never run the experiment for, and dressing an absence up as a discouraging result is the same error in the opposite direction as dressing it up as an encouraging one.
What the outside view says about attempt number eight
Running the gates forward on an unbuilt candidate produced answers, and every one of them was a fact about the world: whether somebody is charging, who pays, where those people already are. This is the other input, and there is only one of it. The record tells you about the person who will be doing the work, and no stranger’s framework can supply that.
The way to use it is the question the curriculum expert was asked next, after he had produced the uncomfortable number — whether he had reason to believe his present team was more skilled than the earlier ones. He said no. They went ahead anyway.
So ask it of yourself, in the strict form. Is there a reason to believe the next one goes differently, and is the reason structural or is it a feeling? Enthusiasm does not count. Every row in that file was started by somebody who wanted to start it, so if enthusiasm were predictive the file would already read differently. What counts is a change to the process that is written down somewhere and would still be there on a day when you did not feel like it: a gate that now gets run before building rather than after, a cap on how many projects can be alive at once, a kill criterion set while you could still think clearly about it, a review with a date on it.
Those are real adjustments and you are entitled to make them. Make them out loud, in writing, one at a time, and be honest that not one of them has yet produced a finished, earning project — they are the argument for a better rate, not evidence of one. An adjustment you can name and defend is the outside view working exactly as intended. An adjustment you cannot name is the inside view, back again, wearing the vocabulary of this lesson.
Hands on
Build the reference class out of the seven
Done when: PORTFOLIO.md carries a filled-in Hours sunk table and a new Reference class section at the foot of the file, holding a stated boundary, one row per project with its state, its hours and one non-defensive sentence naming what stopped it, and at least three base rates written as fractions with their denominators named — including any that come out zero.
- Confirm the Hours sunk table is complete. Triage asked for those numbers and nothing since has been allowed to use them, so this is the first time you have had a reason to look at them properly. Fill any row still blank and correct any you now know was wrong.
- Decide the boundary and write it once at the top of the new section, before any rows. What counts as the start of a project for you — the first commit, the domain purchase, the evening you decided — and what counts as each state label being reached. Rows measured to different boundaries do not compare, and you will not remember which you meant.
- Add a Reference class section at the foot of
learning/monetizing/PORTFOLIO.md. The ledger does not carry one — you are adding it. It goes at the foot rather than beside a project because it is not about any project. - Write one row per project: its state label, its hours, and one sentence naming what stopped it or consumed it. One sentence. If it needs a clause defending you, rewrite it until it does not.
- For every row that stopped short of a stranger being able to reach it, add the second story in a sentence: what made finishing look likely at the start. Aim for structural causes — an unchecked competitor, an invisible maintenance load, a distribution plan deferred to afterwards. If your sentence describes your character, you have written the first story.
- Now the fractions, written with their denominators and never as percentages. A percentage of seven claims a precision seven rows do not have. At minimum: how many reached built-and-live, how many have ever taken money, and how many carry a maintenance figure you actually measured rather than estimated.
- Where a fraction is zero, write the zero and one sentence beside it saying what was never attempted. A zero with its cause named is a finding. A zero on its own reads, later, like a missing entry.
- Write the spread of hours: the smallest, the largest, and the state label sitting next to the largest. Do not average them.
- Put the one measured maintenance figure next to the six estimates and say in one line whether the estimates sit above or below it. That comparison is the only estimate-against-reality row this ledger contains.
- Finish this sentence in the file, in your own words: on this record, a project I start ends up ___. If you cannot complete it from the rows above, the rows are not done.
- Then the adjustment for the next one. Name each structural change you are claiming, and beside each, whether it has yet produced anything. If the honest answer is that none of them has, write that. Then check the candidate section against the pursue cap.
- Bring the whole section into the chat. I will push hardest on two things: any sentence in the rows that is quietly defending you, and any adjustment for the next project that is enthusiasm wearing a process costume. The second one is the harder to catch from the inside, which is the reason to hand it to somebody else.
What you have, stated without inflation
That is the end of the reading, and the version of this ending that flatters you is the one that stops being useful the first time you hold it against the file. So, precisely.
You do not have revenue. Nothing in this course asked you to launch or to charge, and no exercise in it depended on your having done either. That was deliberate and it is the reason the framework is runnable at all: one that only works once money is already arriving is one that would have been read and never used.
What you have is a file and a page. The ledger holds the seven, and a candidate that was judged before it existed, each with a state, a cost, a named payer, a verdict with a date, and now a row in a reference class about you. The gates page holds the questions in an order, each with a pass test underneath it that is answered yes or no.
Two things keep running now that the lessons have stopped, and neither of them is a lesson. The review has a date on it, and the date arrives whether or not anything has happened — which is the entire reason it was a date rather than a trigger. The gates are the list you take every project through when it does, including the ones you have no feelings about, and including the ones whose answers you are sure you already know.
Everything written here was the argument for those two. The argument is finished. The running is not, and the next thing that happens is a date in a calendar rather than a page on this site.
Read this next — primary source
From Nobel Prize to Project Management: Getting Risks RightBent Flyvbjerg — Project Management Journal, vol. 37, no. 3, August 2006, pp. 5–15; the full text is free and open on arXiv, posted 14 February 2013
Read it whole because this lesson borrows the method and then fails its own entry requirements, and you should watch that happen in the original rather than take my word for it. The paper states the procedure as three numbered steps, and the first two ask for a class broad enough to be statistically meaningful and for credible data on enough cases to draw statistically meaningful conclusions from. Hold your ledger against those two sentences as you read them. The other reason to read it is the closing section on potentials and barriers, which splits forecasters into two kinds and treats them completely differently: those making honest mistakes, where Flyvbjerg says barriers are low because nobody has reason to be against a method that improves their own forecasts, and those whose wrong numbers are doing a job for somebody, where his verdict is that the demand for accuracy is simply not there. Read both descriptions and work out which one you are. That answer, not the arithmetic, decides whether the exercise at the foot of this lesson is worth an evening.
Stuck, curious, or think this lesson is wrong? Ask your teaching agent. The lessons are the scaffold; the conversation is where the learning gets unstuck.