The estimate you can defend
A single number is a prediction with its confidence hidden. Here is how the outside view produces a range you can actually argue for.
You have a recommendation: buy the hosted provider, unblock the deal this quarter, accept the per-customer cost. You wrote it in a document with the alternatives laid out beside it. Then the VP of Sales reads the first paragraph and asks the only question they actually care about.
“Great — so how long?”
This is the part of planning where honest engineers most often get themselves into trouble, and it is worth being precise about why. The trouble is not that people lie. It is that the honest answer and the useful answer feel like different answers, and under mild social pressure most of us produce a third thing that is neither.
“Six weeks” is not an estimate
It is a prediction with its confidence hidden. Nobody hearing it can tell whether you mean six weeks, I have done this three times or six weeks, I think, if the thing I am worried about does not happen. Those are wildly different claims and they are being delivered in identical words.
The practical consequence is that a single number cannot be argued with. A reviewer who thinks it is optimistic has nothing to push on except your judgment, which makes the conversation personal. Everything good about the design doc — showing the reasoning so somebody can engage with it — you just threw away at the last step.
How wrong estimates are allowed to be
Before fixing it, it helps to know the size of the problem, because the honest answer is larger than most people expect and knowing the number makes it easier to say out loud.
Steve McConnell’s cone of uncertainty describes how estimate accuracy varies with how well-defined the work is. At initial concept — before requirements, before a design — estimates are wrong by as much as a factor of four in either direction. That is a sixteen-fold spread between the low and high ends (Construx, “The Cone of Uncertainty”). The cone narrows as definition improves, which is the useful part: the precision you can honestly offer is a function of how much you have pinned down.
The inside view is why your number is low
Asked how long something will take, nearly everyone does the same thing: break the work into tasks, guess each one, add them up. This is the inside view, and it fails in a specific, predictable direction.
It fails because a task breakdown can only contain the tasks you thought of. Everything that actually consumes the schedule — the IdP that implements SAML slightly wrong, the customer’s admin who takes two weeks to answer an email, the week you lose to an unrelated incident — is by definition absent from the list, because if you had thought of it you would have listed it. Summing a list of known work produces a confident number for a project that will also contain unknown work.
Kahneman and Tversky named this the planning fallacy. The fix they proposed is the outside view: instead of reasoning forward from this project’s parts, look at how long comparable projects actually took, and start there.
Bent Flyvbjerg turned that into a working method — reference class forecasting — and validated it on infrastructure projects at a scale that is hard to argue with: a dataset of more than 16,000 projects across 136 countries, in which the overwhelming majority of megaprojects come in late, over budget, or short on benefits (Flyvbjerg, “From Nobel Prize to Project Management”). The method is three steps: identify a class of comparable past projects, find the distribution of their actual outcomes, and place your project in that distribution.
Two reasons a number comes out wrong, and only one has a technique
Here is the distinction that separates people who estimate well from people who merely estimate carefully. Flyvbjerg’s paper attributes inaccurate forecasts to two causes, and they are not the same problem wearing different clothes.
Optimism bias is honest. You believe the number. You are simply reasoning from the inside view like everyone else, and the inside view runs low. This is a cognitive error, and the outside view is a genuine fix for it.
Strategic misrepresentation is not honest, and it is usually not personal either. It is what happens when an organisation reliably rewards low numbers — when the project that gets approved is the one that sounded cheapest, and everyone has learned this. The outside view does nothing about it. You can hand someone a beautifully-referenced forecast and watch it get negotiated down to the number that will be approved.
Where people get burned
If your estimate keeps getting pushed and no new information ever appears in the conversation, you are not in a disagreement about duration. You are in an incentive problem, and the fix is a written record of what you said and when — not a better spreadsheet. Notice which conversation you are actually in before reaching for a technique.
Retrieval check
Without scrolling up: name the two causes of an inaccurate estimate, and say which one the outside view actually fixes — and why it does nothing for the other.
Check your answer
Optimism bias and strategic misrepresentation. The outside view fixes optimism bias, because that is a reasoning error — replacing inside-view arithmetic with comparable real outcomes corrects it. It does nothing for strategic misrepresentation, because that is not a reasoning error at all: the number is being shaped by what gets approved, not by what is believed. A technique cannot fix an incentive; only accountability and a written record can.
What is the reference class for SSO?
Applying this to your project means answering one question honestly: what has this team actually done that resembles this?
Note what the reference class is not. It is not “how long does a SAML integration take” in the abstract, and it is not the vendor’s quickstart timing. The useful class is comparable work done by people in your circumstances: the last three third-party integrations your team shipped, measured from kickoff to the first customer actually using it in production — not to code-complete, which is the boundary that quietly hides half the schedule.
If that history does not exist, you have two honest moves and one dishonest one. You can borrow a reference class from elsewhere — another team, published integration timelines — and say that you borrowed it. Or you can state that you have no reference class and that your number is therefore inside-view only, which tells the reader exactly how much to trust it. The dishonest move is to produce the same confident number either way.
The range, and the branch that leads to its top
So you give a range. But a range on its own is worse than useless — it reads as hedging, and a reader who wants a date will simply take the bottom of it and quote that upward.
What makes a range credible is naming the branch. Say what specifically has to happen for the top end to be the real answer.
“Four to nine weeks. Four if their Okta admin is responsive and SCIM behaves the way the vendor documents it. Nine if we hit per-provider quirks in deprovisioning — which is the part the reference class says goes wrong most often, and the part we have never done before. I will know which by the end of week two.”
Read what that does. It is falsifiable. It tells the reader what to watch. It names a date when the range narrows, which converts an uncomfortable spread into a scheduled decision. And if week five arrives and you are in the nine-week branch, you are delivering news you already announced rather than a surprise — which is the entire difference between a schedule slip and a credibility loss.
Now apply it: estimate the SSO build
Hands on
Produce a range you would defend in the room
Done when: You have a range, a named reference class, an explicit branch to the top end, and a date when the range narrows.
- Write your reference class in one sentence: which past projects you are comparing to, and how you are measuring their duration. If you are borrowing or have none, say so here — that sentence is the honest part.
- From that class, write the bottom of your range — the duration if the work resembles the good end of what you have seen before.
- Now the top. Do not compute it by padding the bottom. Ask instead: what is the single thing most likely to go wrong here, and how long does it cost when it does? For this project, deprovisioning across identity providers is the honest candidate.
- Write the branch as a sentence a reader could later check you against: what has to be true for the top end to happen.
- Name the date by which you will know which branch you are on, and what you will have done by then to find out.
- Bring it into the chat. I’ll push the way a VP does — asking for the single number, and asking whether you can “just try for” the bottom of the range. Holding the range under that pressure is the skill this lesson is really about.
What this does not cover
You named one thing that could go wrong in order to build the range. That is not the same as having thought about risk properly — you picked the most likely one and stopped, because a range only needs the top branch. The full question is what else could make this plan wrong, which of those you can cheaply reduce now, and how you tell people about the ones you cannot. That is where this goes next.
Read this next — primary source
From Nobel Prize to Project Management: Getting Risks Right — Bent FlyvbjergProject Management Journal 37(3), 2006 — free on arXiv, a paper rather than a blog post
It is the source of the technique this lesson is built on, and it is short enough for an evening. Read it for the two-causes distinction in particular — the paper is unusually blunt that some bad estimates are honest mistakes and some are not, and that the fixes are different.
Stuck, curious, or think this lesson is wrong? Ask your teaching agent. The lessons are the scaffold; the conversation is where the learning gets unstuck.