Surviving the third why
Most positions hold for one follow-up question and collapse on the second — how to write one that still has something underneath it three levels down.
Your position survives the first question. It usually does — the first question is people checking they heard you right. The second one is where you find out whether there is anything underneath it, and the third one is where most positions in most rooms quietly die.
The death is rarely dramatic. It looks like a slightly louder version of what you already said, or a new adjective, or the phrase “it just depends on the context.” Everyone notices. Nobody says anything. You get less airtime for the rest of the session.
This lesson is a drill for that. It is this course’s own device rather than a published methodology, and it is not borrowed from any named technique — no lineage is being claimed for it.
What a collapse actually looks like
Take the position this whole course is pointed at, and push on it three times.
- The claim. “An advisory system should give a confident answer rather than hedge on everything.”
- First why. “Because a hedge attached to every answer stops carrying information. If everything is uncertain, uncertainty is not a signal any more.” Good. Still standing.
- Second why. “Because users start skipping the hedges entirely, and then the hedge on the one answer that really was shaky gets skipped along with the rest.” Still standing, and now it is a mechanism rather than a preference.
- Third why. “So… being confident is better, because trust.” That is the collapse. The claim got louder and gained an abstract noun, and it stopped being about anything checkable.
The failure is not that the answer was wrong. It is that the third level repeated the first level with more emphasis. There was nothing new underneath.
What surviving looks like instead
The move that works at the third level is almost never a stronger assertion. It is a change of variable — noticing that the thing being argued about was the wrong thing to argue about, and naming the thing that actually governs the outcome.
There is a published example of exactly that move, and it is the study this lesson sends you to read. Fregosi, Vicente, Campagner and Cabitza, at AAAI-26, ran a within-subjects study with 184 participants on a logic-puzzle decision-support task. Their finding, in their words: “well-calibrated confidence scores significantly improved decision accuracy (+20%…), whereas miscalibrated scores yielded minimal accuracy gains (+2%…) and increased vulnerability to automation bias and conservatism bias.”
Do not file this on either side
The tempting move is to grab that study for whichever side of the confident-versus-hedged argument you already hold. It does not belong to either side. It says the operative variable is calibration, and that both extremes have a cost: a confidently wrong answer harms because it is miscalibrated, not because it is confident, and a hedge that is itself miscalibrated — hedging on things the system reliably gets right — produces conservatism bias, which is a cost of its own. Force-fitting it into “evidence against confidence” would misrepresent the paper, which is the failure this whole course was built to prevent. It is a third position, and it is more useful than winning the original argument.
With that in hand, the third level has somewhere to go: the argument I was having was the wrong one. It is not confident versus hedged, it is calibrated versus not. Our review gate had per-field confidence, which is the calibrated-per-item case; the advisory chatbot hedged uniformly, which is the miscalibrated case in the other direction. That is the axis I would design on.
That answer is deeper than the one it replaced, and it is also less convenient — it concedes that your original framing was too simple. Positions that survive probing usually get more specific and less flattering as they go down. That is what depth feels like from the inside.
Three probes that do the killing
Also this course’s own drill. Run them on your own position before anyone else does, in this order:
- “Why does that matter?” Forces you from a statement to a consequence. A position with no consequence attached is a taste.
- “What if the opposite were true?” Forces you to describe the world in which you are wrong. If you cannot describe it, you are not holding a position, and the next lesson is about that specifically.
- “What is the variable?” The one that usually saves you. It asks what quantity the outcome moves with. “Confidence” is not a variable; it is a tone. Calibration is a variable, because it can be measured, and something can have more or less of it.
The counterfeit third level
The most common fake depth is an abstract noun standing in for a mechanism: it is really about trust, it comes down to agency, this is a transparency question. Each of those feels like a level down and is a level sideways. The test is simple: could two people who disagree both say your sentence? If they could, you have not said anything. Nobody argues for less trust.
The transfer question, and having its answer ready
Once you cite the AAAI study, a good room asks the obvious thing: 184 people doing logic puzzles is not brokers picking award flights. That objection is correct, and it is the fourth why.
The answer is to say what you think transfers and what you think does not, rather than defending the whole study. The mechanism transfers — a confidence signal only helps if it tracks correctness, and a uniform one carries no information in either domain. What does not transfer is the effect size. Plus-twenty points on a puzzle task tells me nothing about what to expect on an advisory financial recommendation, and I would not quote the number as if it did.
That is a fourth level. It exists because you read the method section rather than the abstract, which is nearly the whole trick.
Check your recall
Answer from memory — no scrolling back.
Retrieval check
Why does a position that gets more specific and less flattering as it is probed read as stronger than one that stays constant?
Check your answer
Because staying constant under pressure is what a slogan does, and getting more specific is what an understanding does. A person who has actually thought about something has found its edges, and the edges are where the concessions live. When your second answer narrows the claim and your third answer changes the variable, the room learns that you built the position rather than adopted it.
The inverse tells the room the same thing in the other direction. A claim that survives three questions completely unchanged has usually survived by not engaging with any of them.
Hands on
Run the three probes on your own row
Done when: One row in POSITIONS.md has three levels written under its position cell — consequence, the world in which you are wrong, and the variable the outcome moves with — and the third level names something measurable rather than an abstract noun.
- Take the position cell you have been building and write your one-line claim at the top of a scratch block under the row.
- Answer why does that matter in one sentence. Then answer it again for that sentence. Two levels, written down, before you go anywhere near the third.
- Answer what if the opposite were true. Describe the world concretely: what would users be doing, what would the interface look like. If you cannot fill this in, stop and note that — the next lesson is entirely about this gap.
- Answer what is the variable. It has to be something that can be measured and that something can have more or less of. Cross out any answer containing the words trust, agency, transparency or experience unless you can immediately say how you would measure it.
- Bring all three levels into the chat. I will push on the third one specifically, and I will ask the transfer question: what part of your evidence carries to a case you have not shipped, and what part does not.
What this does not cover
The drill tests whether a position has depth. It does not test whether the position is falsifiable, which is a different property and the subject of the last lesson in this course: naming, in advance, the specific evidence that would move you off the claim. That lesson also picks up the second probe above, since “what if the opposite were true” is the question it turns into an exercise. The substantive argument about hedging and confidence itself — the evidence on each side, and where it stops — is the over-hedging lesson in the previous module.
Read this next — primary source
Too Sure for Our Own Good: A User Study on AI Confidence and Human RelianceCaterina Fregosi, Lucia Vicente, Andrea Campagner & Federico Cabitza — University of Milano-Bicocca and IRCCS Ospedale Galeazzi–Sant’Ambrogio; AAAI-26, the Fortieth AAAI Conference on Artificial Intelligence, Singapore, 20–27 January 2026; n = 184. Fetched 2026-09-05.
Read this one because it is what a third-level answer looks like when someone does the work properly. The study does not pick a side in the confident-versus-hedged argument; it changes the variable to calibration and then measures what happens on both sides of it. Read the method section for the task the 184 participants were actually doing — logic puzzles with decision support — because the distance between that and an advisory travel chatbot is exactly the kind of transfer question you will be asked, and having the answer ready is half of surviving a probe.
Stuck, curious, or think this lesson is wrong? Ask your teaching agent. The lessons are the scaffold; the conversation is where the learning gets unstuck.