Over-hedging versus confident wrongness
Scoping services when in doubt is a published guideline and it is sometimes wrong — the case for a confident answer over a hedged one, argued with the trust evidence that cuts against it too.
Here is the sentence this whole course is built to produce, and it lives in this lesson:
The HAX guidelines say to scope services when in doubt — that is G10, filed under “when wrong”. In practice, with brokers, over-hedging destroyed trust faster than a confident wrong answer did.
Two halves. The first is a citation anyone can check. The second is a disagreement from something you shipped. This lesson makes both halves defensible, and it will not let you win the argument, because the evidence does not.
What the canon says, and it says hedge
G10, Scope services when in doubt, verbatim from Table 1 of the CHI 2019 paper: “Engage in disambiguation or gracefully degrade the AI system’s services when uncertain about a user’s goals.” Disambiguation means ask. Graceful degradation means offer less. Both are the system deciding to show its uncertainty rather than resolve it.
Apple’s Generative AI guidance pushes the same way, from a platform vendor that sells the devices this guidance governs: “Raise awareness about and minimize the chance of hallucinations… it’s important to clearly communicate that AI-generated content may contain errors.” Disclose the uncertainty. Say the thing might be wrong.
So the canon is not neutral here. It has a direction, and the direction is toward showing doubt. Disagreeing with that is disagreeing with two of the three largest publishers in the field at once, which is worth knowing before you open your mouth.
The evidence for hedging is real and it is narrow
“I’m Not Sure, But…” (Kim, Liao, Vorvoreanu, Ballard & Vaughan, FAccT ’24, Rio de Janeiro, 3–6 June 2024) is a pre-registered experiment with N = 404 participants using a fictional LLM search engine to answer medical questions. It manipulates how the system expresses uncertainty. The finding, in the authors’ words: first-person expressions of uncertainty such as “I’m not sure, but…”
decrease participants’ confidence in the system and tendency to agree with the system’s answers, while increasing participants’ accuracy
which the paper attributes to “reduced (but not fully eliminated) overreliance on incorrect answers.”
Take that seriously. Hedging cost trust and bought accuracy, and the trade was worth it in that setting. Note two things before you carry it anywhere. Four of the five authors are Microsoft researchers, which is an interest worth naming even though it cuts against the vendor reflex rather than for it. And the task is medical questions, where an incorrect answer is expensive and users have reason to check. Your advisory chatbot is not that.
The evidence against confident wrongness, flagged as thin
There is a 2026 finding pointing the other way, and it needs its qualifiers attached every single time. Sascha Brodsky, writing for IBM Think on 27 July 2026, reports work by Capraro and colleagues in which wrong AI advice left participants “less accurate, more confident, and far less likely to say ‘I don’t know’” — with “accuracy declined by a factor of three, while confidence rose by a factor of 2.5.”
That is an unreviewed arXiv preprint (2607.13562), and this course reached it only through IBM’s secondary writeup. The preprint itself has not been read here. IBM also sells AI products and publishes Think as marketing-adjacent editorial. Cite it exactly that way or leave it out. A number quoted from a summary of a preprint is two hops from anything anyone has checked, and if you present it as a study you have handed the room a reason to distrust everything else you said.
The study that looks like evidence here and is not
NN/g’s “Prioritize Smarts over Sentience to Increase Trust with AI” (Evan Sunwall, 19 September 2025) summarises an academic study by Colombatto, Birch and Fleming with n = 410, and it is the largest number on this topic anywhere on the NN/g site. It reports that “perceptions of intelligence were positively related to taking ChatGPT’s advice” while “perceptions of emotion were negatively related” to it.
It is not evidence about hedging language, and using it as though it were is a mischaracterisation. The variable it measures is perceived intelligence versus perceived emotional capacity. That is a different axis from how a system phrases its uncertainty. The study is real and interesting and it belongs in a different argument. Reaching for it here because the number is big is exactly the failure mode this course exists to prevent.
The finding that dissolves the binary
The most recent and most useful result refuses to take either side. “Too Sure for Our Own Good: A User Study on AI Confidence and Human Reliance” (Fregosi, Vicente, Campagner & Cabitza, University of Milano-Bicocca; AAAI-26, Singapore, 20–27 January 2026, N = 184, a logic-puzzle decision-support task) reports:
well-calibrated confidence scores significantly improved decision accuracy (+20%…), whereas miscalibrated scores yielded minimal accuracy gains (+2%…) and increased vulnerability to automation bias and conservatism bias.
The operative variable is calibration, not tone. A confidently wrong answer is damaging because it is miscalibrated, not because it is confident. And a hedge is not free: hedging on things the system actually gets right produces conservatism bias, which is users rejecting correct answers. Over-hedging has a measured cost in that study, and it is not a cost to your brand voice, it is a cost to decisions.
This is the finding that makes the argument interesting rather than tribal. “Ship confident” and “always hedge” are both answers to the wrong question.
The position, which is a synthesis and not a citation
Everything above is other people’s work, stated as they stated it. What follows is this course’s own reading, and the labelling matters more here than anywhere else in the module, because the temptation to footnote your own opinion is strongest when the footnotes are good.
The defensible position is not that confidence beats hedging. It is that uncertainty should be expressed where it is measured and nowhere else. The review gate does this by construction: confidence is per-field, so a document with two shaky values shows doubt on two values and none on the rest. A blanket disclaimer over the whole extraction would be hedging on the fields the model got right, which is the conservatism-bias failure with a friendlier face.
That synthesis is yours. FAccT does not conclude it, AAAI-26 does not conclude it, and no guideline recommends it. What the evidence gives you is the shape of the claim; what makes it a position is that you built the thing and watched it work.
Retrieval check
Someone quotes G10 at your design and says the agent should ask a clarifying question whenever confidence drops. Give the full answer, both halves, in under sixty words.
Check your answer
G10 is real — scope services when in doubt, disambiguate or gracefully degrade. The 2026 AAAI work says the variable is calibration, not doubt: miscalibrated hedging produces conservatism bias, users rejecting correct answers. So we show uncertainty per-field where we measure it, and not over the whole document. On our advisory flow, blanket hedging cost more trust than a wrong answer did.
Four moves in that. Cite accurately. Introduce evidence the other person probably does not have. Name the mechanism rather than asserting a preference. Then ground it in something you shipped, and stop. The version that fails adds a fifth sentence claiming the research supports your design, which it does not.
Check the citations before you use them
Answer from memory — no scrolling back.
Hands on
Fill the row you actually have evidence for
Done when: The Microsoft row in POSITIONS.md carries G10 verbatim, a position stated as your synthesis rather than as a finding, an evidence cell that names the chatbot’s narrowness, and a change-my-mind cell specific enough that a real study could satisfy it.
- Copy G10’s number, title and qualifier into the citation cell from the paper itself. Add its phase, “when wrong”, since the placement is half the point.
- Write the position cell as a claim about where uncertainty is shown, not about whether. If the sentence works equally well as “be confident”, it is not the position yet.
- Fill the evidence cell from the flight chatbot and state the limits inside the cell: advisory only, no execution, request and response, one domain. Then add one sentence naming what you observed, in terms of what users did rather than what they felt.
- Write the change-my-mind cell. Start from the gap this lesson found: a study on a high-stakes advisory task, n above a hundred, showing hedged answers retained more users. Adjust the numbers to what would genuinely move you, and keep it falsifiable.
- Read the finished row out loud against the target register at the top of
POSITIONS.md. Bring it into the chat and I will push on it with the counter that actually lands — “the FAccT study says hedging improved accuracy, so are you arguing for worse decisions?” — which is the question your row has to already answer.
What this does not cover
This lesson argues one guideline in both directions. It does not teach the reading habit underneath it: telling a platform rule from a research finding, and knowing what weight each can carry. Apple asserted its hallucination guidance with no sample and no methodology, the FAccT paper reported an N and a pre-registration, and both got quoted here in the same paragraph. The platform-and-consultancy lesson takes that apart and is the last in this module. Before this one, the suggest-versus-act and halfway-failure lessons cover the two gaps where the canon says nothing at all rather than saying something you disagree with.
Read this next — primary source
“I’m Not Sure, But…”: Examining the Impact of Large Language Models’ Uncertainty Expression on User Reliance and TrustKim, Liao, Vorvoreanu, Ballard & Vaughan — FAccT ’24, Rio de Janeiro, 3–6 June 2024. Fetched 2026-09-05. N = 404, pre-registered. Four of the five authors are Microsoft researchers, which is worth naming even though the finding cuts against a vendor’s interest in confident-sounding products.
This is the study to have read before you argue either side, because it is the one that actually manipulates hedging language and measures what happens. Read the method section closely: the task is medical questions on a fictional LLM search engine, which is high-stakes and nothing like an advisory product recommendation. Knowing exactly how far the finding travels is what lets you cite it without overclaiming, and knowing where it stops is what gives you your own change-my-mind criterion.
Stuck, curious, or think this lesson is wrong? Ask your teaching agent. The lessons are the scaffold; the conversation is where the learning gets unstuck.