Raising it without becoming the blocker
The architect who blocks every release stops being consulted, so a flag needs a severity, a proposed non-blocking mitigation, and an explicit statement of what you are not claiming.
There is an architect at every company who is right about everything and gets invited to nothing. The flags were correct. Each one arrived as an objection with a release date behind it, and after the third the room learned to schedule the review for a week he was travelling.
You have a correct observation and a named owner. This lesson is about the three things you attach to them so the flag lands as work somebody can do rather than as a veto they have to negotiate.
A standards body says the same thing, less politely
The temptation is to treat “do not block everything” as career advice, which makes it sound like a compromise of your judgement. It is not. NIST states the position directly in its risk-management framework:
“Attempting to eliminate negative risk entirely can be counterproductive in practice because not all incidents and failures can be eliminated.” (NIST AI RMF 1.0, 1.2.3)
The sentence that follows is the one about you specifically: “Unrealistic expectations about risk may lead organizations to allocate resources in a manner that makes risk triage inefficient or impractical or wastes scarce resources.” Every flag you raise spends somebody’s attention, and attention is the scarce resource being allocated.
Quote that passage without its counterweight and you have built a licence to never block anything, which is a different failure with worse consequences. Same section, a paragraph later:
“In cases where an AI system presents unacceptable negative risk levels – such as where significant negative impacts are imminent, severe harms are actually occurring, or catastrophic risks are present – development and deployment should cease in a safe manner until risks can be sufficiently managed.” (NIST AI RMF 1.0, 1.2.3)
Read the two together and you get a usable default rather than a rule: blocking is a real option reserved for imminent, occurring or catastrophic harm, and it costs you the ability to use it again if you spend it on a maybe.
The three things you attach
One: a severity, and an explicit answer on blocking
Severity here means your read, stated in a sentence, plus a direct answer to a question nobody will ask out loud: are you asking us not to ship? Answer it before they wonder. Most of the time the answer is no, and saying so unprompted changes how the rest of the flag is heard.
Do not reach for a number. CVSS and CISA’s stakeholder-specific categorisation both exist, both are real, and both score the exploitability and impact of a vulnerability that has already been found, to prioritise a fix. What you are holding is a governance gap — an unnamed owner, an unchecked masking default, an undecided retention period — which has no exploit, no CVE and no fix to prioritise. Attaching a 7.4 to it borrows precision the observation does not have, and the first person who knows what CVSS actually measures will say so.
What a severity sentence looks like instead:
“My read is medium: the surface shows identity documents, the recorder is running, and nobody has checked what it captures at our installed version. I am not asking to hold the release — I am asking for the answer before the next one.”
Two: a proposed mitigation small enough to say yes to
The proposal is not the fix. The owner owns the fix. The proposal is the smallest change that shrinks the blast radius while the owner decides, and its defining property is that it is cheaper than the argument about it.
OWASP’s LLM03:2026 Excessive Agency entry is a good shape to copy, because its mitigations repeat one phrase: “Limit the tools that LLM agents are allowed to call to only the minimum necessary”, “Limit the functions that are implemented in LLM tools to the minimum necessary”, “Limit the permissions that LLM tools are granted to other systems to the minimum necessary”. None of those is a solution to excessive agency. Each one narrows what a failure reaches, and each one is proposable in a sentence.
Your front-end versions are the same species:
- Turn auto-rendering of model-supplied images off by default in the kit, leaving it available behind an explicit opt-in. Reversible, and it moves a decision from a default to a signature.
- Pin the telemetry SDK’s privacy configuration explicitly rather than inheriting the default, so a routine major-version bump cannot move it silently.
- Exclude one surface from recording while the broader question is answered, rather than proposing the product stop using the tool.
- Render the confirmation from the executed payload, which is a change in one component and not a change to the agent runtime.
Notice what these have in common. Each is inside something you already own, each is reversible, and none requires the owner to agree with your diagnosis before it can ship.
Three: what you are not claiming
The column that keeps you credible with people who do this for a living. It is not modesty, it is precision, and it does more work than the other two combined.
| What you saw | What you are not claiming |
|---|---|
| A sink, and untrusted content reaching it | That you have found an exploit |
| A recorder on a surface showing personal data | That data has leaked, or that anyone was careless |
| No stated retention period for transcripts | That the retention is unlawful. You are not qualified to say that, and saying it costs you the room. |
| A new telemetry vendor on an audited product | That adding it breaks compliance. It probably does not. |
Blocking is one of four options, and the other three close too
The framework’s MANAGE function lists what an organisation can do with a risk it has identified: “Risk response options can include mitigating, transferring, avoiding, or accepting.” Only one of those four looks like holding a release.
This is the sentence that resolves the tension the whole lesson is about. You are not choosing between raising a flag and letting something ship. Accepting a risk is a legitimate response, and a flag that ends in a recorded acceptance has done its job completely. The reason to raise it was never to stop the release — it was to make sure the acceptance happened on purpose, by somebody who was allowed to make it.
Check your recall
Answer from memory — no scrolling back.
Retrieval check
A product owner replies: “Fine, we accept the risk.” Have you lost?
Check your answer
No. Accepting is one of the four response options a risk framework lists, and it is a real outcome rather than a brush-off — provided the person accepting is the role that can, and provided it is written down somewhere other than a thread.
The thing to check is not whether they said yes to your concern. It is whether the acceptance is recorded, and whether the person who recorded it knew what they were accepting. An acceptance based on your one-sentence limit statement is an informed one. An acceptance based on “the architect seemed worried” is not, and it will be re-litigated the first time anything goes wrong.
Hands on
Rewrite one flag so it cannot be heard as a veto
Done when: One flag log row rewritten with three additions: a severity sentence that answers the blocking question explicitly, a proposed mitigation that sits inside something you already own and is reversible, and a “what I am not claiming” sentence. Read aloud to somebody who has not seen the surface, who can then say what you are asking for in one sentence.
- Pick the row you feel most strongly about. That is the one most likely to arrive as an accusation, which makes it the one worth rewriting.
- Write the severity sentence. Your read, the reason for it, and then a direct yes or no on whether you are asking to hold a release. If the answer is yes, name which of NIST’s three conditions you think applies — imminent significant impact, harm occurring, or catastrophic risk. If none applies, the answer is no.
- Write the proposal. One change, in code or configuration you already own, reversible in an afternoon. Check it against the test: could this ship even if the owner disagrees with your diagnosis?
- Write the limit. Start with “I am not claiming” and finish it honestly, including the part where you cannot tell how bad it is.
- Read all three aloud to someone who has never seen the surface. Ask them to say back, in one sentence, what you are asking for. If they say “you want them to stop,” the rewrite has failed and the severity sentence is usually why.
- Update the row. The severity, proposal and limit columns should now be full for at least one surface, with only the pass condition still empty.
What this does not cover
Everything here is about getting the flag heard. It says nothing about how it ends. A well-routed, well-scoped, non-blocking flag with no written pass condition gets a reassuring reply, drops out of the thread, and comes back to you in three months as the same unease with less credibility attached. The last lesson in this course, on what a satisfactory answer sounds like, is about the sentence that closes a row for good — including how to recognise the answers that sound like closure and are not.
It has also not told you how to decide whether a mitigation is technically correct. That is deliberate and it is the ceiling: the moment you start authoring the control rather than proposing the smallest reversible narrowing of it, you have taken ownership of something you cannot maintain across ninety products.
Read this next — primary source
Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 — section 1.2.3, Risk PrioritizationNational Institute of Standards and Technology, January 2023 — free PDF, fetched 2026-09-05. Same document the ownership lesson used; if you read the GOVERN table there, this is a different two pages.
This lesson takes section 1.2.3 and the MANAGE 1 subcategories: the argument that eliminating negative risk entirely is counterproductive, the counterweight naming when deployment should stop, and the four response options that make “accepting” a legitimate outcome rather than a failure. Read 1.2.3 in full — it is under a page — then the MANAGE rows of the Core table. Where it stops: NIST is written for an organisation prioritising its own portfolio of risks, not for one architect deciding how loudly to raise one. It gives you the frame and none of the wording.
Stuck, curious, or think this lesson is wrong? Ask your teaching agent. The lessons are the scaffold; the conversation is where the learning gets unstuck.