FA
ForgeAcademy
← All 21 free courses
ForgeAcademy · Advanced Tier · Free preview
Advanced Scrum Masters
12 sessions · Part Two — Agents
SESSION 9 OF 12 · FREE TO READ
9
When Agents Go Wrong⚠️ Guardrails
Judgment · The failure modes nobody demos
Free preview
🏆
Your win this session
You'll be able to spot the three ways agents fail in Agile work — and you'll leave with a review step that actually survives a busy sprint, instead of 'be careful' as a principle.
🔄
Every vendor demo shows you the agent succeeding. None of them show you the sprint report that read beautifully and quietly dropped the one blocker that mattered. That's the session. If you only take one thing from this course, take this one — it's the difference between using agents and being used by them.
What this session covers
  • The real failure mode: not being wrong, but being wrong fluently — output that reads like a competent colleague wrote it
  • Failure 1, omission: the summary that's accurate about everything it mentions and silent about what mattered most
  • Failure 2, fabrication: agents inventing tickets, releases and names when the data doesn't contain what you asked for
  • Failure 3, stale confidence: answers drawn from a view of the board that was true on Tuesday
  • Why 'tell me what you're unsure about' doesn't work — self-reported confidence doesn't track actual error
  • Four prompt-level guardrails: scope hard, cite ticket IDs, declare absence, flag inference vs. read
  • Blast radius: sorting your workflows by what a wrong answer actually costs
  • Review steps that survive a busy sprint — and why the diligent ones never do

The demo problem

Every agent demo you have ever seen was a success case. The agent reads the board, writes the summary, everyone nods. That's not dishonest — it's just not the part you need to prepare for.

Here's the part you need to prepare for. An agent writes your sprint summary. It's well structured. The tone is right. Every sentence in it is true. And it doesn't mention that the integration work is blocked, because the blocker was recorded as a comment on a subtask rather than as a status change, and the agent didn't look there.

You skim it. It reads fine. You send it to your stakeholders.

Nothing about that output looks wrong, because nothing in it is wrong. That's the failure mode. Not error — omission, delivered fluently.

Three ways it actually breaks

Omission. The most dangerous and the least discussed. Wrong sentences get caught; missing ones don't. Nobody proofreads for absence. And the things most likely to go missing are exactly the things recorded off the happy path — a blocker in a comment, a dependency in a linked ticket, a concern someone raised verbally in standup.

Fabrication. Ask for something the data doesn't contain and you may get it anyway, complete with plausible names and dates. Practitioners in the Atlassian community have reported agents producing entirely fictional releases when the underlying data had nothing to offer. The invented material is not flagged as invented. It looks exactly like the real material.

Stale confidence. The agent answers from whatever slice of the board it pulled, with no sense that the world moved. In a live sprint, "as of some point earlier today" and "now" are different answers, and only one of them is useful in a stakeholder meeting.

One Rising Star in the Atlassian community described this class of tool as being like an intern — brilliant and high-energy, "terrible at taking final responsibility." That's the most accurate summary of the current state of things I've read.

Why asking it doesn't help

The instinct is to add "tell me what you're unsure about" to the prompt. It feels responsible. It doesn't work — the model's self-reported uncertainty doesn't reliably track where it's actually wrong. You get a confident answer with a confident-sounding caveat attached, and you've learned nothing.

You cannot ask the thing that was confidently wrong to notice it was confidently wrong. Verification has to come from outside the output.

Four guardrails that do work

These don't make the agent right. They make checking cheap, which is the thing that determines whether you actually check.

1. Scope hard. Current sprint only. This board. This label. A narrow question has a narrower space to be wrong in.

2. Make it cite. Every claim tagged with the ticket ID it came from. This is the highest-leverage one — it turns verification from "reread everything and hope" into "click three links." Anything that can't be sourced becomes visible immediately.

3. Make it declare absence. "List anything you looked for and could not find." This is the direct counter to the omission failure, and almost nobody does it.

4. Make it flag inference. "Mark anything you inferred rather than read." Separates what's in Jira from what merely sounded plausible.

Four guardrails that make checking cheap Four steps applied to a prompt: scope hard, so there is less ground to be wrong on; make it cite, so verification is clicking three links; make it declare absence, which is the direct counter to omission; make it flag inference, separating what is in Jira from what merely sounded plausible. They don't make the agent right. They make checking cheap. 1 Scope hard Current sprint. This board. Less ground to be wrong on. 2 Make it cite Every claim tagged with its ticket ID. Verify = click 3 links. 3 Declare absence “What did you look for and not find?” Counters omission. 4 Flag inference Marks what it guessed versus what it read straight from Jira. You cannot ask the thing that was confidently wrong to notice it was confidently wrong. Verification has to come from outside the output.

Blast radius: stop reviewing everything

"Review all AI output" is a policy that lasts about three sprints. Then a release goes sideways, everyone's busy, and it quietly stops. Reviewing everything is how you end up reviewing nothing.

Sort by what a wrong answer costs instead:

  • Cheap to be wrong — internal drafts, first-pass refinement notes, discussion prompts. Ship it, glance at it, move on.
  • Expensive to be wrong — anything leaving your team, anything a stakeholder plans against, anything touching governance, compliance, or an audit trail. Always human eyes.

A retro summary that's slightly off wastes five minutes. A capacity forecast that's confidently wrong plans your next sprint into a wall.

Blast radius: sorting AI work by what a wrong answer costs A single axis from cheap to expensive. Cheap: internal drafts, first-pass refinement notes, discussion prompts — ship it, glance at it, move on. Expensive: anything leaving your team, anything a stakeholder plans against, anything touching governance, compliance or an audit trail — always human eyes. Sort by what a wrong answer costs — not by how much output there is CHEAP TO BE WRONG • Internal drafts • First-pass refinement notes • Discussion prompts Ship it. Glance at it. Move on. EXPENSIVE TO BE WRONG • Anything leaving your team • Anything a stakeholder plans against • Governance, compliance, audit trail Always human eyes. cost of a wrong answer a retro summary wastes five minutes · a capacity forecast plans your sprint into a wall
"Review all AI output" lasts about three sprints. Reviewing everything is how you end up reviewing nothing.

Review steps that actually survive

Three properties separate the review habits that last from the ones that die in March.

Structural, not diligent. If it depends on remembering to be careful, it decays. If it's a step in something that already happens, it survives.

Attached to an existing ceremony. The sprint report gets checked in the five minutes before review, because that slot already exists and has a forcing function. A standalone "AI review checkpoint" on your calendar will be gone by the third sprint.

Author separated from verifier. Whoever ran the prompt is the worst person to catch its errors — they already believe it. Where the stakes justify it, someone else looks.

What to never delegate

Short list, and it's short on purpose.

Anything you'd have to defend under scrutiny without being able to say where the number came from. Anything about a specific person — performance, capability, blame. Anything where the failure is discovered by someone other than you, months later.

The test isn't "can AI do this?" It's the one put well in that same community thread: the real shift is from can it do this to how would I know it did it well? If you can't answer the second question, you're not ready to delegate the first.
⚡ Try it with your team
Take a sprint summary an AI has already written for you — or generate one now from your current board. Then run the omission check below against it. You're not looking for wrong sentences. You're looking for the thing that isn't there.
"Here is a sprint summary you generated: [paste it]. Now do three things. First, list every claim in it and the specific ticket ID that supports it — mark any claim you can't source. Second, list anything you looked for in the data and could not find. Third, list anything you inferred rather than read directly. Do not rewrite the summary."
Check yourself
1. What makes agent output in Agile work most dangerous?
2. Why doesn't 'tell me what you're unsure about' work as a safeguard?
3. Which guardrail directly counters the omission failure mode?
4. What makes a review step survive a busy sprint?
The whole course
The other eleven sessions are not written yet. The course ships 15 October 2026 — this list is the shape of what you would be buying, not a menu of pages sitting behind a paywall. Session 9 is the only one you can read today, and it is free.
Part One — Systems, Not Prompts
1
From Prompts to Pipelines
Ships Oct 15
2
Your Ceremony Operating System
Ships Oct 15
3
The Friday Machine — sprint reporting end to end
Ships Oct 15
4
Backlog Hygiene at Scale
Ships Oct 15
Part Two — Agents
5
What an Agent Actually Is
Ships Oct 15
6
Your First Agile Agent
Ships Oct 15
7
Agents Inside Your Stack — Copilot, Rovo, Jira automation
Ships Oct 15
8
Chaining It Together — intake → backlog → report
Ships Oct 15
9
When Agents Go Wrong
You are here
Part Three — Judgment & Defensibility
10
Metrics That Don't Lie
Ships Oct 15
11
Defending AI-Assisted Work
Ships Oct 15
12
Rolling It Out to a Team That Didn't Ask For It
Ships Oct 15
This is session 9 of 12 from Advanced Scrum Masters, a paid course shipping 15 October 2026. The other eleven sessions are not free. This one is.