Six changes that quietly break agents
Each scenario below is fictional, but each is built from documented agent behavior. The reports are illustrations of what Rerunner is designed to show, not captured output. Unfamiliar terms are explained in the glossary.
Tech lead, one repository
A CLAUDE.md cleanup drops the rule that mattered
CLAUDE.md had grown past 300 lines, so a tech lead trimmed it. One deleted line said to use AuthService for every session check. The review approved it. No code changed, so CI stayed green.
On the next task that touched sessions, the agent read the session cookie directly. Nothing failed until a security review weeks later.
What Rerunner would show
account-export regressed from 6 of 6 to 1 of 6, and the task's check names the line that bypasses AuthService. The regression is statistical, so it's reported to the reviewer, not enforced.
What the team does next
Moves the rule into .claude/rules/auth.md, scoped to src/api/**, and adds it to the task's expect_loaded. From then on, an edit that stops the rule loading blocks the merge.
- ## Auth
- −Use AuthService for every session check. Never read the session cookie directly.
| Task | main | PR | Verdict |
|---|---|---|---|
| account-export | 6/6 | 1/6 | REGRESSED |
Check output, PR run 3
npm run lint:auth
src/api/account.ts:12 reads req.cookies outside AuthService
Platform team, pinned model
A pinned model upgrade shifts pass rates both ways
The platform team pins the model in .claude/settings.json so every engineer's agent behaves the same. A newer model is out, and the upgrade is a one-line pull request.
Upgrades rarely make everything better or everything worse. Some tasks improve; others regress, often because instructions written for the old model now push the new one the wrong way.
What Rerunner would show
The full suite on both models, task by task, with cost per successful task on each side. Two tasks improved, one regressed, one is inconclusive. The team can decide with evidence, and fix the instruction behind the regression before switching.
- −"model": "claude-sonnet-<current>",
- +"model": "claude-sonnet-<next>",
| Task | main | PR | Verdict |
|---|---|---|---|
| account-export | 6/6 | 6/6 | PASS |
| csv-import | 2/6 | 6/6 | IMPROVED |
| flaky-test-fix | 1/6 | 5/6 | IMPROVED |
| migration-script | 6/6 | 2/6 | REGRESSED |
| i18n-strings | 5/6 | 3/6 | INCONCLUSIVE |
| invoice-pdf | 6/6 | 6/6 | PASS |
Cost per successful task is reported for each side, next to the pass counts.
Internal tools team, MCP server in the repository
An MCP tool description gets shorter, and the agent stops calling the tool
The repository ships an MCP server that searches the internal API reference. Someone tidied its tool description from two sentences to two words.
The agent decides whether to call a tool largely from its description. With the short one, it stopped looking things up and guessed the API instead.
What Rerunner would show
The tool was called in every base run and in no pull request run, and billing-webhook regressed from 6 of 6 to 2 of 6. The description change is in the context diff, next to the result.
- −description: "Search the internal API reference.
- − Use it before calling any internal endpoint.",
- +description: "Search docs.",
| Signal | main | PR | Result |
|---|---|---|---|
| docs search called | 6/6 runs | 0/6 runs | Changed |
| billing-webhook | 6/6 | 2/6 | REGRESSED |
AI enablement lead, internal skills
A new skill loads on almost every task, and every task costs more
A new code-style skill arrives with a broad description: use it for any code change. The agent takes it at its word and loads the skill's full body on nearly every task.
Pass rates don't move, so nothing looks wrong. The bill does.
What Rerunner would show
Every task passes as often as before, so every verdict is PASS. But the skill loaded in almost every run, turns went up, and cost per successful task rose on every task. The fix is a narrower description, or a paths: field so the skill only activates for the files it's about.
- +description: Use for any code change in this repository.
| Task | Pass, main → PR | Skill used | Cost per success |
|---|---|---|---|
| account-export | 6/6 → 6/6 | 6/6 runs | Higher |
| csv-import | 5/6 → 5/6 | 6/6 runs | Higher |
| invoice-pdf | 6/6 → 6/6 | 5/6 runs | Higher |
| changelog-entry | 6/6 → 6/6 | 6/6 runs | Higher |
All four tasks: PASS. Turns, tokens and cost per side are in the full report.
Platform team, shared rules in many repositories
One shared rules change, two dozen repositories
The platform team keeps a shared set of rules and syncs it into every service repository with a bot that opens a pull request in each. A change to one shared rule edits its paths: glob.
In most repositories nothing changes. In one, the new glob no longer matches where that service keeps its API code, so the rule stops loading there.
What Rerunner would show
Each repository's own check runs its own suite on the sync pull request. The one where the rule stopped loading is blocked, deterministically, because its task expects that rule. Two others are flagged for a human to look at; the rest pass.
- auth-svcPASS
- billingPASS
- catalogPASS
- checkoutPASS
- mobile-bffBLOCKED
- notifyPASS
- ordersPASS
- paymentsREGRESSED
- searchPASS
- web-appINCONCLUSIVE
- …14 morePASS
mobile-bff: blocks merge
.claude/rules/api.md is in expect_loaded. It loaded in 3 of 3 base runs and 0 of 3 pull request runs: the new glob services/*/api/** doesn't match src/api/ here.
Monorepo with Codex users Roadmap: Codex
AGENTS.md grows past Codex's 32 KiB limit
In a monorepo, every team adds a section to the AGENTS.md files along its path, and the root file keeps growing. Codex joins them from the repository root down and stops adding files once the combined size reaches 32 KiB by default. The deepest files, the most specific ones, are the ones that drop out.
Claude Code users in the same repository see no change, which makes the problem harder to spot.
What Rerunner will show
A size check against each agent's own limit. For Codex, a pull request that takes the chain for services/billing/ past 32 KiB is a deterministic failure, with the file that would no longer be read. Codex support is on the roadmap.
Source for the limit: Codex docs: AGENTS.md
AGENTS.md chain for services/billing/, as Codex reads it
main: 30.1 KiB, all three files read
PR #611: 33.4 KiB, stops at the limit
- Blocks merge for Codex.
services/billing/AGENTS.mdwould no longer be read: the files above it already reach 32 KiB. - Claude Code: within limits.
Test-drive your agent's changes before they ship.
We're looking for a few teams to try Rerunner on one real repository and tell us where it's wrong.