Rerunner Early access

Six changes that quietly break agents

Each scenario below is fictional, but each is built from documented agent behavior. The reports are illustrations of what Rerunner is designed to show, not captured output. Unfamiliar terms are explained in the glossary.

Tech lead, one repository

A CLAUDE.md cleanup drops the rule that mattered

CLAUDE.md had grown past 300 lines, so a tech lead trimmed it. One deleted line said to use AuthService for every session check. The review approved it. No code changed, so CI stayed green.

On the next task that touched sessions, the agent read the session cookie directly. Nothing failed until a security review weeks later.

What Rerunner would show

account-export regressed from 6 of 6 to 1 of 6, and the task's check names the line that bypasses AuthService. The regression is statistical, so it's reported to the reviewer, not enforced.

What the team does next

Moves the rule into .claude/rules/auth.md, scoped to src/api/**, and adds it to the task's expect_loaded. From then on, an edit that stops the rule loading blocks the merge.

PR #482 Tidy up CLAUDE.md
CLAUDE.md+3 −9
  1. ## Auth
  2. −Use AuthService for every session check. Never read the session cookie directly.
TaskmainPRVerdict
account-export6/61/6REGRESSED

Check output, PR run 3

npm run lint:auth
src/api/account.ts:12  reads req.cookies outside AuthService
Illustration — not captured output.

Platform team, pinned model

A pinned model upgrade shifts pass rates both ways

The platform team pins the model in .claude/settings.json so every engineer's agent behaves the same. A newer model is out, and the upgrade is a one-line pull request.

Upgrades rarely make everything better or everything worse. Some tasks improve; others regress, often because instructions written for the old model now push the new one the wrong way.

What Rerunner would show

The full suite on both models, task by task, with cost per successful task on each side. Two tasks improved, one regressed, one is inconclusive. The team can decide with evidence, and fix the instruction behind the regression before switching.

PR #517 Upgrade pinned model
.claude/settings.json+1 −1
  1. −"model": "claude-sonnet-<current>",
  2. +"model": "claude-sonnet-<next>",
TaskmainPRVerdict
account-export6/66/6PASS
csv-import2/66/6IMPROVED
flaky-test-fix1/65/6IMPROVED
migration-script6/62/6REGRESSED
i18n-strings5/63/6INCONCLUSIVE
invoice-pdf6/66/6PASS

Cost per successful task is reported for each side, next to the pass counts.

Illustration — not captured output. Model names are placeholders.

Internal tools team, MCP server in the repository

An MCP tool description gets shorter, and the agent stops calling the tool

The repository ships an MCP server that searches the internal API reference. Someone tidied its tool description from two sentences to two words.

The agent decides whether to call a tool largely from its description. With the short one, it stopped looking things up and guessed the API instead.

What Rerunner would show

The tool was called in every base run and in no pull request run, and billing-webhook regressed from 6 of 6 to 2 of 6. The description change is in the context diff, next to the result.

PR #530 Clean up docs server
tools/docs-mcp/server.ts+1 −2
  1. −description: "Search the internal API reference.
  2. − Use it before calling any internal endpoint.",
  3. +description: "Search docs.",
SignalmainPRResult
docs search called6/6 runs0/6 runsChanged
billing-webhook6/62/6REGRESSED
Illustration — not captured output.

AI enablement lead, internal skills

A new skill loads on almost every task, and every task costs more

A new code-style skill arrives with a broad description: use it for any code change. The agent takes it at its word and loads the skill's full body on nearly every task.

Pass rates don't move, so nothing looks wrong. The bill does.

What Rerunner would show

Every task passes as often as before, so every verdict is PASS. But the skill loaded in almost every run, turns went up, and cost per successful task rose on every task. The fix is a narrower description, or a paths: field so the skill only activates for the files it's about.

PR #544 Add code-style skill
.claude/skills/code-style/SKILL.mdadded
  1. +description: Use for any code change in this repository.
TaskPass, main → PRSkill usedCost per success
account-export6/6 → 6/66/6 runsHigher
csv-import5/6 → 5/66/6 runsHigher
invoice-pdf6/6 → 6/65/6 runsHigher
changelog-entry6/6 → 6/66/6 runsHigher

All four tasks: PASS. Turns, tokens and cost per side are in the full report.

Illustration — not captured output.

Platform team, shared rules in many repositories

One shared rules change, two dozen repositories

The platform team keeps a shared set of rules and syncs it into every service repository with a bot that opens a pull request in each. A change to one shared rule edits its paths: glob.

In most repositories nothing changes. In one, the new glob no longer matches where that service keeps its API code, so the rule stops loading there.

What Rerunner would show

Each repository's own check runs its own suite on the sync pull request. The one where the rule stopped loading is blocked, deterministically, because its task expects that rule. Two others are flagged for a human to look at; the rest pass.

rules-sync checks on 24 sync pull requests
  • auth-svcPASS
  • billingPASS
  • catalogPASS
  • checkoutPASS
  • mobile-bffBLOCKED
  • notifyPASS
  • ordersPASS
  • paymentsREGRESSED
  • searchPASS
  • web-appINCONCLUSIVE
  • …14 morePASS

mobile-bff: blocks merge

.claude/rules/api.md is in expect_loaded. It loaded in 3 of 3 base runs and 0 of 3 pull request runs: the new glob services/*/api/** doesn't match src/api/ here.

Illustration — not captured output. Each repository's report lives on its own pull request; this is the view from opening them.

Monorepo with Codex users Roadmap: Codex

AGENTS.md grows past Codex's 32 KiB limit

In a monorepo, every team adds a section to the AGENTS.md files along its path, and the root file keeps growing. Codex joins them from the repository root down and stops adding files once the combined size reaches 32 KiB by default. The deepest files, the most specific ones, are the ones that drop out.

Claude Code users in the same repository see no change, which makes the problem harder to spot.

What Rerunner will show

A size check against each agent's own limit. For Codex, a pull request that takes the chain for services/billing/ past 32 KiB is a deterministic failure, with the file that would no longer be read. Codex support is on the roadmap.

Source for the limit: Codex docs: AGENTS.md

PR #611 Expand the root AGENTS.md

AGENTS.md chain for services/billing/, as Codex reads it

main: 30.1 KiB, all three files read

PR #611: 33.4 KiB, stops at the limit

032 KiB limit
  • Blocks merge for Codex. services/billing/AGENTS.md would no longer be read: the files above it already reach 32 KiB.
  • Claude Code: within limits.
Illustration of a planned check — not captured output.

Test-drive your agent's changes before they ship.

We're looking for a few teams to try Rerunner on one real repository and tell us where it's wrong.