Rerunner Early access

The problem / Essay

Why agent changes need tests

A coding agent is its context, and its context changes every week in pull requests that are reviewed by reading. That's the gap Rerunner exists to close.

Essay

Ask a team how a code change reaches production and you'll hear about types, tests, review and staged rollouts. Ask how a change to CLAUDE.md reaches the agent that now writes a good share of their pull requests, and the answer is shorter: someone reads it and approves.

That gap used to be harmless, when assistants completed lines and a person accepted each one. It isn't harmless now. Agents take a task, read the repository, run commands, edit files and open pull requests. What they do is shaped by files that change as often as the code.

An agent is its context

A coding agent is a model, a harness, and whatever the harness loads: instruction files and the docs they import, rules scoped to parts of the repository, skills, MCP tools, hooks and settings. Change any of it and you've changed the program. The model is the part that changes least often. The context changes every week, in pull requests written by people trying to help.

Those pull requests look like documentation, so they get reviewed like documentation. But reading a CLAUDE.md edit tells you what the file says, not what the agent will do. Whether a rule loads depends on a glob. Whether a skill fires depends on how its description reads to the model. Whether an instruction still matters depends on what else is in context. None of that is visible in a diff.

The evidence cuts both ways

The tempting conclusion is that more instructions make a better agent. The evidence says otherwise. A 2026 study of AGENTS.md files found that context files did not generally improve task success, and raised inference cost by over 20% on average.1 Chroma's work on “context rot” found that model performance degrades as input grows, even on simple tasks.2

And yet Vercel reported that an 8 KB AGENTS.md index of documentation took its agent from 53% to 100% on framework APIs the model had not seen in training.3

Both can be true, and that is the point. The effect of context depends on the repository, the tasks, the model and the agent. A rule that helps in one codebase is noise in another. Trimming a file can save money, or lose the one instruction that mattered. There is no general answer to “is this change good?”, only the answer for your repository.

Cost behaves the same way. Fewer tokens don't guarantee a smaller bill: in one study, the setup that cut tool-output tokens the most, by 38.4%, raised the billed cost by 6.8%, with caching rather than raw token count dominating input-side cost.4 The metric that holds up is cost per successful task, measured rather than estimated.

Silent by default

What makes this a testing problem, not a style problem, is that context fails quietly. In our own testing with Claude Code 2.1.293, a CLAUDE.md import pointing at a file that doesn't exist raised no warning; the file just wasn't loaded. Codex stops adding AGENTS.md files once their combined size reaches 32 KiB by default.5 A typo in a rule's YAML header turns a rule meant for one directory into one that is always on.6

None of this turns a unit test red, because no code changed. You find out later, from a worse pull request or a security review, and by then nobody remembers the edit that caused it.

What a test looks like

If context is code, it deserves what code gets: run it before you merge it. In practice that means three things.

First, diff what the agent will actually load, not just the text of the files. Resolve imports and rules the way the agent does, then check what can be checked without a model: broken references, hidden Unicode, pasted secrets, size limits. These checks are deterministic and cheap, and they are safe to enforce.

Second, run real tasks on both sides of the change: same agent, same pinned model, same starting files, with the pull request as the only difference. Tasks should be small and recent, and scored by a check you already trust, usually your own tests.

Third, be honest about noise. Agents are stochastic. A task that passes three of three times on main and two of three on the branch has told you almost nothing. Report counts, use a test suited to small pass counts (a one-sided Fisher exact test works), and say “inconclusive” when that's the truth. Then enforce only what doesn't depend on chance. A rule that loaded on every run on main and on no run on the branch is a fact, not a fluctuation.

That last part matters more than it looks. A merge gate that fails on noise gets switched off within a week. A gate that blocks only on facts, and reports everything else with its evidence, can stay on.

Neutral, and Claude Code first

Most organisations now run more than one AI coding tool; one 2026 report put it at 91%.7 Each tool reads the same files by different rules, so a test of your agent setup shouldn't belong to any one vendor.

We start with Claude Code anyway, for a practical reason: it reports which instruction files it loaded and why, and it runs headless with machine-readable output and per-run cost. That makes precise tests possible today. Other agents will follow.

What we don't know yet

We don't know how often setup changes cause regressions in practice. As far as we can tell, nobody has measured it well. It may be weekly for a company with a large shared rules set and rare for a small repository. That's the first question our design partners will help answer, and we'll publish what we find, including if the answer is “less often than we thought”.

What we do know is the shape of the problem. Changes to an agent's setup are changes to a program. Programs get tested. The rest is engineering.

Sources

  1. Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?, arXiv 2602.11988 (February 2026). Back
  2. Chroma, Context Rot: How Increasing Input Tokens Impacts LLM Performance (July 2025). Back
  3. Vercel, AGENTS.md outperforms skills in our agent evals (January 2026). Back
  4. Token Reduction Is Not Cost Reduction, arXiv 2607.12161 (July 2026). Back
  5. Codex documentation: AGENTS.md. Back
  6. Claude Code documentation: memory and rules. Back
  7. GitLab 2026 AI Accountability Report, reported by IT Brief (June 2026). Back

Test-drive your agent's changes before they ship.

We're looking for a few teams to try Rerunner on one real repository and tell us where it's wrong.