Rerunner Early access

In development · early access

Your AI coding agent changed. Did anyone notice?

Teams steer AI coding agents with written instructions, tools and models. One small edit can quietly change how the agent works, and normal tests won't catch it. Rerunner is designed to run the same real tasks before and after every change and show what got worse, before it ships.

Starts with Claude Code, with more coding agents to follow.

Proposed change Tidy up the agent's instructions
  1. Write a test for every new page.
  2. −Always use the login service to check who is signed in.
  3. Keep each change small.
Each task ran six times with the old instructions and six times with the new ones.
TaskOld versionNew versionResult
Data download 6/6 1/6 Worse
Invoice PDFs 6/6 6/6 Same
Changelog entry 2/6 6/6 Better
Date handling 4/6 2/6 Not sure yet

Every code test still passes, because no code changed. The agent's behavior changed anyway.

Illustration — not captured output. Each task ran six times on each version. A filled square is a run that passed.

One small edit can quietly change how the agent works.

Companies now let AI coding agents write code. Each agent follows a written handbook the team keeps for it: which rules to follow, which tools to use, which model runs. Teams edit that handbook all the time. Here is what one harmless-looking edit can do.

The edit A tidy-up of the agent's instructions
  1. −Always use the login service to check who is signed in. Never read the login cookie directly.

Review: approved. It looks like a harmless tidy-up.

The same task given to the agent with each version

“Add a page where signed-in users can download their data.”

Old instructions
const user = await authService.requireSession(req);

Uses the login service. Passes the security check in 6 of 6 runs.

New instructions
const sid = req.cookies["session"];
const user = await db.users.bySession(sid);

Reads the login cookie itself. Fails the security check in 5 of 6 runs.

Illustration — not captured output. The code that was being reviewed didn't change, so the usual tests still pass.

Tests check the code, not the agent

A team's automated tests run on the product's code. An edit to the agent's handbook changes no code, so every test stays green.

The edit looks harmless

Deleting a line, adding a tool or upgrading the model reads like housekeeping. Reviewers approve what they read; nobody sees what the agent will do.

The cost shows up later

Weeks later it surfaces as a security hole, a broken feature or a bigger bill, long after anyone remembers the edit.

More on the problem, with sources

How Rerunner will solve it

A test-drive for every change to the agent's handbook, before it's accepted.

  1. You change the agent's instructions

    Someone edits the handbook, adds a tool or switches the model, and proposes the change for review, as they do today.

  2. We'll run the same real tasks on the old and new version

    The agent will do a handful of your real tasks with the old handbook and with the new one, several times each, because an agent doesn't do exactly the same thing every time.

  3. You'll see what got worse, better, or is unclear, before merging

    Each task will get a plain answer: Same Worse Better or Not sure yet, with the counts behind it. Only failures that happen every time can block a change. Anything that could be chance is reported, not enforced.

It's like testing an edited pilot checklist in a simulator before real flights: the change gets tried where a mistake costs nothing.

How it works, in detail

Who it's for

Companies where many teams and repositories already use AI coding agents. There, one failure nobody noticed costs more than the tool that would have caught it.

Platform teams

You keep the shared handbook every team's agent follows. Know a change helps before it reaches every repository at once.

Security teams

Catch the edit that makes an agent skip a security step, or that hides an instruction from reviewers, before it ships.

Engineering leads

Approve changes to the agent's instructions based on what the agent actually did, not on how the edit reads.

AI rollout leads

See whether a new tool or model really helps your teams, and what it costs per task, before you roll it out.

What each team gets

Why now

Agents are doing real work in most companies, and the handbooks that steer them have to be measured, not assumed.

  • AI agents are now part of everyday software work. In Stack Overflow's April 2026 pulse survey, 59% of respondents use AI agents at work, up from 31% a year earlier.

    Source: Stack Overflow blog, May 2026

  • Most companies use more than one. 91% of organisations have at least two AI coding tools in active use. Each agent reads its handbook by its own rules: one cuts a long instruction file off at a size limit, another skips a file entirely when its own is present. The same edit can help one agent and hurt another.

    Sources: GitLab 2026 AI Accountability Report, reported by IT Brief; Codex docs: AGENTS.md; Claude Code docs: memory

  • More instructions don't automatically help. A 2026 study found that instruction files for coding agents did not generally improve task success, and raised cost by over 20% on average. The authors recommend validating any gains rather than assuming them.

    Source: Evaluating AGENTS.md, arXiv 2602.11988

More evidence, with sources

What it is, and isn't

Where things stand today.

What Rerunner is

  • A test for changes to an AI coding agent's setup: its instructions, tools and model.
  • Designed to run as a check in your own CI, under your accounts, for code on GitHub or GitLab.
  • Vendor-neutral. It starts with Claude Code; other agents are on the roadmap.
  • In development, with early access for a few design partners.

What it isn't

  • A coding agent. We don't write your code; we're building a test for the setup you give your agent.
  • A hosted service that holds your code. It's designed to run on your own machines.
  • A single score. Each task will get its own answer, with the counts behind it.
  • A finished product with customers. We're pre-revenue, with no customers yet. Nothing on this site is a customer quote or a benchmark result.

Built on Claude

Claude Code, Anthropic's coding agent, is the first agent Rerunner is built for and will be the engine of every test run. It reports which instructions it loaded and what each run cost, which is what makes a precise test possible. The problem isn't specific to one vendor, and neither is Rerunner.

Where Claude fits: the technical detail

Test-drive your agent's changes before they ship.

We're looking for a few teams to try Rerunner on one real repository and tell us where it's wrong.