Rerunner Early access

Built on Claude Code

This page is the technical detail of where Claude fits. In short: Rerunner is designed to test what happens when a team changes the instructions, tools or model behind its AI coding agent, by having the agent do the same real tasks before and after the change. Claude Code, Anthropic's coding agent, is the first agent we're building for, so every test run will be Claude Code doing that work. If you're new to the field, start with the problem and the glossary.

Below: where Claude fits in the pipeline, which documented features of Claude Code and the Claude API we build on, and how Rerunner complements Anthropic's own tools.

The pipeline

Numbered circles mark where Claude comes in. A dashed outline marks something planned for after the first version.

The Rerunner pipeline and where Claude fitsA pull request that edits agent context goes to the context diff (touchpoints 1 to 3: rules and skills, MCP config and tools, and planned Claude API checks), then to the regression run, where Claude Code runs headless on main and on the pull request (touchpoint 4) and each run reports InstructionsLoaded, system/init, result cost and turns, and benefits from prompt caching (touchpoints 5 to 8). Results go to verdicts and the report (touchpoint 9, optional OpenTelemetry). Root cause analysis, on the roadmap, re-runs tasks with each change restored.Pull requestEdits CLAUDE.md,AGENTS.md, rules,skills, MCP config,settings or the modelContext diffWhat the agent will seedifferently, and thedeterministic checks1Rules and skills2MCP config and tools3Claude API checksRegression runSame tasks, same pinned model, base and PR4Headless runs, three per side by defaultmainclaude -pclaude -pclaude -pPRclaude -pclaude -pclaude -pEach run reports5InstructionsLoaded: which files, and why6system/init: skills, MCP servers, errors7Result: cost, tokens and turns8Prompt caching across repeated runsVerdicts and reportFisher exact test per task,one pull request commentand a check status9OpenTelemetry, optionalRoot causeRoadmap: restore each changeand re-run the task The Rerunner pipeline and where Claude fitsThe same pipeline, top to bottom: pull request, context diff (touchpoints 1 to 3), regression run with Claude Code headless on main and the pull request (touchpoints 4 to 8), verdicts and report (touchpoint 9), and root cause on the roadmap.Pull requestEdits CLAUDE.md, AGENTS.md, rules,skills, MCP config, settings or the modelContext diffWhat the agent will see differently,and the deterministic checks1Rules and skills2MCP config and tools3Claude API checksRegression runSame tasks, same pinned model4Headless runs, three per sidemainclaude -pclaude -pclaude -pPRclaude -pclaude -pclaude -pEach run reports5InstructionsLoaded: which files, and why6system/init: skills, MCP servers, errors7Result: cost, tokens and turns8Prompt caching across repeated runsVerdicts and reportFisher exact test per task, one pullrequest comment and a check status9OpenTelemetry, optionalRoot causeRoadmap: restore each change andre-run the task

Nine touchpoints

Each one is a documented, public part of Claude Code or the Claude API.

  1. Rules and skills: what changes

    Rules in .claude/rules/ with a paths: glob load only when the agent reads or edits a matching file. Skills keep their name and description in context and load their body on demand. These are exactly the pieces that differ between base and pull request, so the context diff is designed to resolve them the way Claude Code does.

    Docs: memory and rules, skills

  2. MCP config and tool descriptions

    Which MCP servers start, and how their tools are described, shape what the agent does. Both are part of the setup under test: a pull request that edits a tool's description will get the same base-against-PR run as one that edits CLAUDE.md.

    Docs: headless mode

  3. The Claude API, for checks that need judgment Planned

    Some problems need a model to spot, such as two instruction files that contradict each other. These checks will run on the Claude API with a model such as Sonnet or Haiku. Every finding must quote the files verbatim, and a finding whose quote can't be found in the file is dropped. They are reported, never blocking.

  4. Headless runs with claude -p

    Every task will run non-interactively, with JSON output, on the base branch and on the pull request. We deliberately won't use --bare: it skips CLAUDE.md and the rest of the project's setup, which is the thing under test.

    Docs: headless mode

  5. The InstructionsLoaded hook

    Claude Code fires this hook whenever an instruction file enters context, with the file and the reason: session_start, nested_traversal, path_glob_match, include or compact. For path-scoped rules it also reports the matching glob and the file that triggered the load. That turns “did the rule load?” into a fact, which is why we designed expect_loaded to block a merge. One gap: an AGENTS.md read through Claude Code's fallback setting doesn't fire the hook, so there we'll rely on the documented loading rules.

    Docs: hooks, memory

  6. The system/init event

    In streaming JSON output, the init event lists the session's model, tools, skills, plugins and MCP servers, with any server or plugin that failed to load. A server that stops connecting on the pull request is a deterministic failure, not a flaky test.

    Docs: headless mode

  7. Cost, tokens and turns

    The result of each run reports its cost, token usage including cache reads and writes, and number of turns. Rerunner will report them per task and per side, so you'll see cost per successful task, not only pass rates.

    Docs: headless mode

  8. Prompt caching

    Repeated runs of a task share the same prefix: system prompt, tool definitions and the context under test. That is what prompt caching is built for. Cache reads are billed at a fraction of the base input price, and the cache lives for five minutes by default, so runs of the same task will be scheduled close together.

    Docs: prompt caching

  9. OpenTelemetry, if you already use it

    Claude Code can export usage and cost metrics, attributable to skills, plugins and MCP servers. Teams that already collect them will be able to see Rerunner's runs next to everyday usage.

    Docs: monitoring usage

A run, as Claude Code reports it

What Rerunner is designed to record for a single run, before any statistics.

account-export PR #482, run 3 of 6
SourceWhatDetail
InstructionsLoaded CLAUDE.md session_start
InstructionsLoaded src/api/CLAUDE.md nested_traversaltriggered by src/api/account.ts
InstructionsLoaded .claude/rules/api.md path_glob_matchsrc/api/**/*.ts, triggered by src/api/account.ts
Not loaded docs/testing.md No include eventloaded in 6 of 6 base runs; the import target is missing on this branch
system/init Skills: release MCP server docs: connected
Tools Read, Edit, Bash Never opened src/services/auth.ts
Result 14 turns Cost and token usage recorded
Check npm run lint:auth Exit 1src/api/account.ts:12 reads req.cookies outside AuthService
Illustration — not captured output. Field names follow Claude Code's documented hook and output formats.

Why Claude Code first

It shows its work

Claude Code reports which instruction files loaded and why. In our review of the documentation for Codex, Cursor and Copilot, we found no equivalent event. Without it, a test can only guess whether a rule was seen.

It's built to run headless

claude -p with JSON output, a startup event and per-run cost is what a CI check needs. Nothing has to be scraped from a terminal.

It has the most to get right

Nested CLAUDE.md files, imports, path-scoped rules, skills, hooks, MCP servers and settings. The richer the setup, the more ways an edit can change behavior, and the more a test is worth.

Alongside Anthropic's tools

Each tool answers a different question. Use them together.

claude plugin eval

Tests a plugin or skill in isolation, with and without it. By design it doesn't load the project's CLAUDE.md, .claude/ directory or .mcp.json. Rerunner is designed to test exactly that project setup, base against pull request, on every pull request that changes it. Plugin evals docs

Code Review

Reviews code changes and posts its findings as comments on the pull request, without approving or blocking it. Rerunner is designed to answer a different question: does a change to the agent's own setup change what the agent does? It's meant to be a pre-merge regression check that teams run alongside Code Review, on the same pull request. Code Review docs

/doctor prompt-audit

Reads your instruction files and flags stale references and contradictions. Rerunner is designed to add the other half: running real tasks to see whether behavior changed. Memory docs

We're building Rerunner to help teams change their Claude Code setup with confidence. That should make rolling Claude Code out across more repositories and teams safer, which is good for the teams using it.

Rerunner is independent. We are not affiliated with, partnered with or endorsed by Anthropic. We build on Claude Code's public, documented interfaces, and we'll update this page when they change.

Test-drive your agent's changes before they ship.

We're looking for a few teams to try Rerunner on one real repository and tell us where it's wrong.