Rerunner Early access

Security

Rerunner is designed to run coding agents on code from pull requests. That is useful, and it is risky: a pull request can change what the agent executes. This page sets out the threat model, how Rerunner is designed around it, and how to run it safely.

Threat model

Treat every run as running untrusted code from the pull request.

What a run can execute

A coding agent does more than read text. During a run it may:

  • execute hooks configured in the repository, for Claude Code in .claude/settings.json;
  • start MCP servers configured in the repository, for example in .mcp.json;
  • run shell commands, install packages and run tests, as the task requires.

Claude Code's documentation is explicit that a headless session runs the project's hooks and connects its MCP servers even in a folder you have never trusted. Its --bare mode avoids that, but it also skips CLAUDE.md and the rest of the project's setup, which is exactly what Rerunner is built to test. So isolation has to come from the environment, not from the agent. Claude Code docs: headless

Who can change it

Anyone who can open a pull request. A pull request can add a hook, add or change an MCP server, edit a rules file to steer the agent, or change the suite's checks.

What's at stake

  • Secrets in the runner's environment, including the model API key and the code host token.
  • Network access from the runner to internal systems.
  • Spending on the model provider account.
  • Instructions hidden from reviewers in context files, for example with invisible Unicode.

How it's designed

Runs in your CI

Rerunner is designed to run on your runners, under your accounts. There is no hosted Rerunner service that receives your code.

Fresh environment per run

Each run is meant to start in a disposable container or runner and be thrown away afterwards, so one run can't affect the next.

No automatic fork runs

The CI check is designed never to run automatically on pull requests from forks.

Least privilege

The code host token needs read access to contents and write access to pull request comments and checks, nothing more.

Bounded cost

A per-run cost guard in the suite, on top of the spending limit on your own key.

MCP servers from the pull request

We're exploring record-and-replay for MCP servers, so a run doesn't have to start a server that the pull request introduced. Until then, the sandbox is the boundary.

Sandbox guidance

What we'll ask design partners to set up.

  • A fresh sandbox for every run: an ephemeral CI runner, container or virtual machine, discarded afterwards.
  • No production secrets: no deploy keys, cloud credentials or customer data in the environment.
  • A dedicated model API key with a spending limit, used only for these runs.
  • A least-privilege code host token: read contents; write pull request comments and checks.
  • Restricted network access where your CI supports it, so a run reaches the model provider and your package registry, and little else.

Pull requests from forks

A fork can change what the agent executes, and a fork's author is outside your organisation. Rerunner is designed to run automatically only on pull requests from branches in the same repository.

Don't use a CI trigger that runs fork pull requests with your repository's secrets. To test a change from a fork, review it first, then start a run yourself, in a sandbox.

Keys and spending

Every run is a real call to your model provider, so cost grows with tasks × runs × two sides. Use a key created only for Rerunner, set a spending limit on it at the provider, and keep the suite's per-run cost guard. If a run is ever compromised, the damage is capped and the key is easy to revoke.

Data handling

Design statements. The product is in development.

  • Rerunner is designed to send nothing to us: no code, no telemetry, no usage data.
  • Results stay where your CI puts them: logs, pull request comments and files on the runner.
  • The coding agent sends prompts and repository content to its model provider, under your account and that provider's terms, just as it does when you use it directly.
  • If we ever add a hosted service or any telemetry, we'll update this page and the privacy page before it ships, and say what is collected and why.

What we can't promise

A sandbox limits the damage untrusted code can do. It doesn't make that code safe. The context checks for hidden Unicode and secrets are meant to protect what your agents read; they don't make the runner safe.

Rerunner has not had an external security audit yet. We'll say so here until it has.

Reporting a vulnerability

Email security@rerunner.dev. Please report privately rather than in public. We'll reply, keep you informed while we fix it, and credit you if you'd like.

Machine-readable contact details: /.well-known/security.txt