Skip to content

Repository files navigation

Claude Review Relay

Claude Review Relay lets Codex ask Claude Code for an independent, read-only review of a local Git change. It is a focused Go MCP server for the workflow:

  1. Codex implements and tests a change.
  2. Claude independently reviews the server-computed diff.
  3. Codex evaluates the findings and fixes only confirmed problems.
  4. Claude resumes the exact same conversation and verifies the corrections.

The server keeps a durable association between its review_id and Claude Code's explicit session_id. Every follow-up uses claude -p --resume <session_id>. It never uses the ambiguous claude --continue command.

The repository is named claude-review-relay. The installed binary, MCP server identifier, tool prefix, configuration directory, and persisted data paths retain the technical name claude-reviewer for compatibility.

Why This Project Exists

Claude Code already provides useful building blocks:

  • claude -p runs a non-interactive prompt and can return JSON;
  • the same CLI can resume a specific session with --resume, while --continue selects the most recent conversation in the current directory;
  • claude mcp serve makes general Claude Code tools available to an MCP client;
  • local /code-review gives quick feedback on a working diff, while /review <pr> reviews a pull request;
  • Ultrareview runs a deeper, independently verified multi-agent review in a remote cloud sandbox and can include staged and uncommitted changes;
  • Claude Code GitHub Actions can review pull requests in GitHub workflows.

Those capabilities remain appropriate for interactive work, general-purpose Claude delegation, quick local feedback, deep cloud review, and hosted pull-request automation. This project covers a different integration problem: making Claude a repeatable second reviewer inside a Codex implementation loop. It is orchestration and safety plumbing, not a claim to outperform Ultrareview's multi-agent bug-finding depth.

Approach Best suited for What the caller still manages
claude -p One-off scripts and prompts Git diff construction, prompt policy, schemas, session IDs, errors, and persistence
claude mcp serve General Claude Code tool access Review-specific scope, lifecycle, durable review identity, and verification workflow
/code-review or /review Quick interactive diff or PR feedback in Claude Code Codex integration and a durable correction-verification workflow
Ultrareview Deep, remotely sandboxed, multi-agent pre-merge review Codex integration, local MCP lifecycle, and same-session verification after Codex fixes
Claude GitHub review workflows Pull requests hosted on GitHub Local uncommitted work and the Codex-side correction loop
Claude Review Relay Local Codex cross-review Codex evaluates findings and remains responsible for all file modifications

Claude Review Relay adds the following review-specific behavior:

  • Codex-native MCP tools with typed inputs and outputs;
  • local review of tracked uncommitted work, staged work, or changes since a Git reference, without requiring a pull request;
  • literal include and exclude path scopes computed by the server;
  • a strict read-only Claude tool policy;
  • structured findings validated against a JSON Schema;
  • durable review metadata and explicit-session continuation;
  • asynchronous reviews that can run longer than a synchronous MCP connector deadline;
  • safe error details, redaction, output limits, process cancellation, and cross-process storage locking.

Choose Ultrareview when independently reproduced multi-agent findings and a cloud task that survives terminal closure are the priority. Choose this MCP server when the priority is an automated Codex policy: narrow a local diff, invoke a read-only reviewer, persist an explicit review identity, let Codex fix confirmed findings, and ask the same Claude conversation to verify those fixes. The two approaches can also be used together for high-risk changes.

This is defense in depth, not a security boundary. Secret detection can miss unknown formats, and any reviewed content is sent to the configured Claude Code service.

What Can Be Reviewed

The server always compares one base_ref with the repository's current working tree. base_ref defaults to HEAD.

Uncommitted tracked changes

Use base_ref: "HEAD". The review includes both staged and unstaged changes to tracked files. Untracked files are reported by repository-relative name, but their contents are not sent automatically.

Typical tool arguments:

{
  "repository_path": "/absolute/path/to/repository",
  "goal": "Review the current uncommitted implementation.",
  "base_ref": "HEAD",
  "include_paths": ["internal", "cmd", "README.md"],
  "effort": "xhigh",
  "timeout_seconds": 1200
}

Everything changed since a commit or branch

Set base_ref to a commit SHA, tag, or branch such as origin/main. The review covers the difference from that reference through the current HEAD, plus any current tracked working-tree changes.

{
  "repository_path": "/absolute/path/to/repository",
  "goal": "Review all changes introduced since the main branch.",
  "base_ref": "origin/main",
  "include_paths": ["internal", "cmd"]
}

One exact commit

The API does not accept a separate to_ref: its right-hand side is always the current working tree. To review one exact non-merge commit, use a clean checkout or temporary Git worktree at that commit and set base_ref to its parent:

current checkout:  <commit-sha>
base_ref:          <commit-sha>^

For a merge commit, select the intended parent explicitly, for example <commit-sha>^1. Any additional working-tree changes at that checkout are also part of the diff, so keep it clean when exact commit isolation matters.

Path-scoped changes

include_paths and exclude_paths contain literal repository-relative files or directories. They apply to both the tracked diff and the untracked filename list. Use a narrow scope for faster, more focused reviews.

How It Works

Codex
  |
  | start_review(repository, base_ref, scope, goal, test results)
  v
Claude Review Relay
  |-- validates the repository and literal path scope
  |-- computes and sanitizes the Git diff locally
  |-- persists the pending review record
  |-- starts a read-only Claude worker
  `-- returns pending immediately
          |
          | captures and persists Claude's explicit session_id
          | get_review_status(review_id)
          v
      structured verdict and findings
          |
          | Codex confirms and fixes valid findings
          v
  start_continue_review(same review_id, refresh_diff: true)
          |
          `-- claude -p --resume <same-session-id>

The MCP server is a local STDIO process started by Codex, not a permanent macOS daemon. It normally remains alive for the lifetime of its Codex client. If the server shuts down during a review, it cancels the Claude subprocess and stores the review as interrupted when a resumable Claude session ID was captured, or failed otherwise.

When Claude rejects a request because its usage quota is exhausted, the server does not keep a process asleep until reset. It persists waiting_for_quota, the structured retry_at timestamp when Claude Code reports one, and the explicit Claude session ID when available. At or after that timestamp—or when capacity is available if no timestamp was reported—Codex calls start_retry_review with the same review_id. The retry uses --resume when a session exists; otherwise it replays the initial review under the same local review identity.

Long reviews use background workers so the initial MCP call returns in a few seconds even when Claude needs several minutes. Polling reads the persisted state; it does not restart Claude or create a new conversation.

Core Guarantees

  • MCP over STDIO, with stdout reserved for the protocol and JSON logs on stderr;
  • MCP annotations identify metadata tools as read-only and close_review as destructive so Codex can apply its configured approval behavior;
  • the server, not the model, computes the Git diff;
  • Git commands use separate arguments, not a shell;
  • Claude receives only Read, Glob, and Grep; editing, Bash, Web, and MCP tools are explicitly denied;
  • prompts and diffs are sent through stdin rather than command-line arguments;
  • responses are constrained by JSON Schema and validated again in Go;
  • session records use atomic writes, 0600 permissions, and advisory file locks;
  • per-review OS-backed leases prevent concurrent continuations of one review;
  • complete private keys reject the request, sensitive filenames are excluded, and common token forms are redacted;
  • a diff larger than the configured limit fails explicitly and is never silently truncated;
  • resumed output must report the same Claude session ID as the requested one.

Prerequisites

  • macOS on Apple Silicon or Intel;
  • Go 1.25 or newer;
  • Git;
  • Claude Code installed and authenticated with claude auth login;
  • Codex CLI or another MCP client with STDIO support.

This guide and the production smoke test were validated with Claude Code 2.1.217. claude-reviewer doctor checks the installed CLI's authentication and required flags directly instead of relying only on a version number.

Build and Test

go build -o ./bin/claude-reviewer ./cmd/claude-reviewer
go test ./...
go vet ./...

Or run:

make check

Install on macOS

Install the binary atomically in the current user's local bin directory:

mkdir -p "$HOME/.local/bin"
cp ./bin/claude-reviewer "$HOME/.local/bin/claude-reviewer.new"
chmod +x "$HOME/.local/bin/claude-reviewer.new"
mv -f "$HOME/.local/bin/claude-reviewer.new" "$HOME/.local/bin/claude-reviewer"

Add the server to Codex with its absolute executable path:

codex mcp add claude-reviewer -- "$HOME/.local/bin/claude-reviewer" serve
codex mcp list

Minimal equivalent Codex server configuration, replacing the username:

[mcp_servers.claude-reviewer]
command = "/Users/USERNAME/.local/bin/claude-reviewer"
args = ["serve"]

Codex also supports optional per-tool approval overrides. To pre-approve the tools that start or finalize review state, extend the server configuration with:

[mcp_servers.claude-reviewer.tools.start_review]
approval_mode = "approve"

[mcp_servers.claude-reviewer.tools.start_continue_review]
approval_mode = "approve"

[mcp_servers.claude-reviewer.tools.start_retry_review]
approval_mode = "approve"

[mcp_servers.claude-reviewer.tools.close_review]
approval_mode = "approve"

The supported values are auto, prompt, writes, and approve; see the official Codex manual's MCP configuration section.

Do not use $HOME literally in the TOML command; shell expansion is not guaranteed there. Restart every running Codex client after replacing the binary. An already running MCP process continues using the executable version it loaded at startup.

Either add $HOME/.local/bin to the shell PATH, or use the absolute commands shown below.

Diagnose the Installation

"$HOME/.local/bin/claude-reviewer" doctor

The JSON report checks the Claude executable and version, authentication, required flags, Git, data-directory access, and session storage.

Run the production review pipeline against an isolated one-line Git diff with:

"$HOME/.local/bin/claude-reviewer" doctor --review-smoke-test

The smoke test calls the configured Claude models. It can incur cost and take several minutes. The equivalent opt-in Go integration test is:

CLAUDE_REVIEWER_INTEGRATION=1 go test ./internal/smoke -run TestInstalledClaudeReview

Use It from Codex

Nothing needs to be started manually after MCP registration. Codex starts the STDIO server when it needs the configured tools. You can ask for a review directly, for example:

Ask Claude Review Relay to review the current uncommitted changes in internal/
and cmd/. Analyze every finding, fix confirmed problems, and ask the same
Claude session to verify the corrections.

For a committed range:

Ask Claude Review Relay to review everything changed since origin/main. Limit
the scope to internal/session and internal/reviewer, and include the Go test and
vet results in the review context.

Normally you do not call the MCP tools or write their JSON arguments yourself: Codex does that. Add the policy block below to AGENTS.md when every non-trivial change in a project should follow the cross-review workflow automatically.

Recommended Asynchronous Workflow

  1. Call start_review with the repository, functional goal, literal file scope, and test results.
  2. Save the returned review_id, claude_session_id when present, expected_response_sequence, and poll_after_seconds.
  3. Poll get_review_status no more frequently than poll_after_seconds until status is no longer pending.
  4. If the status is waiting_for_quota, stop polling. Preserve the returned review_id, claude_session_id, and retry_at when present. At or after retry_at, or when capacity is available if it is absent, call start_retry_review with the same review_id and a message that restates the interrupted operation, then resume polling.
  5. Independently validate every finding. Claude is a reviewer, not an authority.
  6. Apply confirmed corrections and rerun the relevant tests.
  7. Call start_continue_review with the same review_id, a correction summary, and refresh_diff: true.
  8. Poll until the operation is no longer pending. Require the expected response sequence and the same Claude session ID.
  9. Call close_review after accepting the final verdict.

Consumption Reporting

review_diff, continue_review, and get_review_status return a usage object so turn and token budgets can be tuned from measurements rather than guesses:

  • last: what the most recent Claude invocation consumed — num_turns, max_turns, input_tokens, output_tokens, cache_read_input_tokens, cache_creation_input_tokens, total_cost_usd, duration_ms;
  • total: the same counters accumulated over the whole review, including continuations and retries. Turn counts are per invocation, so total reports the peak rather than a sum;
  • turn_budget_used_percent: how much of max_turns the last invocation spent. Values that approach 100 mean the next review of similar size is likely to exhaust its budget before returning a verdict;
  • billed_input_equivalent: the whole review expressed in input-token units, counting cache writes at 2x (the CLI uses a 1h TTL), cache reads at 0.1x, and output at 5x. Cache writes usually dominate, because the diff sits at the head of every prompt and is rewritten on each agentic turn.

Usage is reported on failures too, including claude_max_turns and claude_max_budget: an interrupted review still spent quota, and that is exactly when the number is needed. Compare total.total_cost_usd against the review's max_budget_usd to tell a review that was cut off by the spend cap from one that simply ran long.

Status meanings:

  • pending: a background operation is active;
  • open: a validated structured response is available;
  • waiting_for_quota: Claude rejected the operation for quota or rate limiting; stop polling and retry the same review at retry_at, or when capacity is available if retry_at is absent;
  • interrupted: the operation stopped and the explicit Claude session can be resumed;
  • failed: no resumable Claude session was captured;
  • closed: the review was explicitly finalized.

The model route is:

  • primary model: opus (Claude Opus 5);
  • fallback model: sonnet (Claude Sonnet 5).

--fallback-model only engages when the primary is overloaded or unavailable, so the fallback has to name a different model than the primary to do anything at all. Opus 5 matches the previous fable route on review work at half the price per token, which is why it is the default primary rather than the fallback.

Reviews are the most expensive thing this server does, and depth multiplies with frequency. Reserve reviews for changes where a second opinion changes the outcome, and select depth by the change's risk:

  • moderate-risk changes: effort: high;
  • security-sensitive, architectural, concurrent, persistent-data, authentication, payment, deployment, recovery, or otherwise high-risk changes: effort: xhigh;
  • uncertain classification: effort: high;
  • both profiles: timeout_seconds: 1200.

effort: max is available but rarely the right choice for review work: it yields diminishing returns over xhigh and is prone to overthinking. Use it only when correctness matters more than cost.

Pass effort explicitly. If it is omitted, the server's configured default is xhigh.

Turn budget versus spend cap

The two limits do different jobs and should not be tuned against each other.

max_turns limits Claude's internal agentic turns; it is not a duration and it is not a cost control. Every review that exhausts it returns no verdict at all, so the tokens it spent are a total loss — a turn budget set tightly to save quota is the most reliable way to waste it. Leave it generous (default 30); lower it only to stop a review that is going nowhere sooner.

When max_turns is omitted, the server sizes it from the computed diff rather than serving the configured ceiling to every review: a floor of 8 turns plus one turn per 8 KB of diff, capped at default_max_turns. The budget is stated in the prompt, so it anchors how widely the reviewer reads, and a trivial change given a large budget explores as though it were a large one. An explicit max_turns bypasses the scaling entirely.

max_budget_usd is a runaway guard, not an operating limit. Claude stops when it is reached and preserves its session, so the review continues rather than restarts — but it stops without emitting the structured verdict, so a cap set where reviews normally land converts a finished analysis into nothing. Set it well above the observed cost. It defaults to default_max_budget_usd (20), can be raised per review for a deliberately deep one, and is disabled entirely by setting the server default to 0.

On a subscription the dollar figure is notional: the CLI meters consumption in dollars whether or not any are billed. Read it as a consumption counter — at Opus rates, N dollars is roughly N/5 million input-equivalent tokens, where an input-equivalent token weights what actually depletes the quota (cache writes 2x, cache reads 0.1x, output 5x). The default of 20 is therefore about 4M units, against reviews that land near 1.4M per Claude invocation.

The cap applies per Claude invocation, not per review: a review that runs an initial pass and a continuation may spend up to twice it. Compare it against usage.last.total_cost_usd, never usage.total.total_cost_usd, which accumulates across invocations.

timeout_seconds limits the Claude subprocess. The asynchronous start call does not wait for that duration.

Copy-Paste AGENTS.md Policy

Place the following block in another project's AGENTS.md to make cross-review part of the Codex workflow:

## Cross-Review with Claude

Request a review when a change carries real risk: it touches security,
architecture, concurrency, persistent data, authentication, payments,
deployment, or recovery paths, or it spans enough surface that a regression
would not be obvious from the tests. Routine changes — localized edits,
documentation, test-only changes, mechanical refactors already covered by the
suite — do not need one. Review depth and review frequency both consume quota;
spend them where a second opinion changes the outcome.

For every change that meets that bar:

1. Implement the change.
2. Run the relevant tests, linting, and type checking.
3. Select the review depth according to risk. `effort` is the depth dial; do
   not use `max_turns` as one, because a review that runs out of turns returns
   no verdict and wastes everything it spent.
   - For a moderate-risk change, use `effort: high`.
   - For security-sensitive, architectural, concurrent, persistent-data,
     authentication, payment, deployment, recovery, or otherwise high-risk
     changes, use `effort: xhigh`.
   - When uncertain, use `effort: high`. Reserve `effort: max` for changes
     where correctness matters more than cost: on review work it yields
     diminishing returns over `xhigh` and is prone to overthinking.
4. Call `claude-reviewer.start_review` with the selected `effort`, literal
   `include_paths` for the files under review, and `timeout_seconds: 1200`.
   Leave `max_turns` and `max_budget_usd` at their server defaults; raise
   `max_budget_usd` only for a review you have deliberately scoped as deep.
   Omitting `max_turns` is what lets the server size the budget from the diff,
   so a small change is not handed the budget of a large one.
5. Provide a precise goal and the test results.
6. Poll `claude-reviewer.get_review_status` at the returned
   `poll_after_seconds` interval until the status is no longer `pending`.
7. If the status is `waiting_for_quota`, stop polling and report `retry_at` when
   present. At or after that time, or when capacity is available if it is
   absent, call `claude-reviewer.start_retry_review` with the same `review_id`
   and a message that restates the interrupted operation, then resume polling.
   Do not create a replacement review.
8. If `last_error_code` is `claude_max_turns` or `claude_max_budget`, Claude
   stopped at a budget before returning a verdict. The Claude session is
   preserved in both cases: call `claude-reviewer.start_continue_review` with
   the same `review_id` to obtain the verdict. Do not create a replacement
   review, and do not treat the review as failed. A repeated
   `claude_max_budget` on the same scope means the scope is too wide for the
   cap: narrow `include_paths` rather than raising `max_budget_usd` again.
9. Analyze each finding instead of accepting it blindly.
10. Fix confirmed critical-, high-, and medium-severity findings.
11. Prepare a factual technical response for incorrect findings.
12. If the verdict is `needs_context`, the review is incomplete rather than
    inconclusive: read `questions` for the files Claude could not examine, and
    narrow `include_paths` on a follow-up review of that scope. Do not record
    it as an accepted verdict.
13. Call `claude-reviewer.start_continue_review` with the same `review_id`,
   `refresh_diff: true`, and a request to verify the fixes.
14. Poll until the status is no longer `pending`, then require the returned
    `expected_response_sequence`; report a terminal error instead of polling
    indefinitely if that sequence was not produced.
15. Confirm that the continuation returns the same `claude_session_id`.
16. Stop after two completed review cycles unless a critical issue remains.
17. Call `claude-reviewer.close_review` after the final accepted verdict.
18. Do not treat Claude approval as a substitute for tests.
19. Claude is a read-only reviewer; Codex remains the only agent that modifies
    the repository.

Every review result and status poll carries a `usage` object. Use it instead of
guessing at budgets: when `usage.turn_budget_used_percent` approaches 100, the
next review of comparable scope will run out of turns and return nothing, so
narrow `include_paths` or raise `max_turns`. Compare `usage.last.total_cost_usd`
against `max_budget_usd` the same way — the cap applies per Claude invocation,
so `usage.total.total_cost_usd`, which accumulates across invocations, is not
the figure to compare. Report `usage.total.total_cost_usd` and
`usage.billed_input_equivalent` when the user asks what a review cost.

The project can add stricter language, test, or commit-attribution rules around this block. The filename recognized by Codex is AGENTS.md.

Tool Reference

Tool Behavior
start_review Validate, persist, and start a background review; return immediately
get_review_status Read the latest background status, structured response, or safe error
start_continue_review Resume the persisted explicit Claude session in a background worker
start_retry_review Retry a quota-blocked or interrupted retry with the same review ID
review_diff Synchronous compatibility form of the initial review
continue_review Synchronous compatibility form of the continuation
get_review Read persisted metadata without contacting Claude
list_reviews List review metadata with optional repository and status filters
close_review Mark a review closed or delete only its local association

review_diff and continue_review remain useful for small, bounded reviews. They are not recommended for maximum-effort reviews because an MCP client may stop waiting before Claude finishes. Progress notifications do not guarantee that a client-side deadline will be extended.

Session Persistence

Session metadata is stored at:

~/Library/Application Support/claude-reviewer/sessions.json

The native Claude conversation remains in Claude Code's own storage. This project stores the explicit association needed to resume it safely. Restarting Codex, the MCP server, or the Mac does not intentionally change that mapping.

A quota retry always recomputes the current scoped diff. If Claude provided no session ID, the server rebuilds the initial request from the persisted goal, resolved base commit, review focus, and path scope. Quota retries and ordinary continuations both use that resolved commit, keeping the original range stable if a symbolic reference such as HEAD or origin/main moves. Supply updated additional_context and test_results to start_retry_review when those ephemeral inputs are still relevant; prompt bodies are not persisted locally. When retrying a continuation, message is required so the interrupted verification request is restated explicitly.

An approve verdict leaves the review open so Codex can still request a correction-verification pass. close_review finalizes it explicitly. delete_claude_session: true deletes only the local association; it does not delete Claude Code's native conversation data.

Optional Configuration

Create ~/Library/Application Support/claude-reviewer/config.json:

{
  "claude_binary": "/opt/homebrew/bin/claude",
  "default_model": "opus",
  "default_fallback_model": "sonnet",
  "default_effort": "xhigh",
  "default_max_turns": 30,
  "default_max_budget_usd": 20,
  "timeout_seconds": 600,
  "async_timeout_seconds": 1200,
  "max_diff_bytes": 2097152,
  "max_output_bytes": 8388608,
  "log_level": "info",
  "session_retention_days": 30
}

Without an explicit Claude path, resolution tries PATH, then the Apple Silicon and Intel Homebrew locations. session_retention_days is reserved for future explicit cleanup; V1 does not automatically delete review records.

Errors

Tool errors are actionable JSON with code, message, and safe details. Claude invocation failures can include a correlation ID, failure stage, exit code, terminal reason, bounded redacted stderr, model, turn counts, and argument names without prompt values.

Common codes include:

  • request and Git errors: invalid_repository, invalid_base_ref, invalid_path_scope, empty_review_scope, diff_too_large;
  • lifecycle errors: review_not_found, review_closed, review_not_resumable, review_not_waiting_for_quota, review_waiting_for_quota, quota_retry_not_ready, review_busy, repository_mismatch;
  • Claude errors: claude_not_found, claude_not_authenticated, claude_timeout, claude_canceled, claude_max_turns, claude_max_budget, claude_quota_exceeded, claude_failed, claude_session_id_missing, invalid_claude_output, claude_output_too_large;
  • worker and storage errors: storage_error, worker_failed, background_worker_stopped, server_shutting_down;
  • content policy errors: sensitive_content_detected.

When a failure captured a Claude session ID, details identify the review as resumable. Continue it with the same review_id; never create a replacement conversation and pretend that context was preserved.

Quota rejection is different from an ordinary failure. Claude Code emits a structured rate_limit_event; the server records claude_quota_exceeded, waiting_for_quota, and retry_at plus retry_after_seconds when Claude reports a reset time. A bare HTTP 429 remains immediately retryable because no reset time is known. The server does not retry in a loop or depend on the STDIO process surviving until reset. Use start_retry_review after the reported time, or when capacity is available. force_before_retry_at: true is available only when quota became available early, such as after enabling additional usage.

V1 Limitations

  • macOS only;
  • one base Git reference compared with the current working tree; no independent from_ref and to_ref pair;
  • untracked file contents are not included automatically;
  • no HTTP server, graphical interface, GitHub App, PR comments, network database, multi-Mac synchronization, or telemetry;
  • no automatic split for diffs larger than the configured limit;
  • no automatic review-record retention cleanup;
  • closing a review does not delete the native Claude Code conversation;
  • quota retries are persistent but not automatically scheduled; Codex or another MCP client must call start_retry_review after retry_at.

About

Persistent, read-only Claude code reviews orchestrated by Codex through MCP.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages