Skip to content

Latest commit

 

History

374 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

#+title: sucoder

[[https://doi.org/10.5281/zenodo.21629611][https://img.shields.io/badge/DOI-10.5281%2Fzenodo.21629611-blue.svg]]

* Project Overview
Unix user and group permissions have been battle-tested for over fifty years.  When you bring on a new collaborator you give them their own account, set group read on shared files, and let the filesystem enforce the boundaries.  Why should an AI coding agent be any different?

=sucoder= treats an LLM agent (running through a harness such as Codex, Claude, Gemini, Aider, OpenCode, Goose, or Kimi) as a collaborator with its own unix account.  The human's canonical repository is group-readable but not group-writable; the agent works in a sandboxed mirror clone where it has full write access.  No custom container runtimes, no bespoke sandboxing daemons---just =chown=, =chmod=, and =git=.

* Quick Start


#+begin_src shell
git clone https://github.com/ligon/sucoder && cd sucoder
make quick-start
cd ~/Projects/my-project  # any git repo
sucoder collaborate --harness claude
#+end_src

No config file needed.  =sucoder= detects the git repo, mirrors it under =/var/tmp/coder-mirrors/=, and launches Claude (the default harness).  Use =--harness/-H= to pick a different harness (=codex=, =gemini=, =aider=, =opencode=, =goose=, or =kimi=), and =--model/-m= to choose its LLM independently.  The old =--agent/-a= spelling remains a compatibility alias for =--harness=.  Add =--task fix-login= to start on a dedicated branch.

For example, the same Aider harness can run models from different providers:

#+begin_src shell
sucoder collaborate --harness aider --model openai/gpt-5
sucoder collaborate --harness aider --model anthropic/claude-sonnet-4
sucoder collaborate --harness kimi --model openrouter/moonshotai/kimi-k3
#+end_src

Provider credentials can remain in the human user's =pass= store and be selected by the model's provider prefix; see [[*Provider credentials from =pass=][Provider credentials from =pass=]].  Harnesses must be on the configured =agent_user='s login =PATH=.  On a multi-user host, prefer one root-owned uv tool installation so human and agent accounts cannot drift independently; for example, when =uv= itself is globally available:

#+begin_src sh
sudo env UV_TOOL_DIR=/opt/uv-tools \
  UV_TOOL_BIN_DIR=/usr/local/bin \
  UV_PYTHON_INSTALL_DIR=/opt/uv-python \
  uv tool install --python python3.12 aider-chat

# OpenCode, Codex, and Kimi share the system Node installation.
sudo npm install -g opencode-ai @openai/codex @moonshot-ai/kimi-code

# Goose is a native binary; its official installer accepts a shared target.
curl -fsSL https://github.com/aaif-goose/goose/releases/download/stable/download_cli.sh \
  -o /tmp/goose-download-cli.sh
sudo env GOOSE_BIN_DIR=/opt/goose/bin CONFIGURE=false \
  bash /tmp/goose-download-cli.sh
sudo ln -s /opt/goose/bin/goose /usr/local/bin/goose
#+end_src

The separate shared Python directory matters when uv downloads a managed interpreter: otherwise a root-run install can leave the tool pointing into root's private home directory.

** Provider credentials from =pass=

Keep API keys in the human user's password store and put only entry names in
=~/.sucoder/config.yaml=.  The first line of each entry is the secret:

#+begin_src yaml
credentials:
  openrouter:
    pass: openrouter.ai/apikey
#+end_src

OpenRouter's protocol, endpoint, and conventional environment variable are
built in.  A model such as =openrouter/moonshotai/kimi-k3= therefore selects
both the provider credential and the upstream model.  Sucoder runs =pass show=
as =human_user= at launch; the =coder= account receives neither password-store
access nor the human's GPG keys.

For Aider and other environment-aware harnesses, Sucoder supplies
=OPENROUTER_API_KEY=.  For native Kimi it uses Kimi's temporary
=KIMI_MODEL_*= provider channel, so the key is not copied into
=~coder/.kimi-code/config.toml=.  Environment values are staged over stdin in
a random agent-owned mode-0600 file, then sourced and unlinked immediately
before the harness starts.  They do not appear in the launch command or
Sucoder logs.  An agent with shell access can still inspect its own process
environment; keeping the raw provider key beyond the agent's reach requires a
separate rate-limited model gateway.

Custom providers can be defined once:

#+begin_src yaml
credentials:
  laboratory:
    pass: research/lab-gateway
providers:
  laboratory:
    credential: laboratory
    protocol: openai
    base_url: https://models.example.edu/v1
    env_var: LABORATORY_API_KEY
#+end_src

Literal =agent_launcher.env= and =--agent-env= remain supported and use the
same private-file transport, but command-line values may remain in shell
history.  Prefer named =pass= credentials for secrets.

For multi-repo setups, skills, and system prompts, see [[*Configuration][Configuration]] below.  =make env-setup= does the full host provisioning (=make help= for all targets).

* Requirements
- Linux (user/group sandboxing relies on =useradd=, =groupadd=, and setgid)
- Python >= 3.9 with pip
- git
- git-crypt (for decrypting =.mcp.json= which contains MCP server tokens)
- Node.js and npx (MCP servers are distributed as npm packages)
- sudo
- make (optional; you can run the setup commands by hand)
- bash
- At least one supported harness CLI on the agent user's PATH: =claude=, =codex=, =gemini=, =aider=, =opencode=, =goose=, or =kimi=

* Why Org-mode?
This README is =.org=, the default system prompt is =.org=, and the skills repo leans heavily on Org-mode.  Why not Markdown?

Org-mode is a plain-text format that every LLM reads fluently—but it also has structured features (TODO states, properties, tags, executable source blocks) that Markdown lacks.  For a project about human-agent collaboration, those features pull their weight: a system prompt can carry metadata an agent can act on, a skill file can embed runnable examples, and a handoff note can be a living checklist rather than a static document.

You do not need Emacs to read or edit these files.  Any text editor works; GitHub renders =.org= natively.  But if you /do/ use Emacs, the integration is a bonus, not a requirement.

* What =sucoder collaborate= does

Running =sucoder collaborate --harness claude= (or with =--task=) triggers a multi-step workflow.  Steps marked *auto* happen without intervention; steps marked *human* need you at the keyboard.

** 1. Resolve configuration /auto/
If =~/.sucoder/config.yaml= exists, load it.  Otherwise, derive everything from the environment: =$USER= becomes the human identity, the git repo root becomes the canonical repo, and =/var/tmp/coder-mirrors/= is the mirror root.  The harness CLI is resolved from =--harness= (or legacy =--agent=) > =$SUCODER_AGENT= > =~/.sucoder/agent= > PATH auto-detect > interactive prompt.  =--model= is then applied independently through the selected harness profile.

** 2. Prepare the canonical repository /auto/
Set group-read permissions on your repo so the =coder= user can traverse it.  Adds a =coder= git remote pointing at the mirror (for easy fetching later) and writes a helper script =scripts/fetch-agent-branches.sh=.  Your repo stays read-only to the agent—only the group bits and the remote config change.

** 3. Clone (or verify) the mirror /auto/
If the mirror does not exist under =/var/tmp/coder-mirrors/<repo>/=, clone it from the canonical repo as the =coder= user.  Push access back to canonical is disabled (=no_push=).  If the mirror already exists, verify the remote and enforce permissions.

** 4. Sync or create a task branch /auto/
- Without =--task=: fetch the latest commits from canonical into the mirror.
- With =--task fix-login=: create a branch =coder/fix-login-<timestamp>= in the mirror, based on the default base branch (or =--base=).

** 5. Compose the agent launch command /auto/
Detect the harness type (Claude, Codex, Gemini, Aider, OpenCode, Goose, Kimi) from the command name and inject the appropriate flags: model, write-access (=--dangerously-skip-permissions=, =--sandbox danger-full-access=, =--yolo=, =--yes-always=, =--auto=), writable directories, skills paths, and system prompt.  Aider receives the composed context through a private read-only file so it remains interactive.  Kimi receives a private custom-agent file which includes its native base prompt before the SuCoder context; other harnesses use their native prompt flag or trailing text.

** 6. Launch the agent /human/
The agent process starts inside the mirror working tree.  From here, *you are talking to the agent*.  It has full write access to the mirror but cannot write to your canonical repo.  The agent commits its work to the task branch.

When the agent session ends (you quit it, or it finishes), =sucoder=
auto-commits any changes the agent made to =~coder/.claude/skills/=
(see [[*Agent Skills Tracking][Agent Skills Tracking]]), then control returns to your shell.

** 7. Review and integrate /human/
Fetch the agent's branch into your canonical repo and review it:
#+begin_src shell
git fetch /var/tmp/coder-mirrors/my-project coder/fix-login-20250215180000
git checkout -b review/fix-login FETCH_HEAD
# ... review, test, iterate ...
git checkout main && git merge --ff-only review/fix-login
#+end_src
Or pull directly:
#+begin_src shell
git pull --ff-only /var/tmp/coder-mirrors/my-project coder/fix-login-20250215180000
#+end_src
The canonical repository remains unwritable by the agent throughout.

** Running the steps individually
=collaborate= bundles steps 2–6.  You can also run them separately:
| Command              | Step                                         |
|----------------------+----------------------------------------------|
| =prepare-canonical=  | Set permissions and add agent remote (step 2)|
| =agents-clone=       | Clone the mirror (step 3)                    |
| =push=               | Send canonical commits to the mirror (step 4)|
| =start-task=         | Create a task branch (step 4)                |
| =agents-run=         | Launch the agent (steps 5--6)                |
| =worktrees=          | List active worktrees with status             |
| =attach=             | Reconnect to a remote tmux session            |
| =audit=              | Run compliance review of skills and/or code  |
| =list=               | Discover harnesses, models, mirrors, or skills|

* Key Features
- agents-clone :: Clone a canonical repository into an agent-owned mirror with strict permissions.
- push :: Send the canonical repository's commits to the mirror without granting write access.  Pulls the agent's commits back first, and refuses rather than force-push over a mirror it could not read.  Available as =sync= too.
- start-task :: Create a new agent task branch based on a human branch while preserving branch naming conventions.
- status :: Summarize the state of a mirror, including permissions and git remotes.
- collaborate :: Prepare canonical, ensure mirror, and launch agent in one step.
- worktrees :: List active git worktrees in a mirror with per-worktree status (branch, commits ahead, dirty state).  Supports =--watch= for live monitoring and =--diff= for file-level changes.
- attach :: Reconnect to a remote agent session via tmux after an SSH disconnect.
- nodes :: Show compute-node availability for a SLURM partition (read-only =sinfo= query over the warm gateway tunnel; see [[*Inspecting node availability][Inspecting node availability]]).
- tunnel :: Keep the free SSH hops to a target warm (=up= / =status= / =doctor= / =down=), and forward a compute-node service port to localhost (=forward= / =forwards=) so a web app served on a node (Jupyter, claude-science, ...) is reachable in a local browser without new authentication.
- audit :: Run a compliance review of agent-written skills and/or mirror code changes (see [[*Auditing][Auditing]]).
- list :: Discover built-in and configured harnesses, their capabilities, configured model defaults, harness model catalogs, mirrors, and skills.

** Discover harnesses and models

The =list= command groups discovery by resource:

#+begin_src shell
sucoder list harnesses
sucoder list models
sucoder list models kimi
sucoder list models --provider openrouter kimi
sucoder list models --harness aider gpt
sucoder list models --harness opencode kimi
sucoder list models --harness kimi k3
sucoder list providers
sucoder list mirrors
sucoder list skills
#+end_src

=list harnesses= shows every built-in profile plus custom harness commands found in configuration, where each is configured, its default model, and native shell, file-tool, Agent Skills, MCP, subagent, provider, and approval capabilities.  =suggest= means the harness can print a command but does not expose a model-driven shell tool; =?= means SuCoder has no built-in metadata for a custom harness.  =list models= queries the sole credentialed provider automatically; with several credentialed providers, select one using =--provider=.  The resulting names include the provider prefix and can be passed directly to =--model=.  If no provider is configured, the command falls back to configured model defaults.  =--harness aider=, =--harness opencode=, and =--harness kimi= instead query the selected harness under the configured agent user.  Kimi reports configured aliases.  =list providers= shows endpoints and credential references without invoking =pass= or displaying secrets.  The optional positional argument filters the selected model catalog.

The older =mirrors-list= and =skills-list= spellings remain supported for compatibility.
- --target/-T :: Select a named remote execution target (e.g., =--target savio=) to run the agent on a remote host.
- --version/-V :: Print the installed sucoder version (CalVer, derived from git tags).
- --dry-run :: Available on modifying commands to review logged operations without executing.

* Usage
#+begin_src shell
# Clone the canonical repository into the configured mirror.
sucoder agents-clone project

# Send canonical's commits to the mirror (inverse of `pull`).
sucoder push project

# Create a task branch for the agent starting from ligon/main.
sucoder start-task project rewrite-parser --base main

# Inspect current branch and outstanding changes.
sucoder status project
#+end_src

** Moving commits between canonical and the mirror
- push :: canonical --> mirror.  For remote mirrors the mirror's branches
  are overwritten to match canonical.  For local mirrors canonical's
  commits become *visible* in the mirror as =<remote>/<branch>=; the
  mirror's own branches are left where the agent put them.
- pull :: mirror --> canonical.  Brings the agent's commits back, with a
  prompt when the two histories have diverged.
- sync :: an alias for =push=, kept for compatibility.

=agents-clone= creates the mirror. Remote startup initializes absent or
empty directories and recovers valid repositories in place. It never deletes
a directory because a probe failed or a repository has no default branch.
Unreadable, invalid, symlinked, or non-repository nonempty paths stop startup
with a diagnostic; inspect them before retrying.

** Mirror safety: the pull must succeed before the push
=push= (and =agents-clone=) send to the mirror with =git push --all
--force=, so they first pull whatever the agent committed there.  If that
pull cannot read the mirror --- host unreachable, allocation gone, the
agent's work sitting on a branch other than the configured base --- the
push is refused rather than allowed to overwrite unretrieved commits.

An empty or half-initialised mirror is *not* treated as unreadable, so
first-time bootstrap is unaffected: sucoder asks the mirror whether it
holds any commits instead of guessing from the error text.
Any ref (including feature branches, tags, and WIP refs) counts as content,
even with an unborn HEAD. A repository initialized by this invocation skips
the initial fetch and receives a non-forcing first push. The
=--allow-unverified-mirror= override does not bypass initialization checks.

- --allow-unverified-mirror :: Push anyway, discarding any unpulled
  mirror commits.  Use when the mirror is known to be expendable.

Likewise =pull= exits non-zero when it could not read the mirror,
instead of reporting =Pull complete.= after pulling nothing.

* Configuration
Skip this section if zero-config mode works for you.

Configuration is stored in YAML (default =~/.sucoder/config.yaml=).  To set it up:

1. Copy the sample configuration and adjust paths, users, and mirror names.
   #+begin_src shell
   install -d -m 750 ~/.sucoder
   cp config.example.yaml ~/.sucoder/config.yaml
   #+end_src
2. Update =~/.sucoder/config.yaml= so that:
   - =human_user= matches your login (for example =ligon=).
   - =agent_user= and =agent_group= match the agent Unix account (default =coder=).  This is a shared account used by any supported agent (Codex, Claude Code, Gemini CLI, etc.)---it is not tied to a specific agent binary.
   - =mirror_root= points to the directory where agent mirrors will live.
   - =mirrors.<name>.canonical_repo= points at the human-owned repository.
3. Ensure the agent can read (but not write) the configuration directory and file.
   #+begin_src shell
   sudo chgrp coder ~/.sucoder
   sudo chmod 750 ~/.sucoder
   sudo chgrp coder ~/.sucoder/config.yaml
   sudo chmod 640 ~/.sucoder/config.yaml
   #+end_src
4. (Optional) Install the skills repository for agent capabilities.
   #+begin_src shell
   git clone https://github.com/ligon/sucoder-skills ~/Projects/sucoder-skills
   ln -s ~/Projects/sucoder-skills ~/.sucoder/skills
   sudo chgrp -R coder ~/Projects/sucoder-skills
   sudo chmod -R g+r,g-w ~/Projects/sucoder-skills
   #+end_src

   *Note*: The skills repository is separate from the tool repository for security isolation.
   Skills influence agent behavior and are kept in a dedicated repository to prevent agents
   from modifying skills and tool code in the same commit. The tool enforces semantic
   versioning compatibility between the tool and skills repository.
6. Warm up the shared skills catalog so launches confirm access before work starts.
   #+begin_src shell
   ls ~/.sucoder/skills
   sucoder list skills
   sucoder list mirrors
   #+end_src
   The listings should succeed without permission errors. The CLI helper highlights any unreadable
   paths and =sucoder list mirrors= prints the configured mirrors (with their canonical and mirror directories).
   To preload the catalog into a session, use the agent's file-read command (e.g., =codex read ~/.sucoder/skills/SKILLS.md= for Codex).

** Shell Completion
Enable tab completion for =sucoder= commands and mirror names by running:
#+begin_src shell
sucoder --install-completion
#+end_src

** Sudo Requirement
By default, =sucoder= uses =sudo= to run commands as the agent user (=coder=).
The human user must have sudo access to impersonate the agent user.  If you are
running as the agent user directly (for example inside a container), pass
=--no-agent-sudo= to skip sudo:
#+begin_src shell
sucoder --no-agent-sudo agents-clone project
#+end_src

** Command Overview
- =collaborate= is the all-in-one command: it runs =prepare-canonical= + =agents-clone= (ensure mirror) + =agents-run= (launch agent) in a single step.
- =agents-run= only launches the agent, assuming the mirror already exists.  Use this when the mirror is already prepared and you just want to start a session.

** Provision the Mirror
1. Prepare the canonical repository for collaboration.
   #+begin_src shell
   sucoder prepare-canonical project
   #+end_src
   *Note*: this command modifies the canonical repository — it sets the group to =coder=,
   grants group read permissions, removes group write bits, and optionally adds an agent
   remote and fetch helper script.  Elevate with =--sudo= when root privileges are required.
   By default it also adds a =coder= remote pointing at the agent mirror and writes
   =scripts/fetch-agent-branches.sh= to simplify fetching agent branches (=--no-agent-remote= disables this).
2. Clone the mirror as the human operator; the tool will invoke git as the agent.
   #+begin_src shell
   sucoder agents-clone project --verbose
   #+end_src
   Mirror arguments support shell completion, so pressing Tab after =sucoder <command>= suggests configured mirror names.

** Sync Before Handing Off
1. Refresh the mirror with the latest human commits.
   #+begin_src shell
   sucoder sync project
   #+end_src
2. Create a task branch for the agent based on the desired human branch.
   #+begin_src shell
   sucoder start-task project add-metrics --base main
   #+end_src
   The command prints the full branch name (for example =coder/add-metrics-20251107164028=).

** Agent Workflow
- The agent works directly inside the mirror directory (for example =/var/tmp/coder-mirrors/project=) on the branch created above.
- The agent commits work to =coder/<task>-<timestamp>= without pushing to the canonical repository.
- To run the full workflow in one step (prepare canonical, ensure mirror, launch agent), use:
  #+begin_src shell
  sucoder collaborate project --task add-metrics
  #+end_src
#  Adjust flags such as =--no-agent-remote=, =--sudo=, or extra arguments just as you would with the individual commands.

- To select a harness and model independently, or experiment with an arbitrary binary, use =--harness=/=--model= or =--agent-command= (and optional =--agent-env= for non-secret overrides):
  #+begin_src shell
  sucoder agents-run project --harness aider --model openrouter/deepseek/deepseek-chat
  sucoder agents-run project --agent-command "foo --flag" --agent-env DEBUG=1
  #+end_src

- The CLI automatically injects the appropriate write-access flag for the detected agent (e.g., =--yolo= for Codex, =--dangerously-skip-permissions= for Claude Code; see the flag template table below).  When invoking an agent manually, include the equivalent flag to avoid read-only failures.  The helper will also warn if it detects a read-only sandbox and still attempts a launch.

- Org authoring playground (documentation example): load =docs/examples/org-authoring-demo.org= using your agent's file-read command.
  Demonstrates the in-repo Org skills—structure, markup, tables, capture, citations,
  exporting, agenda integration—so humans and agents have a quick reference.

*** Codex remote control

Use =--agent-command= for a command with subcommands; =--agent= and =--harness= select a single executable.
For example, on a configured Savio target:

#+begin_src shell
sucoder -T savio collaborate project --agent-command "codex remote-control"
#+end_src

=--agent-command "codex remote-control start"= is also accepted.
SuCoder converts =start= to the foreground form and reports that choice, so tmux and Slurm retain ownership of the service process.
Commands such as =stop= and =pair= are management operations; run them separately on the remote host.
The remote Codex installation must support =remote-control= and have the authentication and pairing required by the connected Codex client.
See the [[https://learn.chatgpt.com/docs/developer-commands?surface=cli#codex-remote-control][Codex remote-control reference]].

The service starts in the mirror's working directory, including the existing local-disk clone and Slurm confinement when configured.
The full SuCoder prelude (system prompt, target instructions, and skills catalog) is installed in =AGENTS.override.md= in the service's actual working directory.
This keeps the defaults available when a connected client replaces Codex's =developer_instructions= config value; SuCoder also supplies that config value for clients that retain it.
SuCoder preserves an existing local override's text, or copies =AGENTS.md= into a separate generated region when creating the override.
It refreshes generated regions on each service launch and adds the override to Git's local exclude file.
The original =AGENTS.md= stays unchanged; its copy refreshes at launch, and the agent is directed to read the current file before working.
The launcher needs =python3= on the agent host and allows the prelude size plus 1 MiB for native instruction files; explicit =project_doc_max_bytes= overrides still take precedence.
An edited or ambiguous generated block, a concurrent edit, or a symlinked or Git-tracked override stops launch with an error and preserves the file.
The file remains available to later Codex chats in that mirror.
Choose the mirror's working directory in the connected client: native instructions are scoped to that directory and its children.
After updating SuCoder, restart the service and open a new chat to load these instructions; reattaching to an existing service does not refresh it.
=--model= and =agent_launcher.model= become a =model= config override.
The usual =needs_yolo= intent becomes =sandbox_mode="danger-full-access"= and =approval_policy="never"= config defaults; set =needs_yolo: false= to omit them, or supply explicit =-c= overrides to restrict them.
SuCoder does not apply TUI flag templates to this subcommand; =default_flags= must contain options supported by =remote-control=.

To make this the mirror's default launch:

#+begin_src yaml
agent_launcher:
  command: [codex, remote-control]
#+end_src

=sucoder sessions= labels the pane as a remote-control service, including when the local launcher record is missing.
=sucoder message= skips it even with =--force=; send instructions through the connected Codex client.
=sucoder peek= displays the service log.
=sucoder renew= retains the service command and model across allocation turnover, and regenerates the SuCoder prelude.
Its checkpoint request remains a file sentinel; renewal restarts the server and does not automatically checkpoint or resume active chats.

** Parallel Work with Worktrees
Inside the mirror the agent can use git worktrees to work on multiple
tasks in parallel.  Claude supports this natively via =claude --worktree NAME=;
other agents can achieve the same result with =git worktree add=.

Each worktree gets its own checkout while sharing the same git object store
(no duplicate clone).  This is useful for:

- Parallel feature development :: Fix a bug in one worktree while
  building a feature in another.
- Comparative computation :: Run the same estimation with different
  parameters in separate worktrees, then consolidate results.
- Review isolation :: Keep exploratory changes out of the main
  branch until they are ready.

The human can monitor all worktrees from upstream:

#+begin_src shell
sucoder worktrees project                   # snapshot
sucoder worktrees project --watch 10        # live, refresh every 10s
sucoder worktrees project --diff            # include file-level changes
sucoder worktrees project --main            # include the main worktree
#+end_src

Example output:

#+begin_example
Worktrees for mirror 'MyModel' (/var/tmp/coder-mirrors/MyModel):

  .claude/worktrees/instruments-A
    Branch: worktree-instruments-A @ a1b2c3d
    Ahead:  2 commits (vs main)
    Last:   a1b2c3d GMM estimation with instrument set A (1 minute ago)
    Status: clean

  .claude/worktrees/instruments-B
    Branch: worktree-instruments-B @ e4f5g6h
    Ahead:  2 commits (vs main)
    Last:   e4f5g6h GMM estimation with instrument set B (30 seconds ago)
    Status: dirty (1 modified)
#+end_example

The =worktrees= command uses =git worktree list --porcelain= for discovery,
so it works regardless of which agent created the worktrees or where they
are located.

When the agent consolidates worktree branches back to the mirror's main
branch, prefer =git merge --no-ff= to preserve the individual branch histories.
This keeps the worktree commits individually cherry-pickable from upstream.

*** Shared virtual environment for worktrees

Poetry caches virtual environments by project path, so each worktree
would normally create its own =.venv=.  To share a single environment
across all worktrees, pin =VIRTUAL_ENV= in =.envrc=:

#+begin_example
_git_common="$(git rev-parse --git-common-dir 2>/dev/null)"
if [ "${_git_common}" = ".git" ]; then
  _main_tree="$(pwd)"
else
  _main_tree="$(dirname "${_git_common}")"
fi
VIRTUAL_ENV="${_main_tree}/.venv"
export VIRTUAL_ENV
layout poetry
#+end_example

** Remote Execution
Mirrors can also live on a remote host (e.g., an HPC cluster) where the
local machine cannot run heavyweight computation.  The privilege
separation shifts from "different Unix users on the same box" to "same
user on different machines"---the network boundary enforces isolation.

*** Ordinary Linux host over SSH

A single Linux box needs only a =host= target. SuCoder runs on the launcher;
install Git, Bash, tmux, and the chosen agent harness on the remote host,
with their executables available to non-interactive SSH commands. The agent
runs as the SSH user, without a remote sudo or separate coder account.

#+begin_src yaml
targets:
  workstation:
    host: workstation.example.org  # or an existing ~/.ssh/config alias
    remote_user: ligon             # optional; otherwise SSH chooses the user
    mirror_root: ~/mirrors
    # ssh_options:
    #   Port: 2222
    #   IdentityFile: ~/.ssh/workstation
#+end_src

From the local repository:

#+begin_src sh
sucoder -T workstation collaborate
sucoder -T workstation attach
sucoder -T workstation pull
sucoder -T workstation tunnel status
#+end_src

Execution and Git transport use the configured host directly, including after
connection expiry. Normal SSH config, authentication, and host-key checking
apply. =remote_user= and =ssh_options= are used for both connection startup
and subsequent commands. An explicit =remote_user= overrides any =User=
option, regardless of capitalization. SuCoder evaluates =ssh -G= before
choosing a shared connection, so edits to SSH aliases or included config files
do not reuse a connection to the previous account or host. Prefer an SSH alias for settings that should also
apply to standalone =ssh= and =git= commands.

=host= cannot be combined with =gateway=, =transfer_host=, =slurm=, or the
BRC-specific =cert_file= option. Standard SSH certificates can instead use
=IdentityFile= and =CertificateFile= in SSH config or =ssh_options=.
No scheduler allocation, deadline watchdog, or WIP snapshotter is started.
The mirror uses persistent storage on that host.

=tunnel up= warms the one connection without writing cluster aliases to
=~/.ssh/config=. =tunnel forward 8888= forwards a service on that host;
=tunnel down= closes the SSH connection. =release= is a Slurm-only operation;
end the remote tmux session when finished with an ordinary host. =--node= is
rejected by =collaborate=, =attach=, and =tunnel forward=: a direct target has
one host and no scheduler to request a node from.

*** Cluster configuration

Define a named target in =~/.sucoder/config.yaml=:

#+begin_src yaml
targets:
  savio:
    gateway: brc.berkeley.edu
    transfer_host: dtn.brc.berkeley.edu
    mirror_root: ~/mirrors
    remote_user: ligon        # optional; remote SSH username (defaults to local $USER)
    control_persist: 7d          # warm-tunnel idle lifetime (optional)
    keepalive_interval: 30       # ServerAliveInterval, seconds (optional)
    keepalive_count_max: 120     # ServerAliveCountMax (optional)
    cert_file: ~/.ssh/ssh_certs/brc_cert  # gateway SSH cert (optional)
    x11: false                   # disable the default X11 forwarding (optional)
#+end_src

No =mirrors:= entry is needed---zero-config repo detection works with
targets.  The target can be used with any repo.

The three SSH connection-sharing knobs are optional and default to the
values shown above; tune them per target:

- =control_persist= --- how long a warm =ControlMaster= lingers after the
  last client disconnects (ssh time format: =7d=, =12h=, =90m=...).  This
  is the "tunnel stays open" lever: a longer value means =collaborate=,
  plain =ssh=, and Emacs TRAMP reuse the socket --- no PIN/OTP --- for
  longer.  The cluster's server-side idle policy is the real ceiling, so
  treat a long value as best-effort, not a guarantee.
- =keepalive_interval= / =keepalive_count_max= --- the =ServerAlive*=
  pair.  Their product is the grace budget before a stalled connection is
  torn down (=30 x 120 = 1h= by default), which is what lets a tunnel ride
  out a brief network blip or a laptop nap *on the same network*.  Keep the
  interval short (the probes also keep NAT/firewall mappings warm); raise
  =count_max= for a longer budget.  Neither can revive a connection the
  server already reaped or one whose IP changed --- those still re-auth.
- =cert_file= --- path to a local SSH *certificate* private key, presented on
  the gateway hop so =tunnel up= (and a plain =ssh <target>-gw= / TRAMP)
  authenticate with *no* PIN/OTP for the cert's lifetime.  Mint one with
  =sucoder -T <target> cert=, which prompts for your BRC PIN + one-time code,
  POSTs them to the MSM CA, and writes the cert here (the CA caps a cert at
  12h).  You rarely need to run it explicitly: when the cert is missing or
  expired, an *interactive* =sucoder= command offers to mint a fresh one
  before it connects (so one OTP buys a 12h window instead of an OTP per
  connection); a non-interactive/agent run just falls back to ssh's own
  prompt.  The cert is applied only to the gateway --- the login/DTN hops keep
  their publickey auth through the mux.  =sucoder -T <target> tunnel doctor=
  reports the cert's validity.  (=scripts/brc-cert.sh= still works as a
  standalone minter if you prefer.)
- =x11= --- trusted X11 forwarding (=ssh -Y= semantics) on the
  *interactive* hops (=collaborate= launch, =attach=), so programs started
  inside the remote session can open X windows on your local display.
  *On by default* whenever the local session has a =DISPLAY= (XQuartz on
  macOS); silently skipped when it doesn't, so headless and cron runs stay
  quiet.  Disable with =x11: false= on the target or =sucoder --no-x11=
  per invocation.  The remote node needs =xauth= and =X11Forwarding yes=
  in its =sshd=.  On a confined (sbatch) target, attach reaches the login
  node over ssh and steps into the job with =srun --x11=, which needs the
  cluster's Slurm X11 support --- so that srun flag is added only when x11
  was *explicitly* enabled (=x11: true= or =--x11=), never by the default,
  keeping =attach= safe on clusters without it.
  Caveat: the forward lives and dies with the attaching ssh connection.
  The tmux session created at launch inherits a working =DISPLAY=; after a
  detach/re-attach, *new* tmux windows pick up the fresh =DISPLAY=
  automatically (tmux's default =update-environment= includes it), but
  processes already running keep the stale one --- refresh a shell with
  =eval "$(tmux show-environment -s DISPLAY)"=.

*** Using a target from Emacs / TRAMP

=sucoder -T <target> tunnel up= writes three =Host= aliases into your
=~/.ssh/config= --- =<target>-gw= (gateway), =<target>-ln= (login node),
and =<target>-dtn= (DTN) --- each carrying =HostName=, =ProxyJump=, =User=,
and a =ControlPath= that points at sucoder's warm socket.  TRAMP can ride
those directly, so opening a remote file re-uses the warm =ControlMaster=
with no fresh PIN/OTP (and with a minted =cert_file=, the first hop is
OTP-free too).

One Emacs setting is required, because TRAMP otherwise injects its *own*
=ControlMaster= options that override the alias's =ControlPath= and open a
separate, un-warmed connection:

#+begin_src emacs-lisp
(setq tramp-use-ssh-controlmaster-options nil)   ; defer to ~/.ssh/config
#+end_src

Then warm the tunnel and open files against the =-ln= alias (your NFS
=$HOME= / =~/mirrors= live on the login node, shared with every compute
node):

#+begin_src shell
sucoder -T savio-node tunnel up                  # warms sockets + writes aliases
ssh savio-node-ln hostname                        # sanity check: silent == TRAMP will be too
#+end_src

#+begin_example
C-x C-f /ssh:savio-node-ln:~/mirrors/MyProject/
#+end_example

Notes:
- No alias is written for a *compute* node; multi-hop through the login
  node if you need one: =/ssh:savio-node-ln|ssh:n0032.savio2:~/...=.  But
  since =$HOME=/=~/mirrors= is shared NFS, editing via =-ln= already reaches
  the files a compute job sees.
- If the login-node pin drifted (the alias's =HostName= no longer matches
  the warm node), re-run =tunnel up= before pointing TRAMP at =-ln=.

*** Usage

Use =--target= (=-T=) as a global option before the subcommand:

#+begin_src shell
sucoder -T savio collaborate               # launch agent on cluster
sucoder -T savio worktrees --watch 30      # monitor from local
sucoder -T savio attach                    # reconnect via tmux
sucoder -T savio tunnel forward 8888       # compute-node web app -> http://localhost:8888/
sucoder -T savio tunnel forwards           # list forwards + master liveness
sucoder -T savio tunnel forward 8888 --cancel
#+end_src

=tunnel forward= defaults the node to the one your collaborate session
is on (override with =--node n0030.savio4=) and terminates the forward
*on the node*, so services bound to =127.0.0.1= work.  Open the app's
printed URL with the host replaced by =localhost=; keep the port and any
=?token=...= query intact (use =--local-port= only if the local port is
already taken).

=scripts/claude-science.sh= is a worked end-to-end example of the whole
pattern.  It gives the web app its *own* dedicated thin SLURM slice (a
job named =csci-<target>=, params read from the target's =slurm:= block)
so it never has to share --- and fight --- a compute node with heavy
jobs.  =up= warms the tunnels, reuses-or-allocates that slice (held open
in a detached login-node =tmux=), starts the app on it, forwards its
port, and opens the URL locally; =url= re-fetches a fresh URL from the
*running* app, repairs the forward, and reopens it; =down= releases the
slice (cancels the job --- queued or running --- kills its tmux, removes
the forward).  =up= / =url= are idempotent, and that includes the queue:
a slice still PENDING when the scheduling ceiling (=CS_SCHED_WAIT=,
default 300s) expires *stays queued*, and a later =up= attaches to it
instead of resubmitting, so the job keeps its accrued priority.  A
submission SLURM rejects outright is reported with the =srun= client's
captured stderr rather than a blind timeout.  The same discipline covers
the app itself: a =serve= process that exists but has not yet answered
(e.g. a slow cold start over NFS) is waited on, never double-started ---
two servers would mean two writers on one SQLite DB --- and the timeout
diagnostics say whether the process is still alive (raise =WAIT_SECS=)
or died (with a foreground repro command).  See the header comment for
the =TARGET= / =NODE= / =APP_BIN= / =OPENER= / =WAIT_SECS= and =CS_*=
slice overrides.

=scripts/savio-run= generalizes the *dispatch-and-collect* half of that
pattern for one-off compute: from your laptop it warms the tunnels and
runs a command on a compute node, streaming the output back.  SLURM
resources default to the target's =slurm:= block (flags override).

#+begin_src shell
savio-run -- python3 -c 'print(2**10)'     # srun: blocks, streams stdout back
savio-run -c 8 --mem 32G -- ./crunch.sh    # override resources
id=$(savio-run --batch -t 2:00:00 -- ./long.sh)   # sbatch: prints JOBID, returns
savio-run --fetch "$id"                    # state + stdout/stderr when done
savio-run --status                         # squeue/sacct for your jobs
#+end_src

Batch jobs write to =~/.savio-run/<jobid>.{out,err}= on the target; a lone
quoted argument is run as a shell string (pipes/=&&= work), multiple args
are treated as argv.

*** What happens under the hood

1. *Authenticate once* --- =sucoder= establishes an SSH ControlMaster
   connection to the gateway.  For clusters with OTP/two-factor auth,
   this is the only password prompt; all subsequent SSH commands
   multiplex through the socket.
2. *Pin a login node* --- SSH to the gateway, run =hostname=, store
   the result (e.g., =ln002.brc=) in =~/.sucoder/sessions/<mirror>.yaml=.
   Subsequent connections jump to the same node via =-J=.  For SLURM
   targets --- where the login node is only a routing hop, not where the
   agent lives --- a later command first reconciles this pin with the node
   =tunnel up= keeps warm, and re-pins to a healthy node off the gateway
   if the recorded one is wedged, so a stale pin can't strand you on a dead
   login node while a live tunnel sits idle.
3. *Open a tunnel* --- A local port forward to the data transfer node
   for fast git transport.
4. *Push canonical state* --- =git push= through the tunnel to the
   remote mirror.
5. *Launch the agent* --- SSH to the pinned login node, =cd= into the
   remote mirror, start the agent inside a =tmux= session.
6. *Add a git remote* --- A remote named after the target (e.g.,
   =savio=) is added to the canonical repo so the human can
   =git fetch savio= to pull back agent work.

*** One job per mirror per target

A confined launch never allocates a second job for a mirror that already
has a live one.  It checks twice, because the two checks fail in
different ways:

1. the session record, =~/.sucoder/sessions/<mirror>--<target>.yaml=,
   which holds the job id;
2. failing that, =squeue --me --name=sucoder-<token>= --- the scheduler.

The second exists because the first holds /one/ id and reads as blank
when the file is missing, unreadable, or was written under a different
=-T= spelling.  Blank used to mean "no job", and the launch submitted a
second =sbatch= over a live one.  Slurm has known the answer all along:
every launch is submitted =--job-name=sucoder-<token>=.

When the scheduler finds a job the record lost, =collaborate= adopts it
--- attaches, and writes the id back so =attach=, =release= and =renew=
can reach it again --- rather than allocating beside it.

A job for the same mirror on a /different/ target does not block the
launch: one mirror on two targets is a deliberate configuration.  It is
reported and the new job is allocated.  Targets are told apart by
partition/account/qos, the same signature =sucoder sessions= groups by,
so a target that pins none of the three claims nothing.

A failed =squeue= does not block a launch either.  This check is a safety
net over the record, not a gate; the record-keyed check is the one that
already refuses to resubmit on an unknown answer.

=salloc= targets carry the same =--job-name=.  Nothing reuses by name on
that path yet --- it reuses via the record and adopts by node --- but
without the name an allocation is invisible to =sucoder sessions=, which
filters on the =sucoder-= prefix.

*** Listing your sessions

=sucoder sessions= lists every SuCoder job across all configured
clusters, with its state, time left, node, and what the tmux session is
actually running.  It needs no =-T=: it reports on everything.

#+begin_src shell
sucoder sessions                    # probe each job's tmux pane (a few seconds)
sucoder sessions --fast             # skip the probe; job ids and clocks only
sucoder sessions --no-login-nodes   # skip the login-node sweep only
#+end_src

#+begin_example
carleton-htc   savio4_htc / co_carleton / carleton_htc4_normal
  LSMS_Library  39002769  RUNNING  11-00:13  n0043.savio4  claude  6m
  SuCoder       39025067  RUNNING  11-22:24  n0029.savio4  bash    2h   ! agent exited

savio-node   savio3 / fc_jevons
  K-Aggregators 38991234  RUNNING   3-04:11  n0142.savio3  claude  4m   ! no session record

stale session records (the job is gone; the record is not):
  CFEDemands--savio-node           -> 34739256   sucoder -T savio-node release CFEDemands
  MetricsMiscellany--carleton-htc  -> 38999103   mirror not configured; edit the record by hand
#+end_example

Targets sharing a gateway share a =$HOME= and a scheduler, so they are
queried once between them and sorted out afterwards by
partition/account/qos.  That one query runs on a login node the session
records already pin when there is one, falling back to the gateway ---
both answer =squeue= identically, and the login node is the cheaper and
less contended hop (see [[*What a remote command costs][What a remote command costs]]).  When the same host
also has to be swept for tmux sessions, both questions travel in one
script, because that master carries one session at a time.  The listing
is enumerated from =squeue=, not from the session records under
=~/.sucoder/sessions/=: those hold one
job id per (mirror, target), so a record-driven listing would show
exactly the jobs that are /not/ the problem.  Two flags matter:

- =! no session record= :: nothing local points at this job, so
  =attach=, =release= and =renew= --- which all resolve through that
  record --- cannot reach it.  Only =scancel= can.  It happens when a
  record is overwritten by a later launch, is unreadable, or was written
  under a different =-T= spelling.
- =! agent exited= :: the allocation is alive but its tmux pane is a
  bare shell.  A confined job's window ends in =exec bash -l= so it
  survives a clean =/exit=, and the batch body's keeper polls only
  whether the session exists --- so the job goes on holding its slice for
  the rest of its =--time= with nobody home.  Nothing else reports this.

The trailing sweep is the reverse direction: session records naming a
job the scheduler has forgotten.  Each row carries the invocation that
clears it, because =release= resolves the mirror through
=config.mirrors= and the target through =-T=, and so cannot reach a
record whose mirror has since been dropped from the configuration --- a
common case on a long-lived launcher host.  For those,
=scripts/clear-stale-sessions.py= does in bulk what =release= does to
one record: it nulls =slurm_job_id= and =compute_node=, leaves every
other key (=login_node= included) alone, and never deletes a file.  It
is a dry run unless given =--apply=, which first writes a timestamped
tar backup beside the session directory, and it aborts rather than
touch anything if =squeue= cannot be reached on every configured
cluster --- "the query failed" must never be read as "the job is gone".

Targets with no =slurm:= block hold no allocation, so =squeue= says
nothing about them --- but =collaborate= still launches a tmux session on
the *login node*, with the same =exec bash -l= tail.  A job's allocation
ends at its =--time= and takes the tmux server with it; a login node has
no walltime, so an abandoned session there persists until the node
reboots, on shared infrastructure.  Those are listed too:

#+begin_example
savio   no scheduler (login-node sessions)
  sucoder-SuCoder       ln002.brc  bash    ! agent exited
  sucoder-LSMS_Library  ln001.brc  claude
hhsurveys   no scheduler; no login-node sessions
#+end_example

Enumerated from =tmux list-sessions= filtered on the =sucoder-= prefix ---
the login-node analogue of filtering =squeue --me= on =sucoder-<token>=,
so a session whose local record was lost is still found.  The gateway
round-robins across login nodes, and which one an account reaches depends
on its class (condo and FCA accounts land on different nodes), so the
probe asks each node the session records pin *by name* as well as the
gateway itself; reconnecting to the gateway alone could never see them
all.  =--fast= skips this too, and says =not inspected= rather than
implying absence; =--no-login-nodes= skips *only* this sweep, keeping the
per-job pane probe.

The nodes are asked concurrently, so the wall clock is the slowest node
rather than the sum --- and the sweep is started before the =squeue=
queries rather than after them, since a login-node session has no
allocation and so nothing in it depends on the scheduler.  On a cold
start (no live ControlMasters --- after minting a certificate, say) the
gateway is authenticated first and alone, because it is the only hop that
can prompt, and because several connections racing to authenticate the
same gateway earns =Too many authentication failures= from =sshd=; the
nodes behind it are then brought up together.

The =! agent exited= check is why the probe exists, and why it is on by
default: it reads the pane's child process, not
=#{pane_current_command}=, which reports the pane's /shell/ and would
call every working agent dead.  =--fast= skips it and says so.
Read-only throughout --- nothing submits, cancels or writes.

*** Listing and retiring WIP snapshots

=sucoder snapshots= lists every WIP snapshot ref on every configured
cluster's mirrors, with what the scheduler's accounting says about the
job that took it, and whether it is past retention.  Like =sessions= it
needs no =-T= and is read-only; =--retire= deletes what the listing marks.

#+begin_src shell
sucoder snapshots                       # list; nothing is deleted
sucoder snapshots --retire              # delete the ones past retention
sucoder snapshots --recovery-window 3   # days after a job's end; default 7
#+end_src

#+begin_example
hpc.brc.berkeley.edu   (carleton-htc, savio-htc, savio-node)
  LSMS_Library   wip-job/LSMS_Library/39123514  0h   39123514  running           kept: job still going
  LSMS_Library   wip/LSMS_Library               3d   39002769  cancelled 3d ago  kept: within the 7d window
  SuCoder        wip-job/SuCoder/38710868       41d  38710868  timeout 40d ago   retire: job ended 40d ago, past the 7d window

3 snapshot(s); 1 past retention.  --retire deletes them.
#+end_example

The prepare script already retires ended jobs' snapshots, but only when
a job is launched against that mirror, so a mirror no longer in use kept
its refs forever --- and each ref pins a snapshot tree that =gc= can then
never collect.  This is the sweep that needs no launch and no
allocation.  Mirrors are found on the cluster's filesystem, not in the
config, so a mirror dropped from the configuration (the one =release=
can no longer reach) is swept too.

A snapshot is judged by /when its job ended/, from =sacct=, never by the
age of the ref alone.  The snapshotter leaves the ref untouched while the
tree is unchanged, and a confined window survives the agent's exit, so a
ref can sit frozen for days under a job that is still alive; any
age-only threshold shorter than the allocation would delete a live job's
only snapshot.  Where accounting cannot place a job (aged out, or no
=sacct=), ref age is used instead against the cluster's longest
=slurm.time= plus the window, which no live job can exceed; where a
target has no finite =--time= even that has no safe bound, and nothing
is deleted.  Every deletion is guarded by the hash the listing saw, so a
snapshot rewritten in between is left alone.

*** Messaging a running session

Two agents working one mirror --- one in a job, one on the laptop ---
have no channel to each other; on 2026-09-21 the only link was a human
copying between terminals, and a wrong diagnosis was overturned only
because of it.  =sucoder message= types a line into a running agent's
session from wherever you are.

#+begin_src shell
sucoder message LSMS_Library "ln001's Lustre client is wedged; read from dtn, stop repairing"
sucoder message --all "ln001 is wedged; stop repairing"     # every live session, every cluster
sucoder message LSMS_Library --dry-run "..."                # who would get it, and the exact line
sucoder -T carleton-htc message LSMS_Library "..."          # one target, when a mirror runs on two
#+end_src

For a terminal agent stdin is the API and tmux owns it, so the line is
=tmux send-keys= into the session's pane --- through =srun --overlap=
for a job, where the tmux server lives inside the allocation, exactly
as the pane probe reaches it.  Recipients are the sessions =sucoder
sessions= shows, enumerated from the scheduler and the login nodes, so
a job whose record was lost can still be reached.

The text arrives framed, =[message from you@laptop via sucoder, <time>]
...=, because it lands with the same authority as you typing; the
system prompt's bulletin convention tells the agent to treat it as a
peer's report to check, not an instruction from the human.  A session
whose pane is a bare shell --- an agent that has exited, leaving the
=exec bash -l= behind --- is never sent to, since Enter there runs the
text as a command; one whose pane could not be probed is skipped unless
=--force=.  The message is one line: a newline sent literally is an
Enter, and would submit half of it.

A reply cannot be pushed back --- a cluster has no route to a laptop
behind NAT --- so it is pulled.  The default system prompt asks an agent
to answer on its own screen with a line =REPLY <id>: ...=, naming the id
the message carried, and =--wait= reads the pane after the send for that
line and prints it:

#+begin_src shell
sucoder -T carleton-htc message SuCoder --wait 120 "is n0032's Lustre healthy from where you sit?"
sucoder -T carleton-htc peek SuCoder -n 60      # the pane's last rows, any time
#+end_src

=peek= is the read half on its own, for looking at a session without
sending anything; unlike =message= it will happily show a pane that is
a bare shell.  Whether an agent acts on a message is still up to it.
For a Codex remote-control service, =peek= shows server logs and =message= refuses terminal delivery, including with =--force=.
Read and steer its chats through the connected Codex client.
The cheaper half of the same fix is the bulletin,
=~/.sucoder/bulletin.org= on the host the mirror lives on, which the
default system prompt tells agents to read before concluding that
infrastructure is broken and to write to before attempting a repair.

*** Choosing a login node

The login node is pinned once, from a round-robin =ssh <gateway> hostname=,
and stored in the session record; every later command reuses it.  When that
node goes bad --- BRC's login nodes can lose a Lustre route while the node
itself stays up and answers ssh --- =--login-node= steers off it:

#+begin_src shell
sucoder --login-node ln002.brc -T savio collaborate
#+end_src

It overrides the record's pin and writes the new one back, so subsequent
commands agree.  The sibling of =--node=, which selects a *compute* node.

For a SLURM target the login node is only a routing hop (the work lives on
=compute_node= + job id), so repointing is free.  Without a scheduler the
agent's tmux lives *on* the login node, so changing the pin abandons any
session on the old one --- =attach= will look at the new node and not find
it.  sucoder warns when that is what you are doing; =sucoder sessions= still
lists the session on the old node.

*** Inspecting node availability

Before reserving a node (or targeting one with =--node= /
=--local-disk=), =sucoder nodes= shows which nodes in a partition are
free, reusing the target's warm gateway ControlMaster so it costs no
OTP when a tunnel is already up:

#+begin_src shell
sucoder -T savio-node nodes            # partition defaults to slurm.partition (savio3)
sucoder -T savio-htc  nodes            # -> savio4_htc
sucoder -T savio-node nodes savio3_gpu # positional arg overrides the partition
#+end_src

It prints one row per node --- state, CPUs (Allocated/Idle/Other/Total)
and load --- followed by the drained/down nodes and their reasons
(=sinfo -R=).  The query is read-only: no allocation, no session
changes.

Caveat: =sinfo= reports SLURM state, not Lustre health.  A node can
read =idle= while its filesystem mount is wedged, so weigh the drain
reasons and any anomalous load on an otherwise-idle node --- the query
cannot promise a node's filesystem is healthy.

*** Node pinning

On SLURM clusters, =sucoder= allocates a compute node for the agent
session.  The node is stored in =~/.sucoder/sessions/<mirror>.yaml= so
that subsequent commands (=attach=, =pull=, =worktrees=) reconnect to
the same node.

To request a specific node (e.g., to recover work on local disk from a
previous session):

#+begin_src shell
sucoder -T savio collaborate --node n0047.savio3
sucoder -T savio attach --node n0047.savio3
#+end_src

The =--node= option sets a preferred node for the SLURM allocation.
If that node is unavailable, SLURM may fall back to another node.

**** Sharing a reserved node

An exclusive partition hands out the whole node, but a single agent
session often leaves it underused.  If you already hold a live
allocation on a node, point a *second* mirror at it with =--node= and
=sucoder= adopts the existing job instead of reserving another:

#+begin_src shell
sucoder -T savio collaborate Foo                      # reserves n0020.savio3
sucoder -T savio collaborate Bar --node n0020.savio3  # shares it
#+end_src

Each mirror gets its own tmux session (=sucoder-Foo=, =sucoder-Bar=)
and its own clone, so the two agents run side by side on one node.
Adoption only ever attaches to *your own* RUNNING jobs on that node.

Because the sessions share one SLURM job, =sucoder release= is
job-aware: releasing a mirror while siblings still hold the job
*detaches* it (kills that mirror's tmux session) but keeps the
allocation alive.  Only the last holder's =release= runs =scancel=.

*** Local disk

By default the agent works directly in the remote mirror on the shared
filesystem.  On clusters whose compute nodes have fast local scratch,
=slurm.local_disk= (or =--local-disk= on the command line, with
=--local-disk-root PATH= to pick a root other than =/local=) moves the
agent's /working tree and caches/ there while the shared mirror stays
the durable repository:

#+begin_src shell
sucoder --local-disk -T carleton-htc collaborate SuCoder
sucoder --local-disk-root /scratch/local -T carleton-htc collaborate SuCoder   # implies --local-disk
#+end_src

#+begin_src yaml
targets:
  carleton-htc:
    gateway: hpc.brc.berkeley.edu
    transfer_host: dtn.brc.berkeley.edu
    mirror_root: ~/mirrors
    slurm:
      partition: savio4_htc
      account: co_carleton
      confined: true
      local_disk: true        # or a path; true means /local
#+end_src

This is /tiering/ (design and measurements in
[[file:docs/local-disk-tiering.org][docs/local-disk-tiering.org]]), and it works the same way
on =confined= (=sbatch=) and =salloc= targets:

- the launch clones =~/mirrors/<name>= to
  =/local/job<ID>/mirrors/<name>= and starts the agent there, with
  =UV_CACHE_DIR=, =PIP_CACHE_DIR=, =npm_config_cache=, and =TMPDIR=
  under =/local/job<ID>/=;
- a =post-commit= hook publishes every commit to the shared mirror the
  moment it exists, so =sucoder pull= and the login-node view see it
  immediately; the hook never forces, and a push refused because the
  laptop pushed first is reported for the agent to =git pull
  --ff-only=;
- the deadline watchdog snapshots uncommitted work to
  =refs/sucoder/wip-job/<name>/<jobid>= on the shared mirror every
  =wip_snapshot_minutes= and restores it on the next launch when the
  branch tip has not moved and the job that took it has ended; a clean
  =sucoder release= retires the job's own snapshot (=--keep-wip= keeps
  it), and the next launch retires those of other ended jobs;
- SLURM wipes =/local/job<ID>/= when the job ends; nothing is orphaned
  and =--node= pinning is unnecessary.

The shared mirror becomes a mailbox: do not edit its working tree by
hand (an uncommitted tracked change there makes =updateInstead= refuse
every publish from the clone; the prepare step warns loudly if it
finds one).  Ignored files (=.venv=, =node_modules=) are rebuilt each
job; committed work is durable instantly, dirty work to within
=wip_snapshot_minutes=.  The agent is told all of this in a
=WORKSPACE (local-disk tiering)= block of its prelude, including the
command that shows the last snapshot.

Earlier releases put the /whole/ mirror at =/local/mirrors= on one
node.  That layout is retired: a session that still records it gets a
warning naming the node, so any commits made there can be pushed by
hand before =sucoder pull= is trusted.

*** Shared partitions (fractional allocations)

Exclusive partitions (e.g. =savio3=) hand out an entire compute node.
Shared/HTC partitions (e.g. =savio4_htc=) expect each job to declare how
much of the node it wants; without =cpus_per_task= and =mem= you get the
partition's minimal default (typically 1 core / a few hundred MB).

#+begin_src yaml
targets:
  savio-htc:
    gateway: hpc.brc.berkeley.edu
    transfer_host: dtn.brc.berkeley.edu
    mirror_root: ~/mirrors
    control_persist: 24h
    slurm:
      partition: savio4_htc
      account: fc_jevons
      qos: savio_normal
      cpus_per_task: 4   # cores
      mem: 16G           # per-job memory
      time: "24:00:00"
#+end_src

=cpus_per_task= must be a positive integer; =mem= is passed verbatim to
=salloc --mem=, so use SLURM's own units (=16G=, =4000M=, etc.).  Both
fields are optional --- omitting them keeps the partition default, which
is the right behaviour for whole-node partitions.

*** Deadline watchdog and WIP snapshots

Every SLURM-backed session (=salloc= or =confined= =sbatch=) starts a
small watchdog on the compute node.  It warns at 30, 15, and 5 minutes
before the allocation's =--time= via =tmux display-message= and by
writing a warning under
=/tmp/sucoder-<uid>/timers/<scope>/<node>-<job>/= (the
un-suffixed =slurm-deadline.warn= and per-mirror
=slurm-deadline-<mirror>.warn= are also written in =$HOME/.cache/sucoder/= for
older prompts; these compatibility copies show the last writer).
It never cancels the job; that stays with
=sucoder release=.

The scope is a digest of the mirror and target names. The timer uses Linux
=/proc= and =flock= to verify/reuse a live watchdog for that allocation;
repeated startup does not restart it or reset warning state. Startup reports
whether it started or reused a timer. Failures explicitly report that
deadline warnings and periodic snapshots are unavailable; diagnostics remain
in =timer.log= beside its =owner= and =status= files. Timers from the older
implementation are not killed by a broad process-name match; they exit with
their old allocation.

Locks, owner/readiness records, threshold markers, and timer logs live in a
private, owner-verified mode-700 directory on node-local =/tmp=. This transient
state only needs to survive for the watchdog lifetime. =TMPDIR= is deliberately
not used: it may point at shared storage. Staged scripts and compatibility
warnings stay on NFS; NFS need not support =flock=. Working clones may use
=/local/job<jobid>/=, with snapshots pushed to durable storage. Lustre is not
used for the timer's frequent small-file operations.

The same script snapshots the mirror's dirty working tree (tracked and
untracked files, not ignored ones) to
=refs/sucoder/wip-job/<mirror>/<jobid>= on the mirror's =origin= at each
warning and every =wip_snapshot_minutes= (default 10; =0= disables the
periodic run).  A mirror with no =origin= is never snapshotted, so on a
shared-filesystem mirror this is a no-op today; it becomes live with the
local-disk tiering described in [[file:docs/local-disk-tiering.org][docs/local-disk-tiering.org]].

The ref carries the job id because one mirror can be running in two jobs
at once, on two nodes, with two different uncommitted trees.  A relaunch
restores this job's own snapshot if it has one, otherwise the newest whose
job has ended -- never one belonging to a job still running, which would
silently import another node's work.  Snapshotted repositories should
=.gitignore= their generated outputs: every snapshot that changes the tree
writes a new commit to the shared mirror, and the superseded one becomes
unreachable there until it is collected.

#+begin_src yaml
    slurm:
      partition: savio4_htc
      account: co_carleton
      confined: true
      wip_snapshot_minutes: 10
#+end_src

*** Reconnecting

If the SSH connection drops, the tmux session on the login node
survives.  Reconnect with:

#+begin_src shell
sucoder -T savio attach
#+end_src

If the ControlMaster socket expires (default 12 hours), =sucoder=
detects the stale socket and re-prompts for authentication.

**** Orphaned sessions (SLURM)

If you lose the SSH connection but the tmux server is still running on
the compute node --- and =sucoder attach= fails (e.g. the session JSON
is missing, the compute-node hostname wasn't recorded, or the site
blocks direct SSH login -> compute) --- attach via the allocation itself:

#+begin_src shell
sucoder -T savio attach --via-srun
#+end_src

This stops at the login node and joins the running job with =srun
--jobid=<JOB> --overlap --pty=, which lands you on the compute node
*inside the job's cgroup*.  =--overlap= is the key flag: it lets you
add a step to an existing allocation without trying to consume extra
resources.

If =sucoder= itself isn't available (different machine, no install),
the manual recipe is:

#+begin_src shell
ssh <gateway>
ssh ln001                                  # or whichever login node
squeue -u $USER                            # find your JOBID
srun --jobid=<JOBID> --overlap --pty bash -l
tmux attach -t sucoder-<mirror>            # or just: tmux a
#+end_src

*** Persistent sessions (auto-renew)

On a condo QOS with no wall-clock cap (e.g. =carleton_htc4_normal= on
=savio4_htc=), a session can outlive any single allocation.  =sucoder
renew= watches the current job and, as it nears its courtesy =--time=
(or if it vanishes at a maintenance reboot), re-allocates a fresh node
and relaunches the agent in a *detached* tmux --- so the presence
survives turnover hands-off.

#+begin_src shell
sucoder -T carleton-htc collaborate          # start the session
sucoder -T carleton-htc renew                # another terminal: keep it alive
#+end_src

The loop is conservative: a transient probe failure (SSH blip) never
triggers a relaunch --- only a *successful* probe reporting the job
gone/terminal does.  Tunables: =--drain-minutes= (lead time before
=--time=, default 20), =--poll-interval= (seconds, default 60),
=--checkpoint-grace= (seconds to let the agent commit/push after the
drain nudge).  =--once= runs a single probe/act cycle, for cron-driven
renewal.

Before each turnover the loop writes
=$HOME/.cache/sucoder/renew-requested= on the compute node; the agent
(per its =system_prompt_extra=) commits, pushes, and writes a handoff
note, then rehydrates from that note after relaunch.  Stop renewal with
Ctrl-C (the allocation is left running) or =sucoder release= (frees it).

For Codex remote control, renewal preserves the original service command and model, including =--agent-command= and =--model= overrides.
The prelude and credential environment are rebuilt for the replacement allocation.
The checkpoint sentinel is not a chat message: checkpoint active work through the connected Codex client before turnover.
Renewal restarts the server; it does not automatically resume its chats.

See [[file:docs/persistent-presence.org][docs/persistent-presence.org]]
for the design and the open cgroup-confinement caveat.

*** SSH robustness

=sucoder= handles several common failure modes transparently:

- *Stale ControlMaster sockets* are detected and recycled
  automatically.
- *DTN timeouts* --- if the data transfer node is unreachable,
  =sucoder= falls back to pushing via the compute node.
- *Debug mode changes* --- switching =--debug-ssh= on or off
  automatically cycles the ControlMaster socket so the new setting
  takes effect.
- *Busy masters* --- a =Session open refused by peer= is a /busy/
  signal, not a dead connection (see below); it is waited out and
  retried, never re-authenticated.

*** What a remote command costs

Two site facts shape every remote command, and both were measured on
=hpc.brc.berkeley.edu= (2026-09-20):

1. *A session open costs ~8s*, over an already-warm ControlMaster,
   before the remote command runs at all --- ~4.5s on a login node, and
   the same for an =sftp= subsystem open, so it is the session setup
   itself and not your login shell.  Multiplexing saves the
   /authentication/, not the time.
2. *A ControlMaster carries one session channel at a time*
   (=MaxSessions 1=).  A second simultaneous command on the same host is
   not queued but *refused*, and on a refusal =ssh= falls back to dialling
   the host directly --- a fresh authentication, which on this gateway can
   earn =Too many authentication failures=.  Port-forward channels are
   exempt, which is why reaching several hosts /through/ one gateway
   concurrently is fine while running two commands /on/ it is not.

So: connections are verified once per run rather than once per command;
hosts are asked one question at a time, with the answers to two
questions fused into one script where they share a host; a refusal is
waited out rather than reconnected; and a command that can run on a
login node does, because that hop is cheaper and is not the one every
jumped connection already contends for.

Because the waiting is unavoidable, it is at least visible: every remote
round trip announces itself on *stderr* before it is made, naming the
host and what it is about to run.

#+begin_example
· ln001.brc  squeue --me --noheader -o '%i|%j|%P|%a|%q|%T|%L|%N'
· hpc.brc.berkeley.edu  tmux list-sessions -F "#{session_name}" 2>/dev/null ...
#+end_example

Durations go to the log (=~/.sucoder/logs/=), and =-v= puts them on the
console along with the connection machinery's own reasoning --- which
probe timed out, why a master was declared dead, whether a refusal was
read as busy.

** Human: Review and Integrate Agent Work

*** Local mirrors
Fetch the agent branch back into the canonical repository:
#+begin_src shell
git fetch coder                              # uses the coder remote added by prepare-canonical
git log --oneline coder/coder/add-metrics-20251107164028 -5
git checkout -b review/add-metrics FETCH_HEAD
# ... review, test, iterate ...
git checkout main && git merge --ff-only review/add-metrics
#+end_src

Or pull directly:
#+begin_src shell
git pull --ff-only /var/tmp/coder-mirrors/project coder/add-metrics-20251107164028
#+end_src

*** Remote mirrors
When using =--target=, =sucoder= automatically adds a git remote
named after the target.  Fetch agent work back over SSH:
#+begin_src shell
git fetch savio                              # fetches from brc.berkeley.edu:~/mirrors/project
git log --oneline savio/master -5
git merge --ff-only savio/master
#+end_src

The canonical repository remains unwritable by the agent throughout
both local and remote flows.

* Configuration
Configuration is stored in YAML (default path is =~/.sucoder/config.yaml= for the human operator).  Grant the agent group read access to this file while keeping it group-writable disabled.

#+begin_src yaml
human_user: ligon
agent_user: coder
agent_group: coder
mirror_root: /var/tmp/coder-mirrors
log_dir: ~/.sucoder/logs
skills:
  - ~/.sucoder/skills  # optional global default for all mirrors (overridable per-mirror)
system_prompt: ~/.sucoder/system_prompt.org

# Named remote execution targets (optional).
# Use sucoder -T <name> to select one.
targets:
  savio:
    gateway: brc.berkeley.edu
    transfer_host: dtn.brc.berkeley.edu
    mirror_root: ~/mirrors
    control_persist: 7d

# Mirrors are optional when using zero-config repo detection.
mirrors:
  project:
    canonical_repo: ~/src/project.git
    mirror_name: project
    # Optional; omit to auto-detect the canonical repo's default branch
    # (origin/HEAD, then a local main, then master, then current HEAD).
    # default_base_branch: main
    task_branch_prefix: task
    branch_prefixes:
      human: ligon
      agent: coder
    # Optional: override or extend global skills for this mirror
    skills:
      - ~/.sucoder/skills/org-style
      - ~/.sucoder/skills/document-skill
    agent_launcher:
      # Harness executable and optional default model can be configured independently.
      command: [aider]
      model: openrouter/deepseek/deepseek-chat
      # Alternatives: [opencode], [goose, run, --interactive], or [kimi].
      # Remote Codex service: [codex, remote-control] (start is normalized to foreground).
      # Optional: map generic intents to agent-specific flags
      needs_yolo: true
      writable_dirs:
        - "~"
      flags:
        skills: "--skills {path}"  # Uses mirror skills or ~/.sucoder/skills by default
#+end_src
You can also maintain a global catalog at =~/.sucoder/skills/SKILLS.md=, which can
point to additional skill files or directories using =file:= links or bullet-listed paths. Any
catalog discovered this way is loaded automatically before launching the agent.

The =agent_launcher.flags= mapping lets you translate generic intents (write access, writable
directories, workdir, default flags, skills paths, model, context file) into harness-specific switches. If a skills flag
template is provided, sucoder will pass any configured skills plus the default
=~/.sucoder/skills= directory when it exists.

** Agent flag templates by CLI
The table below documents the current default flag templates used when the agent type is detected.
Templates can be overridden per-mirror or globally via =agent_launcher.flags=, with precedence:
per-mirror > global > profile > UNKNOWN.

| Harness  | Model            | Write access                                         | Writable directory            | Prompt content       | Prompt file   | MCP config          |
|----------+------------------+------------------------------------------------------+-------------------------------+----------------------+---------------+---------------------|
| Claude   | --model {model}  | --dangerously-skip-permissions                       | --add-dir {path}              | --system-prompt      | (none)        | --mcp-config {path} |
| Codex    | --model {model}  | --sandbox danger-full-access --ask-for-approval never| (none)                        | trailing text        | (none)        | (none)              |
| Gemini   | --model {model}  | --yolo                                               | --include-directories {path}  | --prompt-interactive | (none)        | (none)              |
| Aider    | --model {model}  | --yes-always                                         | (none)                        | (none)               | --read {path} | (none)              |
| OpenCode | --model {model}  | --auto                                               | (none)                        | --prompt             | (none)        | native config       |
| Goose    | --model {model}  | native policy                                        | (none)                        | --text               | (none)        | native config       |
| Kimi     | --model {model}  | --auto                                               | --add-dir {path}              | (none)               | --agent-file {path} | native config    |

The common =default_flag= template is ={flag}=.  Native Agent Skills discovery
means OpenCode, Goose, and Kimi do not need a =skills= flag.  Goose launches as
=goose run --interactive= so the injected =--text= context is processed before
the interactive session begins.  Kimi's generated custom-agent file includes
=${base_prompt}= so its native tools and skills remain available.

If you change any defaults, update =AGENT_PROFILES= in =sucoder/config.py= alongside the
README so the documentation stays in sync.

** MCP servers (repo-specific tools)

MCP (Model Context Protocol) servers give agents access to external services---GitHub, web pages, databases---via a standardised tool interface.  Claude Code on the web injects GitHub tools automatically, but the CLI (=claude=) does not.  Since =sucoder= launches agents via the CLI, MCP servers bridge the gap.

*** How it works

There are two independent sources of MCP configuration; Claude merges them automatically:

1. *Repo-shipped =.mcp.json=* --- lives in the repository root.  Claude discovers it natively when launched in the mirror.  This is the simplest approach: add the file, commit it, and every agent session gets those tools.

2. *Sucoder-config =mcp_servers=* --- defined in =~/.sucoder/config.yaml= at the global or per-mirror level.  Sucoder generates a =.sucoder-mcp.json= file in the mirror and passes it via =--mcp-config=.  Use this for infrastructure tools that are not part of any single repo.

*** Shipped MCP servers

This repository includes a =.mcp.json= with three servers:

| Server   | Purpose                                              |
|----------+------------------------------------------------------|
| =github= | Read/write GitHub issues, PRs, CI status, comments   |
| =fetch=  | Fetch web pages and URLs (documentation, references)  |
| =memory= | Persistent knowledge graph across agent sessions      |

The =github= server requires a =GITHUB_TOKEN= environment variable.  Because =.mcp.json= contains this secret, the file is encrypted with =git-crypt=.

*** Setting up git-crypt (first time)

After cloning the repository, unlock the encrypted files:

#+begin_src shell
# Install git-crypt if needed
brew install git-crypt   # macOS
sudo apt install git-crypt  # Debian/Ubuntu

# If you are initialising git-crypt for the first time in this repo:
git-crypt init
git-crypt add-gpg-user <your-gpg-key-id>

# Edit .mcp.json to add your real GITHUB_TOKEN, then commit.
# The file will be stored encrypted in the repo but readable in your checkout.

# If git-crypt is already initialised, just unlock:
git-crypt unlock
#+end_src

Collaborators who have been added via =git-crypt add-gpg-user= can unlock with their GPG key.  The agent user (=coder=) works in a mirror cloned from an already-unlocked checkout, so the plaintext =.mcp.json= is available without additional setup.

*** Configuring MCP servers via =config.yaml=

For MCP servers that should be available across multiple repos (not shipped in any single repo), add them to the sucoder configuration:

#+begin_src yaml
# Global MCP servers (available to all mirrors)
mcp_servers:
  my-database:
    command: npx
    args: ["-y", "@modelcontextprotocol/server-postgres"]
    env:
      DATABASE_URL: "postgresql://localhost/mydb"

# Or per-mirror (overrides global for that mirror)
mirrors:
  project:
    canonical_repo: ~/src/project.git
    mcp_servers:
      custom-tool:
        command: my-mcp-server
        args: ["--port", "8080"]
#+end_src

Sucoder writes these to =.sucoder-mcp.json= in the mirror (excluded from git) and passes =--mcp-config= to Claude.  If the repo also ships a =.mcp.json=, Claude merges servers from both sources.

** Agent capability comparison

The core =sucoder= workflow (mirror, sync, permissions, branch management) is
fully agent-agnostic.  The table below summarises where agents differ in
features that =sucoder= can exploit.

| Capability                   | Claude                          | Codex                          | Gemini                  | Aider                 | OpenCode       | Goose                     | Kimi                   |
|------------------------------+---------------------------------+--------------------------------+-------------------------+-----------------------+----------------+---------------------------+------------------------|
| Native =--worktree= flag     | Yes (=claude --worktree NAME=) | No                             | No                      | No                    | No             | No                        | No                     |
| Subagent worktree isolation  | Yes (=isolation: worktree=)    | No                             | No                      | No                    | No             | No                        | No                     |
| Session resume               | Yes (=claude --resume=)        | Yes                            | Yes                     | Chat history support  | Yes            | Yes (=--resume=)          | Yes (=--session=)      |
| System prompt mechanism      | =--system-prompt=               | Trailing text                  | =--prompt-interactive=  | Private =--read= file | =--prompt=     | =run --interactive --text= | Private =--agent-file=|
| Write-access flag            | =--dangerously-skip-permissions=| =--sandbox danger-full-access= | =--yolo=                | =--yes-always=        | =--auto=       | Native policy             | =--auto=               |
| TTY requirement              | Works with subprocess           | Works with subprocess          | Requires =exec= (TTY)   | Works with subprocess | Subprocess     | Subprocess                | Subprocess             |
| Default harness              | *Yes*                           | Supported                      | Supported               | Supported             | Supported      | Supported                 | Supported              |

Agents without native =--worktree= support can still use git worktrees
manually.  The agent just needs to run =git worktree add= inside the mirror;
=sucoder worktrees= will discover and display them regardless of which agent
created them.

Similarly, the remote execution support (=--target=, SSH tunnels,
ControlMaster) is entirely agent-agnostic---any agent CLI can be launched
over SSH.

When adding a new harness, implement its profile in =AGENT_PROFILES= in
=sucoder/config.py= and add a row to these tables.

** Which agent binary actually launches
=sucoder= hands the agent name to =execvp= (or to =sudo -u <agent_user>=), so the
binary is chosen by a PATH lookup---but the PATH that decides is the *agent user's
login* PATH, not yours.  Both launch paths resolve inside a shell that has sourced
the agent's login profile, and a stock Debian =~/.profile= prepends
=$HOME/.local/bin=.  What such a shell does /not/ pick up is =~/.bashrc=, which
returns early when non-interactive, so version managers wired up there---nvm and
friends---never load.  A shell you type =codex= into and a shell =sucoder= launches
can therefore resolve to different installs, and the stale one wins silently.

Before each local launch, =sucoder= asks that same shell (=bash -lc 'command -v
<agent>'=, through the same =sudo= escalation the launch uses) and logs the path
and =--version= it gets back:

#+begin_example
INFO Agent binary: codex -> /usr/local/bin/codex (codex-cli 0.147.0)
#+end_example

Resolving through the launch's own shell rather than =sucoder='s PATH is what makes
the report trustworthy: it accounts for the agent user's login profile, and for
=sudo='s =secure_path= when sudo is in play.

If it finds a *newer* build of the same name elsewhere---anywhere on PATH, or under
the agent user's =~/.local/bin= or =~/.nvm/versions/node/*/bin=---it warns that the
resolved one is shadowing it.  The usual fix is a symlink from a directory that
comes earlier on the agent user's login PATH, or an =agent_launcher.nvm= block to
pin resolution outright.

The whole check is best-effort and contained: probes are bounded, and any failure
inside it is logged at debug level and swallowed rather than allowed to abort a
launch that would otherwise succeed.

The check is skipped for remote launches (PATH belongs to the remote host), under
=--dry-run= (it would spawn =--version= subprocesses), and when an
=agent_launcher.nvm= block is configured, since that pins resolution deliberately
and it happens inside the nvm shell.

** Tool versions on the target (launch preflight)
The section above answers /which agent binary/ runs.  This one answers a
neighbouring question the shipped prompts and skills make unavoidable: the
agent is told to use =gh=, =git=, =jq=, =rg= and =tmux=, and a session inherits
whatever happens to be in the target's =$HOME=.  Until now =sucoder= never said
how old any of it was.

Before every launch, =sucoder= runs one probe on the target (a generated bash
script fed to =bash -l -s=, so it resolves under the same login PATH the agent
will get) and logs what it found:

#+begin_example
INFO Tool versions on n0043.savio4:
  gh: gh version 2.67.0 (2025-02-11), floor 2.90.0 -- BELOW FLOOR [/global/home/users/u/bin/gh]
  git: git version 2.43.7, floor 2.34.0 [/usr/bin/git]
  jq: jq-1.6, floor 1.6 [/usr/bin/jq]
  rg: not found, floor 13.0.0
  tmux: tmux 3.7c, floor 3.0 [/global/home/users/u/bin/tmux]
WARNING Tool preflight: gh: gh version 2.67.0 (2025-02-11), floor 2.90.0 -- BELOW FLOOR [...]
        (this is a HOST tooling fact, not a property of the repository; sucoder
        does not manage these binaries and did not block the launch)
#+end_example

*Why this exists.*  A target's =gh= was 2.67.0 (February 2025).  Every =gh pr
edit= and =gh issue view= failed with =GraphQL: Projects (classic) is being
deprecated ... (repository.pullRequest.projectCards)=---an error naming a GitHub
product sunset, emitted for a command carrying no project flags.  An agent
concluded /"gh pr edit is broken on this repo"/ and wrote that up as a
repo-specific gotcha for others to inherit.  The misattribution, not the outage,
is what the preflight prevents: the version is in the log before the agent
starts.  It also makes visible, for the first time, the case where two targets
carry different versions and a workflow verified on one silently fails on the
other.

*What it deliberately does not do.*  It never gates---a stale tool must not stop
a session starting, and nothing in the check can fail a launch, including its own
bugs.  It never touches the network: the floor is a constant in the config file,
because a preflight that needs the internet fails on exactly the constrained
targets it is meant to help.  And it does not manage the tools; they live in the
user's =$HOME=, shared across targets, and =sucoder= is not a package manager.
Replacing a running, user-owned static binary is =mv= (which replaces the
directory entry and leaves running processes on the old inode), never =cp= over
the top---several targets may share one =$HOME= and another session may be
mid-call.  Auth and config are unaffected; for =gh= they live in =~/.config/gh/=.

A version string it cannot read (=tmux 3.3a=, =jq-1.6= and =gh version 2.67.0
(2025-02-11)= are all handled, but the world is larger than five tools) is
reported as unreadable---recorded, and neither a pass nor a warning.

*** Floors and which tools are probed
Both live in one place, =tool_preflight= in =config.yaml=:

#+begin_src yaml
tool_preflight:
  # enabled: false          # skip the probe at launch; `sucoder doctor` still runs it
  floors:
    gh: "2.95.0"            # override a shipped floor
    tmux: null              # record the version, stop judging it
    shellcheck: "0.9.0"     # add a tool (probed with --version)
#+end_src

The keys of =floors= are what gets probed, so one knob both adds a tool and sets
its floor, and =null= silences a floor without losing the reading.  Entries are
merged over the shipped defaults (=gh 2.90.0=, =git 2.34.0=, =jq 1.6=, =rg
13.0.0=, =tmux 3.0=), so naming one tool does not quietly stop checking the rest.
A floor =sucoder= cannot parse is rejected when the config loads, not at launch.

*A floor is a staleness threshold, not a correctness boundary.*  It says "old
enough that a server-shaped failure should be suspected of being the client",
which is the whole job; it does not certify that a passing version is free of
any particular incompatibility---and for the incident above there is no version
that would.  [[https://github.com/cli/cli/issues/11983][cli/cli#11983]] is an /issue/, closed 2025-10-21 by its own reporter
with no linked commit; duplicates kept arriving through March 2026 (#12476,
#12640, #13069); the two PRs that would have fixed it (#13083, #13282) were both
closed *unmerged*; and =projectCards= is still referenced across cli/cli today,
=pkg/cmd/pr/edit/edit.go= included.  What /is/ established is two readings:
2.67.0 failed on the affected target and 2.101.0 succeeded on it.  The shipped
=gh= floor sits between them---near enough to current to catch a genuinely stale
client, far enough back that a fortnightly release does not re-fire it.  All five
defaults are conservative choices, not findings.

*** =sucoder doctor=
The same check on demand:

#+begin_src sh
sucoder doctor [<mirror>]
#+end_src

Unlike the launch-time preflight it *exits non-zero* when a tool is missing or
below its floor (the same convention as =sucoder tunnel doctor=), and it prints
the =mv=-not-=cp= note above.  It refuses a non-confined SLURM target rather than
quietly running =salloc=: a diagnostic must not bill a compute allocation.  Run
it without =-T= (or against a confined target) to probe the login node, or read
the report the launcher already wrote to the session log.

*Caveat worth knowing.*  For a confined (=sbatch=) target the preflight runs on
the login node at launch time, while the agent runs on a compute node.  Tools
under a shared =$HOME= (=gh=, =rg=, =jq=) are the same binary; a system =git= or
=tmux= need not be.  The probed hostname is printed in the report so that
difference is visible where it is read.

** Pin agent Node version with nvm
Some agents (e.g., Codex) bundle their own Node.js runtime.  =sucoder= can wrap the
launch so that the agent runs under a newer nvm-managed runtime instead (for example Node 22).

1. Install/activate the desired version for the agent user:
   #+begin_src shell
   sudo -u coder bash -lc 'export NVM_DIR=/home/coder/.nvm; . "$NVM_DIR/nvm.sh"; nvm install 22.11.0'
   sudo -u coder bash -lc 'export NVM_DIR=/home/coder/.nvm; . "$NVM_DIR/nvm.sh"; nvm alias default 22.11.0'
   #+end_src
   Replace the version string with the release you need (direct versions, =lts/*= aliases, etc.).

2. Update the mirror configuration so launches always source nvm before running the agent:
   #+begin_src yaml
   mirrors:
     project:
       canonical_repo: ~/src/project.git
       agent_launcher:
         command:
           - codex
         nvm:
           version: "22.11.0"   # Any version/alias `nvm use` understands
           dir: /home/coder/.nvm  # Optional; defaults to <agent home>/.nvm
   #+end_src

When the =nvm= block is present, the launcher injects:
- =export NVM_DIR=…; source "$NVM_DIR/nvm.sh"=
- =nvm use <version>= (fails fast if the version is missing)
- =exec <agent> …= so the rest of the CLI arguments (including any injected flags and
  context prelude) are preserved.

If =dir= is omitted the helper assumes the agent’s home directory contains =~.nvm=.
Any existing =agent_launcher.command= and =agent_launcher.env= settings continue to work.

Workspace skills follow Anthropic’s Agent Skills Spec:
- Each skill directory must contain a =SKILL.md= file with YAML frontmatter (`name`, `description`,
  and `license`; add `allowed-tools`/`metadata` as needed). If you copy a third-party skill, keep its
  license value and include the upstream notice in the directory.

** Unified skills catalog (local + curated)
- Generate a tool-agnostic catalog of local skills (from the configured skills directory) plus the openai/skills curated list:
  #+begin_src shell
  python scripts/generate_unified_skills_catalog.py \
    --output ../Skills/UNIFIED_SKILLS_CATALOG.md   # choose any writable path
  #+end_src
  The script reads =../Skills/trusted_skill_sources.yaml= to determine which sources to include (local catalogs plus curated repos; defaults include =openai/skills= and =anthropics/skills=). If a target is not writable (for example =/home/ligon/.sucoder/skills=), use =--output= to pick a writable path.
- The generated catalog shows entries per source and flags conflicts when the same skill name appears in multiple sources.
- Any LLM can read the catalog and follow linked =SKILL.md= files. Curated entries are documentation-first; helper scripts are optional.
- To install a curated skill manually, download =skills/.curated/<name>= from https://github.com/openai/skills and place it under your skills directory (e.g., =~/.sucoder/skills=).
- The frontmatter `name` must match the directory name exactly.
- Long-form references, scripts, and assets live under =references/=, =scripts/=, and =assets/= to
  keep the entrypoint concise.
- The shared catalog (=~/.sucoder/skills/SKILLS.md=) in the skills repository lists every bundled skill so agents discover them automatically.
- A reusable system prompt template (`default_system_prompt.org`) mirrors these expectations, including a reminder to confirm today's date at session start.
- The upstream workflow is available in the =sucoder-skills= repository under =skill-creator/=; use it alongside =document-skill= when you need initialization or packaging scripts.
- A starter configuration (`default_config.yaml`) points at the bundled skills directory and default prompt; copy or merge it into =~/.sucoder/config.yaml= as needed.

Resource directories are summarized automatically when present:
- =references/= — documentation to load on demand (each file includes a suggested load command).
- =scripts/= — helper executables (never run automatically; humans can execute after review).
- =assets/= — supporting files such as templates or images that may be referenced in outputs.

* Skills Repository

Skills are maintained in a separate repository (=sucoder-skills=) for security isolation.
This separation prevents agents from modifying skills and tool code in the same commit, reducing
the attack surface and simplifying security reviews.

** Why Skills Are Separate

1. *Security Isolation* — An agent working on tool code cannot simultaneously modify skills
   that influence agent behavior.
2. *Review Clarity* — Code reviews and content reviews use different security mindsets and
   can be performed separately.
3. *Blast Radius Reduction* — Malicious skill instructions cannot be hidden among tool code changes.

** Version Compatibility

The tool enforces semantic versioning compatibility between itself and the skills repository:

- Tool validates skills repository =VERSION= file on startup
- Compatible: Tool requires =1.0.0=, skills has =1.x.x= ✓
- Incompatible: Tool requires =1.x.x=, skills has =2.0.0= ✗

Skills repository: https://github.com/ligon/sucoder-skills

To update skills:
#+begin_src shell
cd ~/.sucoder/skills  # Or ~/Projects/sucoder-skills
git pull
#+end_src

To skip version checking (not recommended):
#+begin_src shell
export SUCODER_SKIP_SKILLS_VERSION=1
#+end_src

** Version Bumping Strategy

*PATCH (1.0.0 → 1.0.1)*: Typo fixes, clarifications, examples
*MINOR (1.0.0 → 1.1.0)*: New skills, new optional features, backward-compatible improvements
*MAJOR (1.0.0 → 2.0.0)*: Breaking changes, requires tool update

* Agent-Agnostic Project Instructions

Projects can provide agent instructions and skills using agent-agnostic
names.  During mirror setup, =sucoder= creates symlinks so that Claude
discovers them natively while other agents receive the content via
prompt injection.

| Canonical file     | Symlink created          | Purpose                        |
|--------------------+--------------------------+--------------------------------|
| =AGENT.md= / =.org= | =CLAUDE.md -> AGENT.md=  | Project-level instructions     |
| =.skills/=          | =.claude/skills -> .skills= | Project-level skill files   |

- Claude discovers =CLAUDE.md= and =.claude/skills= natively via the symlinks.
- Non-Claude agents receive =AGENT.md= content in the system prompt.
- One source of truth; no duplication across agent types.

Symlinks are created during =ensure_clone= (both fresh clones and
subsequent launches).  If =CLAUDE.md= or =.claude/skills= already
exist independently, =sucoder= leaves them in place and does not
overwrite.

* Agent Skills Tracking

Agents may write skill files to =~coder/.claude/skills/= during sessions.
These persist across sessions and influence future agent behavior.
=sucoder= git-tracks this directory to maintain an audit trail:

- The directory is initialized as a git repo on first use.
- After each subprocess-mode session, any changes are auto-committed
  with a message referencing the mirror that produced them.
- The commit history provides forensics: when a skill was added,
  which session produced it, and what it looked like before modification.

* Auditing

A dedicated compliance agent reviews changes made by the working
agent---both skill files and code.  The auditor runs as a separate
Unix user (=auditor=) that is intentionally /not/ in the =coder=
group, so write access to agent files is prevented by Unix permissions
rather than prompt instructions.  Read access is granted via
world-readable bits (=o+r=).

** Setup

#+begin_src shell
make create-auditor-user   # create auditor user (requires sudo)
make auditor-perms         # set o+r on paths the auditor needs
# Install Claude CLI and authenticate (one-time):
sudo -u auditor bash -c 'curl -fsSL https://claude.ai/install.sh | bash'
sudo -u auditor claude login
# Copy auditor prompts:
cp default_auditor_prompt.org ~/.sucoder/auditor_prompt.org
cp default_code_auditor_prompt.org ~/.sucoder/code_auditor_prompt.org
#+end_src

** Usage

The =--scope= flag controls what is audited: =skills= (default), =code=, or =all=.

#+begin_src shell
# Skills audit (default) — reviews ~coder/.claude/skills/
sucoder audit              # diff review since last approved baseline
sucoder audit --full       # review all skills from scratch
sucoder audit --approve    # advance the baseline after review

# Code audit — reviews mirror repo changes
sucoder audit MIRROR --scope code         # diff since last code baseline
sucoder audit MIRROR --scope code --full  # review all tracked code
sucoder audit MIRROR --scope code --approve

# Both at once
sucoder audit MIRROR --scope all
#+end_src

** Audit scopes

*** Skills audit (=--scope skills=)

Reviews agent-written skill files in =~coder/.claude/skills/=.
Uses =~/.sucoder/auditor_prompt.org= as the system prompt.
Tracks its baseline via =refs/audited= in the skills git repo.

The auditor checks for instruction injection, scope creep,
contradictions, sensitive data, and self-modification loops.

*** Code audit (=--scope code=)

Reviews code changes in a mirror repository.  Requires a mirror
name argument.  Uses =~/.sucoder/code_auditor_prompt.org= as the
system prompt.  Tracks its baseline via =refs/audited-code= in the
mirror.

The auditor checks for:
1. Dependency injection (malicious or unexpected packages)
2. Credential leakage (hardcoded tokens, API keys, passwords)
3. Unsafe subprocess calls (=shell=True=, =eval=, =exec=)
4. Permission escalation (=chmod 777=, setuid, sudoers edits)
5. Unexpected network calls (outbound connections to unknown hosts)
6. Overly broad file operations (recursive deletes, out-of-tree writes)
7. Supply-chain risks (typosquatting, unusual version pinning)
8. Obfuscated code (base64-encoded strings, minified inline scripts)

** How it works

1. *Permissions check* --- Before examining content, the auditor verifies
   that all files are world-readable.  Files the working agent has made
   unreadable (=chmod o-r=) are flagged immediately as suspicious.

2. *Audit mode* --- If no baseline exists (=refs/audited= for skills,
   =refs/audited-code= for code), a full review runs.  Otherwise, only
   the diff since the last approved commit is examined.

3. *Auditor invocation* --- =claude -p= is run as the =auditor= user with
   a scope-specific system prompt.  The skills auditor and code auditor
   use different prompts tuned to their respective concerns.

4. *Report* --- The auditor prints a summary.  If concerns are found, each
   is listed with filename and severity.  Pass =--approve= to advance the
   baseline; omit it to leave the baseline unchanged for human review.

** Auto-trigger after each session

Auditing on demand only catches issues when the human remembers to run
=sucoder audit=.  To close that loop, opt in to a post-session auto
trigger via =~/.sucoder/config.yaml=:

#+begin_src yaml
audit:
  auto_after_session: true   # default false; explicit opt-in
  scope: all                  # skills | code | all
#+end_src

When enabled, =sucoder collaborate= (and other launch paths) call the
audit subsystem after =_auto_commit_agent_skills= runs.  The auditor
operates exactly as if you had typed =sucoder audit MIRROR --scope <X>=,
with three differences:

  - Reports are saved to =<log_dir>/audits/<mirror>-<kind>-<timestamp>.log=
    rather than printed to stdout (the session has ended; nobody is
    watching).  When =log_dir= is unset, =~/.sucoder/logs/audits/= is
    used.

  - A one-line summary is logged: =INFO: Post-session code audit: no
    concerns (<path>)= when the report says so, =WARNING: Post-session
    code audit produced findings: <path>= otherwise.  Watch =log_dir=
    or pipe sessions through =tee= if you want a durable record.

  - Failures are non-blocking.  An expired auditor token, missing
    auditor user, network blip --- all log a warning and let session
    teardown succeed.  Audit infrastructure must never make a
    successful agent session look failed.

*Pre-flight*: the first auto-audit against a never-audited mirror has
no =refs/audited= / =refs/audited-code= baseline, so it runs in *full*
mode.  That's the most expensive variant in LLM tokens.  Run
=sucoder audit MIRROR --scope all --approve= once after you're happy
with the initial review; subsequent auto-audits then run in cheap diff
mode (or skip entirely when there are no changes).

** Relationship with cq

The =cq= knowledge commons and agent skills serve complementary roles:

| Aspect    | Skills                          | cq                              |
|-----------+---------------------------------+---------------------------------|
| Content   | Curated instructions/procedures | Discovered learnings/pitfalls   |
| Lifecycle | Human-authored or promoted      | Agent-proposed, human-confirmed |
| Scope     | Per-project or shared           | Cross-project                   |

The intended flow: agents propose learnings to =cq=; knowledge that proves
itself over time gets promoted to a skill by a human.  Agent-written skills
in =~coder/.claude/skills/= are a fast path for immediate session context,
subject to compliance audit.

* Testing
Run the automated tests with pytest.

#+begin_src shell
pytest
#+end_src

About

Sandbox for coding agents based on unix filesystem permissions

Resources

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages