Folders and files
| Name | Name | Last commit date | ||
|---|---|---|---|---|
Repository files navigation
#+title: sucoder [[https://doi.org/10.5281/zenodo.21629611][https://img.shields.io/badge/DOI-10.5281%2Fzenodo.21629611-blue.svg]] * Project Overview Unix user and group permissions have been battle-tested for over fifty years. When you bring on a new collaborator you give them their own account, set group read on shared files, and let the filesystem enforce the boundaries. Why should an AI coding agent be any different? =sucoder= treats an LLM agent (running through a harness such as Codex, Claude, Gemini, Aider, OpenCode, Goose, or Kimi) as a collaborator with its own unix account. The human's canonical repository is group-readable but not group-writable; the agent works in a sandboxed mirror clone where it has full write access. No custom container runtimes, no bespoke sandboxing daemons---just =chown=, =chmod=, and =git=. * Quick Start #+begin_src shell git clone https://github.com/ligon/sucoder && cd sucoder make quick-start cd ~/Projects/my-project # any git repo sucoder collaborate --harness claude #+end_src No config file needed. =sucoder= detects the git repo, mirrors it under =/var/tmp/coder-mirrors/=, and launches Claude (the default harness). Use =--harness/-H= to pick a different harness (=codex=, =gemini=, =aider=, =opencode=, =goose=, or =kimi=), and =--model/-m= to choose its LLM independently. The old =--agent/-a= spelling remains a compatibility alias for =--harness=. Add =--task fix-login= to start on a dedicated branch. For example, the same Aider harness can run models from different providers: #+begin_src shell sucoder collaborate --harness aider --model openai/gpt-5 sucoder collaborate --harness aider --model anthropic/claude-sonnet-4 sucoder collaborate --harness kimi --model openrouter/moonshotai/kimi-k3 #+end_src Provider credentials can remain in the human user's =pass= store and be selected by the model's provider prefix; see [[*Provider credentials from =pass=][Provider credentials from =pass=]]. Harnesses must be on the configured =agent_user='s login =PATH=. On a multi-user host, prefer one root-owned uv tool installation so human and agent accounts cannot drift independently; for example, when =uv= itself is globally available: #+begin_src sh sudo env UV_TOOL_DIR=/opt/uv-tools \ UV_TOOL_BIN_DIR=/usr/local/bin \ UV_PYTHON_INSTALL_DIR=/opt/uv-python \ uv tool install --python python3.12 aider-chat # OpenCode, Codex, and Kimi share the system Node installation. sudo npm install -g opencode-ai @openai/codex @moonshot-ai/kimi-code # Goose is a native binary; its official installer accepts a shared target. curl -fsSL https://github.com/aaif-goose/goose/releases/download/stable/download_cli.sh \ -o /tmp/goose-download-cli.sh sudo env GOOSE_BIN_DIR=/opt/goose/bin CONFIGURE=false \ bash /tmp/goose-download-cli.sh sudo ln -s /opt/goose/bin/goose /usr/local/bin/goose #+end_src The separate shared Python directory matters when uv downloads a managed interpreter: otherwise a root-run install can leave the tool pointing into root's private home directory. ** Provider credentials from =pass= Keep API keys in the human user's password store and put only entry names in =~/.sucoder/config.yaml=. The first line of each entry is the secret: #+begin_src yaml credentials: openrouter: pass: openrouter.ai/apikey #+end_src OpenRouter's protocol, endpoint, and conventional environment variable are built in. A model such as =openrouter/moonshotai/kimi-k3= therefore selects both the provider credential and the upstream model. Sucoder runs =pass show= as =human_user= at launch; the =coder= account receives neither password-store access nor the human's GPG keys. For Aider and other environment-aware harnesses, Sucoder supplies =OPENROUTER_API_KEY=. For native Kimi it uses Kimi's temporary =KIMI_MODEL_*= provider channel, so the key is not copied into =~coder/.kimi-code/config.toml=. Environment values are staged over stdin in a random agent-owned mode-0600 file, then sourced and unlinked immediately before the harness starts. They do not appear in the launch command or Sucoder logs. An agent with shell access can still inspect its own process environment; keeping the raw provider key beyond the agent's reach requires a separate rate-limited model gateway. Custom providers can be defined once: #+begin_src yaml credentials: laboratory: pass: research/lab-gateway providers: laboratory: credential: laboratory protocol: openai base_url: https://models.example.edu/v1 env_var: LABORATORY_API_KEY #+end_src Literal =agent_launcher.env= and =--agent-env= remain supported and use the same private-file transport, but command-line values may remain in shell history. Prefer named =pass= credentials for secrets. For multi-repo setups, skills, and system prompts, see [[*Configuration][Configuration]] below. =make env-setup= does the full host provisioning (=make help= for all targets). * Requirements - Linux (user/group sandboxing relies on =useradd=, =groupadd=, and setgid) - Python >= 3.9 with pip - git - git-crypt (for decrypting =.mcp.json= which contains MCP server tokens) - Node.js and npx (MCP servers are distributed as npm packages) - sudo - make (optional; you can run the setup commands by hand) - bash - At least one supported harness CLI on the agent user's PATH: =claude=, =codex=, =gemini=, =aider=, =opencode=, =goose=, or =kimi= * Why Org-mode? This README is =.org=, the default system prompt is =.org=, and the skills repo leans heavily on Org-mode. Why not Markdown? Org-mode is a plain-text format that every LLM reads fluently—but it also has structured features (TODO states, properties, tags, executable source blocks) that Markdown lacks. For a project about human-agent collaboration, those features pull their weight: a system prompt can carry metadata an agent can act on, a skill file can embed runnable examples, and a handoff note can be a living checklist rather than a static document. You do not need Emacs to read or edit these files. Any text editor works; GitHub renders =.org= natively. But if you /do/ use Emacs, the integration is a bonus, not a requirement. * What =sucoder collaborate= does Running =sucoder collaborate --harness claude= (or with =--task=) triggers a multi-step workflow. Steps marked *auto* happen without intervention; steps marked *human* need you at the keyboard. ** 1. Resolve configuration /auto/ If =~/.sucoder/config.yaml= exists, load it. Otherwise, derive everything from the environment: =$USER= becomes the human identity, the git repo root becomes the canonical repo, and =/var/tmp/coder-mirrors/= is the mirror root. The harness CLI is resolved from =--harness= (or legacy =--agent=) > =$SUCODER_AGENT= > =~/.sucoder/agent= > PATH auto-detect > interactive prompt. =--model= is then applied independently through the selected harness profile. ** 2. Prepare the canonical repository /auto/ Set group-read permissions on your repo so the =coder= user can traverse it. Adds a =coder= git remote pointing at the mirror (for easy fetching later) and writes a helper script =scripts/fetch-agent-branches.sh=. Your repo stays read-only to the agent—only the group bits and the remote config change. ** 3. Clone (or verify) the mirror /auto/ If the mirror does not exist under =/var/tmp/coder-mirrors/<repo>/=, clone it from the canonical repo as the =coder= user. Push access back to canonical is disabled (=no_push=). If the mirror already exists, verify the remote and enforce permissions. ** 4. Sync or create a task branch /auto/ - Without =--task=: fetch the latest commits from canonical into the mirror. - With =--task fix-login=: create a branch =coder/fix-login-<timestamp>= in the mirror, based on the default base branch (or =--base=). ** 5. Compose the agent launch command /auto/ Detect the harness type (Claude, Codex, Gemini, Aider, OpenCode, Goose, Kimi) from the command name and inject the appropriate flags: model, write-access (=--dangerously-skip-permissions=, =--sandbox danger-full-access=, =--yolo=, =--yes-always=, =--auto=), writable directories, skills paths, and system prompt. Aider receives the composed context through a private read-only file so it remains interactive. Kimi receives a private custom-agent file which includes its native base prompt before the SuCoder context; other harnesses use their native prompt flag or trailing text. ** 6. Launch the agent /human/ The agent process starts inside the mirror working tree. From here, *you are talking to the agent*. It has full write access to the mirror but cannot write to your canonical repo. The agent commits its work to the task branch. When the agent session ends (you quit it, or it finishes), =sucoder= auto-commits any changes the agent made to =~coder/.claude/skills/= (see [[*Agent Skills Tracking][Agent Skills Tracking]]), then control returns to your shell. ** 7. Review and integrate /human/ Fetch the agent's branch into your canonical repo and review it: #+begin_src shell git fetch /var/tmp/coder-mirrors/my-project coder/fix-login-20250215180000 git checkout -b review/fix-login FETCH_HEAD # ... review, test, iterate ... git checkout main && git merge --ff-only review/fix-login #+end_src Or pull directly: #+begin_src shell git pull --ff-only /var/tmp/coder-mirrors/my-project coder/fix-login-20250215180000 #+end_src The canonical repository remains unwritable by the agent throughout. ** Running the steps individually =collaborate= bundles steps 2–6. You can also run them separately: | Command | Step | |----------------------+----------------------------------------------| | =prepare-canonical= | Set permissions and add agent remote (step 2)| | =agents-clone= | Clone the mirror (step 3) | | =push= | Send canonical commits to the mirror (step 4)| | =start-task= | Create a task branch (step 4) | | =agents-run= | Launch the agent (steps 5--6) | | =worktrees= | List active worktrees with status | | =attach= | Reconnect to a remote tmux session | | =audit= | Run compliance review of skills and/or code | | =list= | Discover harnesses, models, mirrors, or skills| * Key Features - agents-clone :: Clone a canonical repository into an agent-owned mirror with strict permissions. - push :: Send the canonical repository's commits to the mirror without granting write access. Pulls the agent's commits back first, and refuses rather than force-push over a mirror it could not read. Available as =sync= too. - start-task :: Create a new agent task branch based on a human branch while preserving branch naming conventions. - status :: Summarize the state of a mirror, including permissions and git remotes. - collaborate :: Prepare canonical, ensure mirror, and launch agent in one step. - worktrees :: List active git worktrees in a mirror with per-worktree status (branch, commits ahead, dirty state). Supports =--watch= for live monitoring and =--diff= for file-level changes. - attach :: Reconnect to a remote agent session via tmux after an SSH disconnect. - nodes :: Show compute-node availability for a SLURM partition (read-only =sinfo= query over the warm gateway tunnel; see [[*Inspecting node availability][Inspecting node availability]]). - tunnel :: Keep the free SSH hops to a target warm (=up= / =status= / =doctor= / =down=), and forward a compute-node service port to localhost (=forward= / =forwards=) so a web app served on a node (Jupyter, claude-science, ...) is reachable in a local browser without new authentication. - audit :: Run a compliance review of agent-written skills and/or mirror code changes (see [[*Auditing][Auditing]]). - list :: Discover built-in and configured harnesses, their capabilities, configured model defaults, harness model catalogs, mirrors, and skills. ** Discover harnesses and models The =list= command groups discovery by resource: #+begin_src shell sucoder list harnesses sucoder list models sucoder list models kimi sucoder list models --provider openrouter kimi sucoder list models --harness aider gpt sucoder list models --harness opencode kimi sucoder list models --harness kimi k3 sucoder list providers sucoder list mirrors sucoder list skills #+end_src =list harnesses= shows every built-in profile plus custom harness commands found in configuration, where each is configured, its default model, and native shell, file-tool, Agent Skills, MCP, subagent, provider, and approval capabilities. =suggest= means the harness can print a command but does not expose a model-driven shell tool; =?= means SuCoder has no built-in metadata for a custom harness. =list models= queries the sole credentialed provider automatically; with several credentialed providers, select one using =--provider=. The resulting names include the provider prefix and can be passed directly to =--model=. If no provider is configured, the command falls back to configured model defaults. =--harness aider=, =--harness opencode=, and =--harness kimi= instead query the selected harness under the configured agent user. Kimi reports configured aliases. =list providers= shows endpoints and credential references without invoking =pass= or displaying secrets. The optional positional argument filters the selected model catalog. The older =mirrors-list= and =skills-list= spellings remain supported for compatibility. - --target/-T :: Select a named remote execution target (e.g., =--target savio=) to run the agent on a remote host. - --version/-V :: Print the installed sucoder version (CalVer, derived from git tags). - --dry-run :: Available on modifying commands to review logged operations without executing. * Usage #+begin_src shell # Clone the canonical repository into the configured mirror. sucoder agents-clone project # Send canonical's commits to the mirror (inverse of `pull`). sucoder push project # Create a task branch for the agent starting from ligon/main. sucoder start-task project rewrite-parser --base main # Inspect current branch and outstanding changes. sucoder status project #+end_src ** Moving commits between canonical and the mirror - push :: canonical --> mirror. For remote mirrors the mirror's branches are overwritten to match canonical. For local mirrors canonical's commits become *visible* in the mirror as =<remote>/<branch>=; the mirror's own branches are left where the agent put them. - pull :: mirror --> canonical. Brings the agent's commits back, with a prompt when the two histories have diverged. - sync :: an alias for =push=, kept for compatibility. =agents-clone= creates the mirror. Remote startup initializes absent or empty directories and recovers valid repositories in place. It never deletes a directory because a probe failed or a repository has no default branch. Unreadable, invalid, symlinked, or non-repository nonempty paths stop startup with a diagnostic; inspect them before retrying. ** Mirror safety: the pull must succeed before the push =push= (and =agents-clone=) send to the mirror with =git push --all --force=, so they first pull whatever the agent committed there. If that pull cannot read the mirror --- host unreachable, allocation gone, the agent's work sitting on a branch other than the configured base --- the push is refused rather than allowed to overwrite unretrieved commits. An empty or half-initialised mirror is *not* treated as unreadable, so first-time bootstrap is unaffected: sucoder asks the mirror whether it holds any commits instead of guessing from the error text. Any ref (including feature branches, tags, and WIP refs) counts as content, even with an unborn HEAD. A repository initialized by this invocation skips the initial fetch and receives a non-forcing first push. The =--allow-unverified-mirror= override does not bypass initialization checks. - --allow-unverified-mirror :: Push anyway, discarding any unpulled mirror commits. Use when the mirror is known to be expendable. Likewise =pull= exits non-zero when it could not read the mirror, instead of reporting =Pull complete.= after pulling nothing. * Configuration Skip this section if zero-config mode works for you. Configuration is stored in YAML (default =~/.sucoder/config.yaml=). To set it up: 1. Copy the sample configuration and adjust paths, users, and mirror names. #+begin_src shell install -d -m 750 ~/.sucoder cp config.example.yaml ~/.sucoder/config.yaml #+end_src 2. Update =~/.sucoder/config.yaml= so that: - =human_user= matches your login (for example =ligon=). - =agent_user= and =agent_group= match the agent Unix account (default =coder=). This is a shared account used by any supported agent (Codex, Claude Code, Gemini CLI, etc.)---it is not tied to a specific agent binary. - =mirror_root= points to the directory where agent mirrors will live. - =mirrors.<name>.canonical_repo= points at the human-owned repository. 3. Ensure the agent can read (but not write) the configuration directory and file. #+begin_src shell sudo chgrp coder ~/.sucoder sudo chmod 750 ~/.sucoder sudo chgrp coder ~/.sucoder/config.yaml sudo chmod 640 ~/.sucoder/config.yaml #+end_src 4. (Optional) Install the skills repository for agent capabilities. #+begin_src shell git clone https://github.com/ligon/sucoder-skills ~/Projects/sucoder-skills ln -s ~/Projects/sucoder-skills ~/.sucoder/skills sudo chgrp -R coder ~/Projects/sucoder-skills sudo chmod -R g+r,g-w ~/Projects/sucoder-skills #+end_src *Note*: The skills repository is separate from the tool repository for security isolation. Skills influence agent behavior and are kept in a dedicated repository to prevent agents from modifying skills and tool code in the same commit. The tool enforces semantic versioning compatibility between the tool and skills repository. 6. Warm up the shared skills catalog so launches confirm access before work starts. #+begin_src shell ls ~/.sucoder/skills sucoder list skills sucoder list mirrors #+end_src The listings should succeed without permission errors. The CLI helper highlights any unreadable paths and =sucoder list mirrors= prints the configured mirrors (with their canonical and mirror directories). To preload the catalog into a session, use the agent's file-read command (e.g., =codex read ~/.sucoder/skills/SKILLS.md= for Codex). ** Shell Completion Enable tab completion for =sucoder= commands and mirror names by running: #+begin_src shell sucoder --install-completion #+end_src ** Sudo Requirement By default, =sucoder= uses =sudo= to run commands as the agent user (=coder=). The human user must have sudo access to impersonate the agent user. If you are running as the agent user directly (for example inside a container), pass =--no-agent-sudo= to skip sudo: #+begin_src shell sucoder --no-agent-sudo agents-clone project #+end_src ** Command Overview - =collaborate= is the all-in-one command: it runs =prepare-canonical= + =agents-clone= (ensure mirror) + =agents-run= (launch agent) in a single step. - =agents-run= only launches the agent, assuming the mirror already exists. Use this when the mirror is already prepared and you just want to start a session. ** Provision the Mirror 1. Prepare the canonical repository for collaboration. #+begin_src shell sucoder prepare-canonical project #+end_src *Note*: this command modifies the canonical repository — it sets the group to =coder=, grants group read permissions, removes group write bits, and optionally adds an agent remote and fetch helper script. Elevate with =--sudo= when root privileges are required. By default it also adds a =coder= remote pointing at the agent mirror and writes =scripts/fetch-agent-branches.sh= to simplify fetching agent branches (=--no-agent-remote= disables this). 2. Clone the mirror as the human operator; the tool will invoke git as the agent. #+begin_src shell sucoder agents-clone project --verbose #+end_src Mirror arguments support shell completion, so pressing Tab after =sucoder <command>= suggests configured mirror names. ** Sync Before Handing Off 1. Refresh the mirror with the latest human commits. #+begin_src shell sucoder sync project #+end_src 2. Create a task branch for the agent based on the desired human branch. #+begin_src shell sucoder start-task project add-metrics --base main #+end_src The command prints the full branch name (for example =coder/add-metrics-20251107164028=). ** Agent Workflow - The agent works directly inside the mirror directory (for example =/var/tmp/coder-mirrors/project=) on the branch created above. - The agent commits work to =coder/<task>-<timestamp>= without pushing to the canonical repository. - To run the full workflow in one step (prepare canonical, ensure mirror, launch agent), use: #+begin_src shell sucoder collaborate project --task add-metrics #+end_src # Adjust flags such as =--no-agent-remote=, =--sudo=, or extra arguments just as you would with the individual commands. - To select a harness and model independently, or experiment with an arbitrary binary, use =--harness=/=--model= or =--agent-command= (and optional =--agent-env= for non-secret overrides): #+begin_src shell sucoder agents-run project --harness aider --model openrouter/deepseek/deepseek-chat sucoder agents-run project --agent-command "foo --flag" --agent-env DEBUG=1 #+end_src - The CLI automatically injects the appropriate write-access flag for the detected agent (e.g., =--yolo= for Codex, =--dangerously-skip-permissions= for Claude Code; see the flag template table below). When invoking an agent manually, include the equivalent flag to avoid read-only failures. The helper will also warn if it detects a read-only sandbox and still attempts a launch. - Org authoring playground (documentation example): load =docs/examples/org-authoring-demo.org= using your agent's file-read command. Demonstrates the in-repo Org skills—structure, markup, tables, capture, citations, exporting, agenda integration—so humans and agents have a quick reference. *** Codex remote control Use =--agent-command= for a command with subcommands; =--agent= and =--harness= select a single executable. For example, on a configured Savio target: #+begin_src shell sucoder -T savio collaborate project --agent-command "codex remote-control" #+end_src =--agent-command "codex remote-control start"= is also accepted. SuCoder converts =start= to the foreground form and reports that choice, so tmux and Slurm retain ownership of the service process. Commands such as =stop= and =pair= are management operations; run them separately on the remote host. The remote Codex installation must support =remote-control= and have the authentication and pairing required by the connected Codex client. See the [[https://learn.chatgpt.com/docs/developer-commands?surface=cli#codex-remote-control][Codex remote-control reference]]. The service starts in the mirror's working directory, including the existing local-disk clone and Slurm confinement when configured. The full SuCoder prelude (system prompt, target instructions, and skills catalog) is installed in =AGENTS.override.md= in the service's actual working directory. This keeps the defaults available when a connected client replaces Codex's =developer_instructions= config value; SuCoder also supplies that config value for clients that retain it. SuCoder preserves an existing local override's text, or copies =AGENTS.md= into a separate generated region when creating the override. It refreshes generated regions on each service launch and adds the override to Git's local exclude file. The original =AGENTS.md= stays unchanged; its copy refreshes at launch, and the agent is directed to read the current file before working. The launcher needs =python3= on the agent host and allows the prelude size plus 1 MiB for native instruction files; explicit =project_doc_max_bytes= overrides still take precedence. An edited or ambiguous generated block, a concurrent edit, or a symlinked or Git-tracked override stops launch with an error and preserves the file. The file remains available to later Codex chats in that mirror. Choose the mirror's working directory in the connected client: native instructions are scoped to that directory and its children. After updating SuCoder, restart the service and open a new chat to load these instructions; reattaching to an existing service does not refresh it. =--model= and =agent_launcher.model= become a =model= config override. The usual =needs_yolo= intent becomes =sandbox_mode="danger-full-access"= and =approval_policy="never"= config defaults; set =needs_yolo: false= to omit them, or supply explicit =-c= overrides to restrict them. SuCoder does not apply TUI flag templates to this subcommand; =default_flags= must contain options supported by =remote-control=. To make this the mirror's default launch: #+begin_src yaml agent_launcher: command: [codex, remote-control] #+end_src =sucoder sessions= labels the pane as a remote-control service, including when the local launcher record is missing. =sucoder message= skips it even with =--force=; send instructions through the connected Codex client. =sucoder peek= displays the service log. =sucoder renew= retains the service command and model across allocation turnover, and regenerates the SuCoder prelude. Its checkpoint request remains a file sentinel; renewal restarts the server and does not automatically checkpoint or resume active chats. ** Parallel Work with Worktrees Inside the mirror the agent can use git worktrees to work on multiple tasks in parallel. Claude supports this natively via =claude --worktree NAME=; other agents can achieve the same result with =git worktree add=. Each worktree gets its own checkout while sharing the same git object store (no duplicate clone). This is useful for: - Parallel feature development :: Fix a bug in one worktree while building a feature in another. - Comparative computation :: Run the same estimation with different parameters in separate worktrees, then consolidate results. - Review isolation :: Keep exploratory changes out of the main branch until they are ready. The human can monitor all worktrees from upstream: #+begin_src shell sucoder worktrees project # snapshot sucoder worktrees project --watch 10 # live, refresh every 10s sucoder worktrees project --diff # include file-level changes sucoder worktrees project --main # include the main worktree #+end_src Example output: #+begin_example Worktrees for mirror 'MyModel' (/var/tmp/coder-mirrors/MyModel): .claude/worktrees/instruments-A Branch: worktree-instruments-A @ a1b2c3d Ahead: 2 commits (vs main) Last: a1b2c3d GMM estimation with instrument set A (1 minute ago) Status: clean .claude/worktrees/instruments-B Branch: worktree-instruments-B @ e4f5g6h Ahead: 2 commits (vs main) Last: e4f5g6h GMM estimation with instrument set B (30 seconds ago) Status: dirty (1 modified) #+end_example The =worktrees= command uses =git worktree list --porcelain= for discovery, so it works regardless of which agent created the worktrees or where they are located. When the agent consolidates worktree branches back to the mirror's main branch, prefer =git merge --no-ff= to preserve the individual branch histories. This keeps the worktree commits individually cherry-pickable from upstream. *** Shared virtual environment for worktrees Poetry caches virtual environments by project path, so each worktree would normally create its own =.venv=. To share a single environment across all worktrees, pin =VIRTUAL_ENV= in =.envrc=: #+begin_example _git_common="$(git rev-parse --git-common-dir 2>/dev/null)" if [ "${_git_common}" = ".git" ]; then _main_tree="$(pwd)" else _main_tree="$(dirname "${_git_common}")" fi VIRTUAL_ENV="${_main_tree}/.venv" export VIRTUAL_ENV layout poetry #+end_example ** Remote Execution Mirrors can also live on a remote host (e.g., an HPC cluster) where the local machine cannot run heavyweight computation. The privilege separation shifts from "different Unix users on the same box" to "same user on different machines"---the network boundary enforces isolation. *** Ordinary Linux host over SSH A single Linux box needs only a =host= target. SuCoder runs on the launcher; install Git, Bash, tmux, and the chosen agent harness on the remote host, with their executables available to non-interactive SSH commands. The agent runs as the SSH user, without a remote sudo or separate coder account. #+begin_src yaml targets: workstation: host: workstation.example.org # or an existing ~/.ssh/config alias remote_user: ligon # optional; otherwise SSH chooses the user mirror_root: ~/mirrors # ssh_options: # Port: 2222 # IdentityFile: ~/.ssh/workstation #+end_src From the local repository: #+begin_src sh sucoder -T workstation collaborate sucoder -T workstation attach sucoder -T workstation pull sucoder -T workstation tunnel status #+end_src Execution and Git transport use the configured host directly, including after connection expiry. Normal SSH config, authentication, and host-key checking apply. =remote_user= and =ssh_options= are used for both connection startup and subsequent commands. An explicit =remote_user= overrides any =User= option, regardless of capitalization. SuCoder evaluates =ssh -G= before choosing a shared connection, so edits to SSH aliases or included config files do not reuse a connection to the previous account or host. Prefer an SSH alias for settings that should also apply to standalone =ssh= and =git= commands. =host= cannot be combined with =gateway=, =transfer_host=, =slurm=, or the BRC-specific =cert_file= option. Standard SSH certificates can instead use =IdentityFile= and =CertificateFile= in SSH config or =ssh_options=. No scheduler allocation, deadline watchdog, or WIP snapshotter is started. The mirror uses persistent storage on that host. =tunnel up= warms the one connection without writing cluster aliases to =~/.ssh/config=. =tunnel forward 8888= forwards a service on that host; =tunnel down= closes the SSH connection. =release= is a Slurm-only operation; end the remote tmux session when finished with an ordinary host. =--node= is rejected by =collaborate=, =attach=, and =tunnel forward=: a direct target has one host and no scheduler to request a node from. *** Cluster configuration Define a named target in =~/.sucoder/config.yaml=: #+begin_src yaml targets: savio: gateway: brc.berkeley.edu transfer_host: dtn.brc.berkeley.edu mirror_root: ~/mirrors remote_user: ligon # optional; remote SSH username (defaults to local $USER) control_persist: 7d # warm-tunnel idle lifetime (optional) keepalive_interval: 30 # ServerAliveInterval, seconds (optional) keepalive_count_max: 120 # ServerAliveCountMax (optional) cert_file: ~/.ssh/ssh_certs/brc_cert # gateway SSH cert (optional) x11: false # disable the default X11 forwarding (optional) #+end_src No =mirrors:= entry is needed---zero-config repo detection works with targets. The target can be used with any repo. The three SSH connection-sharing knobs are optional and default to the values shown above; tune them per target: - =control_persist= --- how long a warm =ControlMaster= lingers after the last client disconnects (ssh time format: =7d=, =12h=, =90m=...). This is the "tunnel stays open" lever: a longer value means =collaborate=, plain =ssh=, and Emacs TRAMP reuse the socket --- no PIN/OTP --- for longer. The cluster's server-side idle policy is the real ceiling, so treat a long value as best-effort, not a guarantee. - =keepalive_interval= / =keepalive_count_max= --- the =ServerAlive*= pair. Their product is the grace budget before a stalled connection is torn down (=30 x 120 = 1h= by default), which is what lets a tunnel ride out a brief network blip or a laptop nap *on the same network*. Keep the interval short (the probes also keep NAT/firewall mappings warm); raise =count_max= for a longer budget. Neither can revive a connection the server already reaped or one whose IP changed --- those still re-auth. - =cert_file= --- path to a local SSH *certificate* private key, presented on the gateway hop so =tunnel up= (and a plain =ssh <target>-gw= / TRAMP) authenticate with *no* PIN/OTP for the cert's lifetime. Mint one with =sucoder -T <target> cert=, which prompts for your BRC PIN + one-time code, POSTs them to the MSM CA, and writes the cert here (the CA caps a cert at 12h). You rarely need to run it explicitly: when the cert is missing or expired, an *interactive* =sucoder= command offers to mint a fresh one before it connects (so one OTP buys a 12h window instead of an OTP per connection); a non-interactive/agent run just falls back to ssh's own prompt. The cert is applied only to the gateway --- the login/DTN hops keep their publickey auth through the mux. =sucoder -T <target> tunnel doctor= reports the cert's validity. (=scripts/brc-cert.sh= still works as a standalone minter if you prefer.) - =x11= --- trusted X11 forwarding (=ssh -Y= semantics) on the *interactive* hops (=collaborate= launch, =attach=), so programs started inside the remote session can open X windows on your local display. *On by default* whenever the local session has a =DISPLAY= (XQuartz on macOS); silently skipped when it doesn't, so headless and cron runs stay quiet. Disable with =x11: false= on the target or =sucoder --no-x11= per invocation. The remote node needs =xauth= and =X11Forwarding yes= in its =sshd=. On a confined (sbatch) target, attach reaches the login node over ssh and steps into the job with =srun --x11=, which needs the cluster's Slurm X11 support --- so that srun flag is added only when x11 was *explicitly* enabled (=x11: true= or =--x11=), never by the default, keeping =attach= safe on clusters without it. Caveat: the forward lives and dies with the attaching ssh connection. The tmux session created at launch inherits a working =DISPLAY=; after a detach/re-attach, *new* tmux windows pick up the fresh =DISPLAY= automatically (tmux's default =update-environment= includes it), but processes already running keep the stale one --- refresh a shell with =eval "$(tmux show-environment -s DISPLAY)"=. *** Using a target from Emacs / TRAMP =sucoder -T <target> tunnel up= writes three =Host= aliases into your =~/.ssh/config= --- =<target>-gw= (gateway), =<target>-ln= (login node), and =<target>-dtn= (DTN) --- each carrying =HostName=, =ProxyJump=, =User=, and a =ControlPath= that points at sucoder's warm socket. TRAMP can ride those directly, so opening a remote file re-uses the warm =ControlMaster= with no fresh PIN/OTP (and with a minted =cert_file=, the first hop is OTP-free too). One Emacs setting is required, because TRAMP otherwise injects its *own* =ControlMaster= options that override the alias's =ControlPath= and open a separate, un-warmed connection: #+begin_src emacs-lisp (setq tramp-use-ssh-controlmaster-options nil) ; defer to ~/.ssh/config #+end_src Then warm the tunnel and open files against the =-ln= alias (your NFS =$HOME= / =~/mirrors= live on the login node, shared with every compute node): #+begin_src shell sucoder -T savio-node tunnel up # warms sockets + writes aliases ssh savio-node-ln hostname # sanity check: silent == TRAMP will be too #+end_src #+begin_example C-x C-f /ssh:savio-node-ln:~/mirrors/MyProject/ #+end_example Notes: - No alias is written for a *compute* node; multi-hop through the login node if you need one: =/ssh:savio-node-ln|ssh:n0032.savio2:~/...=. But since =$HOME=/=~/mirrors= is shared NFS, editing via =-ln= already reaches the files a compute job sees. - If the login-node pin drifted (the alias's =HostName= no longer matches the warm node), re-run =tunnel up= before pointing TRAMP at =-ln=. *** Usage Use =--target= (=-T=) as a global option before the subcommand: #+begin_src shell sucoder -T savio collaborate # launch agent on cluster sucoder -T savio worktrees --watch 30 # monitor from local sucoder -T savio attach # reconnect via tmux sucoder -T savio tunnel forward 8888 # compute-node web app -> http://localhost:8888/ sucoder -T savio tunnel forwards # list forwards + master liveness sucoder -T savio tunnel forward 8888 --cancel #+end_src =tunnel forward= defaults the node to the one your collaborate session is on (override with =--node n0030.savio4=) and terminates the forward *on the node*, so services bound to =127.0.0.1= work. Open the app's printed URL with the host replaced by =localhost=; keep the port and any =?token=...= query intact (use =--local-port= only if the local port is already taken). =scripts/claude-science.sh= is a worked end-to-end example of the whole pattern. It gives the web app its *own* dedicated thin SLURM slice (a job named =csci-<target>=, params read from the target's =slurm:= block) so it never has to share --- and fight --- a compute node with heavy jobs. =up= warms the tunnels, reuses-or-allocates that slice (held open in a detached login-node =tmux=), starts the app on it, forwards its port, and opens the URL locally; =url= re-fetches a fresh URL from the *running* app, repairs the forward, and reopens it; =down= releases the slice (cancels the job --- queued or running --- kills its tmux, removes the forward). =up= / =url= are idempotent, and that includes the queue: a slice still PENDING when the scheduling ceiling (=CS_SCHED_WAIT=, default 300s) expires *stays queued*, and a later =up= attaches to it instead of resubmitting, so the job keeps its accrued priority. A submission SLURM rejects outright is reported with the =srun= client's captured stderr rather than a blind timeout. The same discipline covers the app itself: a =serve= process that exists but has not yet answered (e.g. a slow cold start over NFS) is waited on, never double-started --- two servers would mean two writers on one SQLite DB --- and the timeout diagnostics say whether the process is still alive (raise =WAIT_SECS=) or died (with a foreground repro command). See the header comment for the =TARGET= / =NODE= / =APP_BIN= / =OPENER= / =WAIT_SECS= and =CS_*= slice overrides. =scripts/savio-run= generalizes the *dispatch-and-collect* half of that pattern for one-off compute: from your laptop it warms the tunnels and runs a command on a compute node, streaming the output back. SLURM resources default to the target's =slurm:= block (flags override). #+begin_src shell savio-run -- python3 -c 'print(2**10)' # srun: blocks, streams stdout back savio-run -c 8 --mem 32G -- ./crunch.sh # override resources id=$(savio-run --batch -t 2:00:00 -- ./long.sh) # sbatch: prints JOBID, returns savio-run --fetch "$id" # state + stdout/stderr when done savio-run --status # squeue/sacct for your jobs #+end_src Batch jobs write to =~/.savio-run/<jobid>.{out,err}= on the target; a lone quoted argument is run as a shell string (pipes/=&&= work), multiple args are treated as argv. *** What happens under the hood 1. *Authenticate once* --- =sucoder= establishes an SSH ControlMaster connection to the gateway. For clusters with OTP/two-factor auth, this is the only password prompt; all subsequent SSH commands multiplex through the socket. 2. *Pin a login node* --- SSH to the gateway, run =hostname=, store the result (e.g., =ln002.brc=) in =~/.sucoder/sessions/<mirror>.yaml=. Subsequent connections jump to the same node via =-J=. For SLURM targets --- where the login node is only a routing hop, not where the agent lives --- a later command first reconciles this pin with the node =tunnel up= keeps warm, and re-pins to a healthy node off the gateway if the recorded one is wedged, so a stale pin can't strand you on a dead login node while a live tunnel sits idle. 3. *Open a tunnel* --- A local port forward to the data transfer node for fast git transport. 4. *Push canonical state* --- =git push= through the tunnel to the remote mirror. 5. *Launch the agent* --- SSH to the pinned login node, =cd= into the remote mirror, start the agent inside a =tmux= session. 6. *Add a git remote* --- A remote named after the target (e.g., =savio=) is added to the canonical repo so the human can =git fetch savio= to pull back agent work. *** One job per mirror per target A confined launch never allocates a second job for a mirror that already has a live one. It checks twice, because the two checks fail in different ways: 1. the session record, =~/.sucoder/sessions/<mirror>--<target>.yaml=, which holds the job id; 2. failing that, =squeue --me --name=sucoder-<token>= --- the scheduler. The second exists because the first holds /one/ id and reads as blank when the file is missing, unreadable, or was written under a different =-T= spelling. Blank used to mean "no job", and the launch submitted a second =sbatch= over a live one. Slurm has known the answer all along: every launch is submitted =--job-name=sucoder-<token>=. When the scheduler finds a job the record lost, =collaborate= adopts it --- attaches, and writes the id back so =attach=, =release= and =renew= can reach it again --- rather than allocating beside it. A job for the same mirror on a /different/ target does not block the launch: one mirror on two targets is a deliberate configuration. It is reported and the new job is allocated. Targets are told apart by partition/account/qos, the same signature =sucoder sessions= groups by, so a target that pins none of the three claims nothing. A failed =squeue= does not block a launch either. This check is a safety net over the record, not a gate; the record-keyed check is the one that already refuses to resubmit on an unknown answer. =salloc= targets carry the same =--job-name=. Nothing reuses by name on that path yet --- it reuses via the record and adopts by node --- but without the name an allocation is invisible to =sucoder sessions=, which filters on the =sucoder-= prefix. *** Listing your sessions =sucoder sessions= lists every SuCoder job across all configured clusters, with its state, time left, node, and what the tmux session is actually running. It needs no =-T=: it reports on everything. #+begin_src shell sucoder sessions # probe each job's tmux pane (a few seconds) sucoder sessions --fast # skip the probe; job ids and clocks only sucoder sessions --no-login-nodes # skip the login-node sweep only #+end_src #+begin_example carleton-htc savio4_htc / co_carleton / carleton_htc4_normal LSMS_Library 39002769 RUNNING 11-00:13 n0043.savio4 claude 6m SuCoder 39025067 RUNNING 11-22:24 n0029.savio4 bash 2h ! agent exited savio-node savio3 / fc_jevons K-Aggregators 38991234 RUNNING 3-04:11 n0142.savio3 claude 4m ! no session record stale session records (the job is gone; the record is not): CFEDemands--savio-node -> 34739256 sucoder -T savio-node release CFEDemands MetricsMiscellany--carleton-htc -> 38999103 mirror not configured; edit the record by hand #+end_example Targets sharing a gateway share a =$HOME= and a scheduler, so they are queried once between them and sorted out afterwards by partition/account/qos. That one query runs on a login node the session records already pin when there is one, falling back to the gateway --- both answer =squeue= identically, and the login node is the cheaper and less contended hop (see [[*What a remote command costs][What a remote command costs]]). When the same host also has to be swept for tmux sessions, both questions travel in one script, because that master carries one session at a time. The listing is enumerated from =squeue=, not from the session records under =~/.sucoder/sessions/=: those hold one job id per (mirror, target), so a record-driven listing would show exactly the jobs that are /not/ the problem. Two flags matter: - =! no session record= :: nothing local points at this job, so =attach=, =release= and =renew= --- which all resolve through that record --- cannot reach it. Only =scancel= can. It happens when a record is overwritten by a later launch, is unreadable, or was written under a different =-T= spelling. - =! agent exited= :: the allocation is alive but its tmux pane is a bare shell. A confined job's window ends in =exec bash -l= so it survives a clean =/exit=, and the batch body's keeper polls only whether the session exists --- so the job goes on holding its slice for the rest of its =--time= with nobody home. Nothing else reports this. The trailing sweep is the reverse direction: session records naming a job the scheduler has forgotten. Each row carries the invocation that clears it, because =release= resolves the mirror through =config.mirrors= and the target through =-T=, and so cannot reach a record whose mirror has since been dropped from the configuration --- a common case on a long-lived launcher host. For those, =scripts/clear-stale-sessions.py= does in bulk what =release= does to one record: it nulls =slurm_job_id= and =compute_node=, leaves every other key (=login_node= included) alone, and never deletes a file. It is a dry run unless given =--apply=, which first writes a timestamped tar backup beside the session directory, and it aborts rather than touch anything if =squeue= cannot be reached on every configured cluster --- "the query failed" must never be read as "the job is gone". Targets with no =slurm:= block hold no allocation, so =squeue= says nothing about them --- but =collaborate= still launches a tmux session on the *login node*, with the same =exec bash -l= tail. A job's allocation ends at its =--time= and takes the tmux server with it; a login node has no walltime, so an abandoned session there persists until the node reboots, on shared infrastructure. Those are listed too: #+begin_example savio no scheduler (login-node sessions) sucoder-SuCoder ln002.brc bash ! agent exited sucoder-LSMS_Library ln001.brc claude hhsurveys no scheduler; no login-node sessions #+end_example Enumerated from =tmux list-sessions= filtered on the =sucoder-= prefix --- the login-node analogue of filtering =squeue --me= on =sucoder-<token>=, so a session whose local record was lost is still found. The gateway round-robins across login nodes, and which one an account reaches depends on its class (condo and FCA accounts land on different nodes), so the probe asks each node the session records pin *by name* as well as the gateway itself; reconnecting to the gateway alone could never see them all. =--fast= skips this too, and says =not inspected= rather than implying absence; =--no-login-nodes= skips *only* this sweep, keeping the per-job pane probe. The nodes are asked concurrently, so the wall clock is the slowest node rather than the sum --- and the sweep is started before the =squeue= queries rather than after them, since a login-node session has no allocation and so nothing in it depends on the scheduler. On a cold start (no live ControlMasters --- after minting a certificate, say) the gateway is authenticated first and alone, because it is the only hop that can prompt, and because several connections racing to authenticate the same gateway earns =Too many authentication failures= from =sshd=; the nodes behind it are then brought up together. The =! agent exited= check is why the probe exists, and why it is on by default: it reads the pane's child process, not =#{pane_current_command}=, which reports the pane's /shell/ and would call every working agent dead. =--fast= skips it and says so. Read-only throughout --- nothing submits, cancels or writes. *** Listing and retiring WIP snapshots =sucoder snapshots= lists every WIP snapshot ref on every configured cluster's mirrors, with what the scheduler's accounting says about the job that took it, and whether it is past retention. Like =sessions= it needs no =-T= and is read-only; =--retire= deletes what the listing marks. #+begin_src shell sucoder snapshots # list; nothing is deleted sucoder snapshots --retire # delete the ones past retention sucoder snapshots --recovery-window 3 # days after a job's end; default 7 #+end_src #+begin_example hpc.brc.berkeley.edu (carleton-htc, savio-htc, savio-node) LSMS_Library wip-job/LSMS_Library/39123514 0h 39123514 running kept: job still going LSMS_Library wip/LSMS_Library 3d 39002769 cancelled 3d ago kept: within the 7d window SuCoder wip-job/SuCoder/38710868 41d 38710868 timeout 40d ago retire: job ended 40d ago, past the 7d window 3 snapshot(s); 1 past retention. --retire deletes them. #+end_example The prepare script already retires ended jobs' snapshots, but only when a job is launched against that mirror, so a mirror no longer in use kept its refs forever --- and each ref pins a snapshot tree that =gc= can then never collect. This is the sweep that needs no launch and no allocation. Mirrors are found on the cluster's filesystem, not in the config, so a mirror dropped from the configuration (the one =release= can no longer reach) is swept too. A snapshot is judged by /when its job ended/, from =sacct=, never by the age of the ref alone. The snapshotter leaves the ref untouched while the tree is unchanged, and a confined window survives the agent's exit, so a ref can sit frozen for days under a job that is still alive; any age-only threshold shorter than the allocation would delete a live job's only snapshot. Where accounting cannot place a job (aged out, or no =sacct=), ref age is used instead against the cluster's longest =slurm.time= plus the window, which no live job can exceed; where a target has no finite =--time= even that has no safe bound, and nothing is deleted. Every deletion is guarded by the hash the listing saw, so a snapshot rewritten in between is left alone. *** Messaging a running session Two agents working one mirror --- one in a job, one on the laptop --- have no channel to each other; on 2026-09-21 the only link was a human copying between terminals, and a wrong diagnosis was overturned only because of it. =sucoder message= types a line into a running agent's session from wherever you are. #+begin_src shell sucoder message LSMS_Library "ln001's Lustre client is wedged; read from dtn, stop repairing" sucoder message --all "ln001 is wedged; stop repairing" # every live session, every cluster sucoder message LSMS_Library --dry-run "..." # who would get it, and the exact line sucoder -T carleton-htc message LSMS_Library "..." # one target, when a mirror runs on two #+end_src For a terminal agent stdin is the API and tmux owns it, so the line is =tmux send-keys= into the session's pane --- through =srun --overlap= for a job, where the tmux server lives inside the allocation, exactly as the pane probe reaches it. Recipients are the sessions =sucoder sessions= shows, enumerated from the scheduler and the login nodes, so a job whose record was lost can still be reached. The text arrives framed, =[message from you@laptop via sucoder, <time>] ...=, because it lands with the same authority as you typing; the system prompt's bulletin convention tells the agent to treat it as a peer's report to check, not an instruction from the human. A session whose pane is a bare shell --- an agent that has exited, leaving the =exec bash -l= behind --- is never sent to, since Enter there runs the text as a command; one whose pane could not be probed is skipped unless =--force=. The message is one line: a newline sent literally is an Enter, and would submit half of it. A reply cannot be pushed back --- a cluster has no route to a laptop behind NAT --- so it is pulled. The default system prompt asks an agent to answer on its own screen with a line =REPLY <id>: ...=, naming the id the message carried, and =--wait= reads the pane after the send for that line and prints it: #+begin_src shell sucoder -T carleton-htc message SuCoder --wait 120 "is n0032's Lustre healthy from where you sit?" sucoder -T carleton-htc peek SuCoder -n 60 # the pane's last rows, any time #+end_src =peek= is the read half on its own, for looking at a session without sending anything; unlike =message= it will happily show a pane that is a bare shell. Whether an agent acts on a message is still up to it. For a Codex remote-control service, =peek= shows server logs and =message= refuses terminal delivery, including with =--force=. Read and steer its chats through the connected Codex client. The cheaper half of the same fix is the bulletin, =~/.sucoder/bulletin.org= on the host the mirror lives on, which the default system prompt tells agents to read before concluding that infrastructure is broken and to write to before attempting a repair. *** Choosing a login node The login node is pinned once, from a round-robin =ssh <gateway> hostname=, and stored in the session record; every later command reuses it. When that node goes bad --- BRC's login nodes can lose a Lustre route while the node itself stays up and answers ssh --- =--login-node= steers off it: #+begin_src shell sucoder --login-node ln002.brc -T savio collaborate #+end_src It overrides the record's pin and writes the new one back, so subsequent commands agree. The sibling of =--node=, which selects a *compute* node. For a SLURM target the login node is only a routing hop (the work lives on =compute_node= + job id), so repointing is free. Without a scheduler the agent's tmux lives *on* the login node, so changing the pin abandons any session on the old one --- =attach= will look at the new node and not find it. sucoder warns when that is what you are doing; =sucoder sessions= still lists the session on the old node. *** Inspecting node availability Before reserving a node (or targeting one with =--node= / =--local-disk=), =sucoder nodes= shows which nodes in a partition are free, reusing the target's warm gateway ControlMaster so it costs no OTP when a tunnel is already up: #+begin_src shell sucoder -T savio-node nodes # partition defaults to slurm.partition (savio3) sucoder -T savio-htc nodes # -> savio4_htc sucoder -T savio-node nodes savio3_gpu # positional arg overrides the partition #+end_src It prints one row per node --- state, CPUs (Allocated/Idle/Other/Total) and load --- followed by the drained/down nodes and their reasons (=sinfo -R=). The query is read-only: no allocation, no session changes. Caveat: =sinfo= reports SLURM state, not Lustre health. A node can read =idle= while its filesystem mount is wedged, so weigh the drain reasons and any anomalous load on an otherwise-idle node --- the query cannot promise a node's filesystem is healthy. *** Node pinning On SLURM clusters, =sucoder= allocates a compute node for the agent session. The node is stored in =~/.sucoder/sessions/<mirror>.yaml= so that subsequent commands (=attach=, =pull=, =worktrees=) reconnect to the same node. To request a specific node (e.g., to recover work on local disk from a previous session): #+begin_src shell sucoder -T savio collaborate --node n0047.savio3 sucoder -T savio attach --node n0047.savio3 #+end_src The =--node= option sets a preferred node for the SLURM allocation. If that node is unavailable, SLURM may fall back to another node. **** Sharing a reserved node An exclusive partition hands out the whole node, but a single agent session often leaves it underused. If you already hold a live allocation on a node, point a *second* mirror at it with =--node= and =sucoder= adopts the existing job instead of reserving another: #+begin_src shell sucoder -T savio collaborate Foo # reserves n0020.savio3 sucoder -T savio collaborate Bar --node n0020.savio3 # shares it #+end_src Each mirror gets its own tmux session (=sucoder-Foo=, =sucoder-Bar=) and its own clone, so the two agents run side by side on one node. Adoption only ever attaches to *your own* RUNNING jobs on that node. Because the sessions share one SLURM job, =sucoder release= is job-aware: releasing a mirror while siblings still hold the job *detaches* it (kills that mirror's tmux session) but keeps the allocation alive. Only the last holder's =release= runs =scancel=. *** Local disk By default the agent works directly in the remote mirror on the shared filesystem. On clusters whose compute nodes have fast local scratch, =slurm.local_disk= (or =--local-disk= on the command line, with =--local-disk-root PATH= to pick a root other than =/local=) moves the agent's /working tree and caches/ there while the shared mirror stays the durable repository: #+begin_src shell sucoder --local-disk -T carleton-htc collaborate SuCoder sucoder --local-disk-root /scratch/local -T carleton-htc collaborate SuCoder # implies --local-disk #+end_src #+begin_src yaml targets: carleton-htc: gateway: hpc.brc.berkeley.edu transfer_host: dtn.brc.berkeley.edu mirror_root: ~/mirrors slurm: partition: savio4_htc account: co_carleton confined: true local_disk: true # or a path; true means /local #+end_src This is /tiering/ (design and measurements in [[file:docs/local-disk-tiering.org][docs/local-disk-tiering.org]]), and it works the same way on =confined= (=sbatch=) and =salloc= targets: - the launch clones =~/mirrors/<name>= to =/local/job<ID>/mirrors/<name>= and starts the agent there, with =UV_CACHE_DIR=, =PIP_CACHE_DIR=, =npm_config_cache=, and =TMPDIR= under =/local/job<ID>/=; - a =post-commit= hook publishes every commit to the shared mirror the moment it exists, so =sucoder pull= and the login-node view see it immediately; the hook never forces, and a push refused because the laptop pushed first is reported for the agent to =git pull --ff-only=; - the deadline watchdog snapshots uncommitted work to =refs/sucoder/wip-job/<name>/<jobid>= on the shared mirror every =wip_snapshot_minutes= and restores it on the next launch when the branch tip has not moved and the job that took it has ended; a clean =sucoder release= retires the job's own snapshot (=--keep-wip= keeps it), and the next launch retires those of other ended jobs; - SLURM wipes =/local/job<ID>/= when the job ends; nothing is orphaned and =--node= pinning is unnecessary. The shared mirror becomes a mailbox: do not edit its working tree by hand (an uncommitted tracked change there makes =updateInstead= refuse every publish from the clone; the prepare step warns loudly if it finds one). Ignored files (=.venv=, =node_modules=) are rebuilt each job; committed work is durable instantly, dirty work to within =wip_snapshot_minutes=. The agent is told all of this in a =WORKSPACE (local-disk tiering)= block of its prelude, including the command that shows the last snapshot. Earlier releases put the /whole/ mirror at =/local/mirrors= on one node. That layout is retired: a session that still records it gets a warning naming the node, so any commits made there can be pushed by hand before =sucoder pull= is trusted. *** Shared partitions (fractional allocations) Exclusive partitions (e.g. =savio3=) hand out an entire compute node. Shared/HTC partitions (e.g. =savio4_htc=) expect each job to declare how much of the node it wants; without =cpus_per_task= and =mem= you get the partition's minimal default (typically 1 core / a few hundred MB). #+begin_src yaml targets: savio-htc: gateway: hpc.brc.berkeley.edu transfer_host: dtn.brc.berkeley.edu mirror_root: ~/mirrors control_persist: 24h slurm: partition: savio4_htc account: fc_jevons qos: savio_normal cpus_per_task: 4 # cores mem: 16G # per-job memory time: "24:00:00" #+end_src =cpus_per_task= must be a positive integer; =mem= is passed verbatim to =salloc --mem=, so use SLURM's own units (=16G=, =4000M=, etc.). Both fields are optional --- omitting them keeps the partition default, which is the right behaviour for whole-node partitions. *** Deadline watchdog and WIP snapshots Every SLURM-backed session (=salloc= or =confined= =sbatch=) starts a small watchdog on the compute node. It warns at 30, 15, and 5 minutes before the allocation's =--time= via =tmux display-message= and by writing a warning under =/tmp/sucoder-<uid>/timers/<scope>/<node>-<job>/= (the un-suffixed =slurm-deadline.warn= and per-mirror =slurm-deadline-<mirror>.warn= are also written in =$HOME/.cache/sucoder/= for older prompts; these compatibility copies show the last writer). It never cancels the job; that stays with =sucoder release=. The scope is a digest of the mirror and target names. The timer uses Linux =/proc= and =flock= to verify/reuse a live watchdog for that allocation; repeated startup does not restart it or reset warning state. Startup reports whether it started or reused a timer. Failures explicitly report that deadline warnings and periodic snapshots are unavailable; diagnostics remain in =timer.log= beside its =owner= and =status= files. Timers from the older implementation are not killed by a broad process-name match; they exit with their old allocation. Locks, owner/readiness records, threshold markers, and timer logs live in a private, owner-verified mode-700 directory on node-local =/tmp=. This transient state only needs to survive for the watchdog lifetime. =TMPDIR= is deliberately not used: it may point at shared storage. Staged scripts and compatibility warnings stay on NFS; NFS need not support =flock=. Working clones may use =/local/job<jobid>/=, with snapshots pushed to durable storage. Lustre is not used for the timer's frequent small-file operations. The same script snapshots the mirror's dirty working tree (tracked and untracked files, not ignored ones) to =refs/sucoder/wip-job/<mirror>/<jobid>= on the mirror's =origin= at each warning and every =wip_snapshot_minutes= (default 10; =0= disables the periodic run). A mirror with no =origin= is never snapshotted, so on a shared-filesystem mirror this is a no-op today; it becomes live with the local-disk tiering described in [[file:docs/local-disk-tiering.org][docs/local-disk-tiering.org]]. The ref carries the job id because one mirror can be running in two jobs at once, on two nodes, with two different uncommitted trees. A relaunch restores this job's own snapshot if it has one, otherwise the newest whose job has ended -- never one belonging to a job still running, which would silently import another node's work. Snapshotted repositories should =.gitignore= their generated outputs: every snapshot that changes the tree writes a new commit to the shared mirror, and the superseded one becomes unreachable there until it is collected. #+begin_src yaml slurm: partition: savio4_htc account: co_carleton confined: true wip_snapshot_minutes: 10 #+end_src *** Reconnecting If the SSH connection drops, the tmux session on the login node survives. Reconnect with: #+begin_src shell sucoder -T savio attach #+end_src If the ControlMaster socket expires (default 12 hours), =sucoder= detects the stale socket and re-prompts for authentication. **** Orphaned sessions (SLURM) If you lose the SSH connection but the tmux server is still running on the compute node --- and =sucoder attach= fails (e.g. the session JSON is missing, the compute-node hostname wasn't recorded, or the site blocks direct SSH login -> compute) --- attach via the allocation itself: #+begin_src shell sucoder -T savio attach --via-srun #+end_src This stops at the login node and joins the running job with =srun --jobid=<JOB> --overlap --pty=, which lands you on the compute node *inside the job's cgroup*. =--overlap= is the key flag: it lets you add a step to an existing allocation without trying to consume extra resources. If =sucoder= itself isn't available (different machine, no install), the manual recipe is: #+begin_src shell ssh <gateway> ssh ln001 # or whichever login node squeue -u $USER # find your JOBID srun --jobid=<JOBID> --overlap --pty bash -l tmux attach -t sucoder-<mirror> # or just: tmux a #+end_src *** Persistent sessions (auto-renew) On a condo QOS with no wall-clock cap (e.g. =carleton_htc4_normal= on =savio4_htc=), a session can outlive any single allocation. =sucoder renew= watches the current job and, as it nears its courtesy =--time= (or if it vanishes at a maintenance reboot), re-allocates a fresh node and relaunches the agent in a *detached* tmux --- so the presence survives turnover hands-off. #+begin_src shell sucoder -T carleton-htc collaborate # start the session sucoder -T carleton-htc renew # another terminal: keep it alive #+end_src The loop is conservative: a transient probe failure (SSH blip) never triggers a relaunch --- only a *successful* probe reporting the job gone/terminal does. Tunables: =--drain-minutes= (lead time before =--time=, default 20), =--poll-interval= (seconds, default 60), =--checkpoint-grace= (seconds to let the agent commit/push after the drain nudge). =--once= runs a single probe/act cycle, for cron-driven renewal. Before each turnover the loop writes =$HOME/.cache/sucoder/renew-requested= on the compute node; the agent (per its =system_prompt_extra=) commits, pushes, and writes a handoff note, then rehydrates from that note after relaunch. Stop renewal with Ctrl-C (the allocation is left running) or =sucoder release= (frees it). For Codex remote control, renewal preserves the original service command and model, including =--agent-command= and =--model= overrides. The prelude and credential environment are rebuilt for the replacement allocation. The checkpoint sentinel is not a chat message: checkpoint active work through the connected Codex client before turnover. Renewal restarts the server; it does not automatically resume its chats. See [[file:docs/persistent-presence.org][docs/persistent-presence.org]] for the design and the open cgroup-confinement caveat. *** SSH robustness =sucoder= handles several common failure modes transparently: - *Stale ControlMaster sockets* are detected and recycled automatically. - *DTN timeouts* --- if the data transfer node is unreachable, =sucoder= falls back to pushing via the compute node. - *Debug mode changes* --- switching =--debug-ssh= on or off automatically cycles the ControlMaster socket so the new setting takes effect. - *Busy masters* --- a =Session open refused by peer= is a /busy/ signal, not a dead connection (see below); it is waited out and retried, never re-authenticated. *** What a remote command costs Two site facts shape every remote command, and both were measured on =hpc.brc.berkeley.edu= (2026-09-20): 1. *A session open costs ~8s*, over an already-warm ControlMaster, before the remote command runs at all --- ~4.5s on a login node, and the same for an =sftp= subsystem open, so it is the session setup itself and not your login shell. Multiplexing saves the /authentication/, not the time. 2. *A ControlMaster carries one session channel at a time* (=MaxSessions 1=). A second simultaneous command on the same host is not queued but *refused*, and on a refusal =ssh= falls back to dialling the host directly --- a fresh authentication, which on this gateway can earn =Too many authentication failures=. Port-forward channels are exempt, which is why reaching several hosts /through/ one gateway concurrently is fine while running two commands /on/ it is not. So: connections are verified once per run rather than once per command; hosts are asked one question at a time, with the answers to two questions fused into one script where they share a host; a refusal is waited out rather than reconnected; and a command that can run on a login node does, because that hop is cheaper and is not the one every jumped connection already contends for. Because the waiting is unavoidable, it is at least visible: every remote round trip announces itself on *stderr* before it is made, naming the host and what it is about to run. #+begin_example · ln001.brc squeue --me --noheader -o '%i|%j|%P|%a|%q|%T|%L|%N' · hpc.brc.berkeley.edu tmux list-sessions -F "#{session_name}" 2>/dev/null ... #+end_example Durations go to the log (=~/.sucoder/logs/=), and =-v= puts them on the console along with the connection machinery's own reasoning --- which probe timed out, why a master was declared dead, whether a refusal was read as busy. ** Human: Review and Integrate Agent Work *** Local mirrors Fetch the agent branch back into the canonical repository: #+begin_src shell git fetch coder # uses the coder remote added by prepare-canonical git log --oneline coder/coder/add-metrics-20251107164028 -5 git checkout -b review/add-metrics FETCH_HEAD # ... review, test, iterate ... git checkout main && git merge --ff-only review/add-metrics #+end_src Or pull directly: #+begin_src shell git pull --ff-only /var/tmp/coder-mirrors/project coder/add-metrics-20251107164028 #+end_src *** Remote mirrors When using =--target=, =sucoder= automatically adds a git remote named after the target. Fetch agent work back over SSH: #+begin_src shell git fetch savio # fetches from brc.berkeley.edu:~/mirrors/project git log --oneline savio/master -5 git merge --ff-only savio/master #+end_src The canonical repository remains unwritable by the agent throughout both local and remote flows. * Configuration Configuration is stored in YAML (default path is =~/.sucoder/config.yaml= for the human operator). Grant the agent group read access to this file while keeping it group-writable disabled. #+begin_src yaml human_user: ligon agent_user: coder agent_group: coder mirror_root: /var/tmp/coder-mirrors log_dir: ~/.sucoder/logs skills: - ~/.sucoder/skills # optional global default for all mirrors (overridable per-mirror) system_prompt: ~/.sucoder/system_prompt.org # Named remote execution targets (optional). # Use sucoder -T <name> to select one. targets: savio: gateway: brc.berkeley.edu transfer_host: dtn.brc.berkeley.edu mirror_root: ~/mirrors control_persist: 7d # Mirrors are optional when using zero-config repo detection. mirrors: project: canonical_repo: ~/src/project.git mirror_name: project # Optional; omit to auto-detect the canonical repo's default branch # (origin/HEAD, then a local main, then master, then current HEAD). # default_base_branch: main task_branch_prefix: task branch_prefixes: human: ligon agent: coder # Optional: override or extend global skills for this mirror skills: - ~/.sucoder/skills/org-style - ~/.sucoder/skills/document-skill agent_launcher: # Harness executable and optional default model can be configured independently. command: [aider] model: openrouter/deepseek/deepseek-chat # Alternatives: [opencode], [goose, run, --interactive], or [kimi]. # Remote Codex service: [codex, remote-control] (start is normalized to foreground). # Optional: map generic intents to agent-specific flags needs_yolo: true writable_dirs: - "~" flags: skills: "--skills {path}" # Uses mirror skills or ~/.sucoder/skills by default #+end_src You can also maintain a global catalog at =~/.sucoder/skills/SKILLS.md=, which can point to additional skill files or directories using =file:= links or bullet-listed paths. Any catalog discovered this way is loaded automatically before launching the agent. The =agent_launcher.flags= mapping lets you translate generic intents (write access, writable directories, workdir, default flags, skills paths, model, context file) into harness-specific switches. If a skills flag template is provided, sucoder will pass any configured skills plus the default =~/.sucoder/skills= directory when it exists. ** Agent flag templates by CLI The table below documents the current default flag templates used when the agent type is detected. Templates can be overridden per-mirror or globally via =agent_launcher.flags=, with precedence: per-mirror > global > profile > UNKNOWN. | Harness | Model | Write access | Writable directory | Prompt content | Prompt file | MCP config | |----------+------------------+------------------------------------------------------+-------------------------------+----------------------+---------------+---------------------| | Claude | --model {model} | --dangerously-skip-permissions | --add-dir {path} | --system-prompt | (none) | --mcp-config {path} | | Codex | --model {model} | --sandbox danger-full-access --ask-for-approval never| (none) | trailing text | (none) | (none) | | Gemini | --model {model} | --yolo | --include-directories {path} | --prompt-interactive | (none) | (none) | | Aider | --model {model} | --yes-always | (none) | (none) | --read {path} | (none) | | OpenCode | --model {model} | --auto | (none) | --prompt | (none) | native config | | Goose | --model {model} | native policy | (none) | --text | (none) | native config | | Kimi | --model {model} | --auto | --add-dir {path} | (none) | --agent-file {path} | native config | The common =default_flag= template is ={flag}=. Native Agent Skills discovery means OpenCode, Goose, and Kimi do not need a =skills= flag. Goose launches as =goose run --interactive= so the injected =--text= context is processed before the interactive session begins. Kimi's generated custom-agent file includes =${base_prompt}= so its native tools and skills remain available. If you change any defaults, update =AGENT_PROFILES= in =sucoder/config.py= alongside the README so the documentation stays in sync. ** MCP servers (repo-specific tools) MCP (Model Context Protocol) servers give agents access to external services---GitHub, web pages, databases---via a standardised tool interface. Claude Code on the web injects GitHub tools automatically, but the CLI (=claude=) does not. Since =sucoder= launches agents via the CLI, MCP servers bridge the gap. *** How it works There are two independent sources of MCP configuration; Claude merges them automatically: 1. *Repo-shipped =.mcp.json=* --- lives in the repository root. Claude discovers it natively when launched in the mirror. This is the simplest approach: add the file, commit it, and every agent session gets those tools. 2. *Sucoder-config =mcp_servers=* --- defined in =~/.sucoder/config.yaml= at the global or per-mirror level. Sucoder generates a =.sucoder-mcp.json= file in the mirror and passes it via =--mcp-config=. Use this for infrastructure tools that are not part of any single repo. *** Shipped MCP servers This repository includes a =.mcp.json= with three servers: | Server | Purpose | |----------+------------------------------------------------------| | =github= | Read/write GitHub issues, PRs, CI status, comments | | =fetch= | Fetch web pages and URLs (documentation, references) | | =memory= | Persistent knowledge graph across agent sessions | The =github= server requires a =GITHUB_TOKEN= environment variable. Because =.mcp.json= contains this secret, the file is encrypted with =git-crypt=. *** Setting up git-crypt (first time) After cloning the repository, unlock the encrypted files: #+begin_src shell # Install git-crypt if needed brew install git-crypt # macOS sudo apt install git-crypt # Debian/Ubuntu # If you are initialising git-crypt for the first time in this repo: git-crypt init git-crypt add-gpg-user <your-gpg-key-id> # Edit .mcp.json to add your real GITHUB_TOKEN, then commit. # The file will be stored encrypted in the repo but readable in your checkout. # If git-crypt is already initialised, just unlock: git-crypt unlock #+end_src Collaborators who have been added via =git-crypt add-gpg-user= can unlock with their GPG key. The agent user (=coder=) works in a mirror cloned from an already-unlocked checkout, so the plaintext =.mcp.json= is available without additional setup. *** Configuring MCP servers via =config.yaml= For MCP servers that should be available across multiple repos (not shipped in any single repo), add them to the sucoder configuration: #+begin_src yaml # Global MCP servers (available to all mirrors) mcp_servers: my-database: command: npx args: ["-y", "@modelcontextprotocol/server-postgres"] env: DATABASE_URL: "postgresql://localhost/mydb" # Or per-mirror (overrides global for that mirror) mirrors: project: canonical_repo: ~/src/project.git mcp_servers: custom-tool: command: my-mcp-server args: ["--port", "8080"] #+end_src Sucoder writes these to =.sucoder-mcp.json= in the mirror (excluded from git) and passes =--mcp-config= to Claude. If the repo also ships a =.mcp.json=, Claude merges servers from both sources. ** Agent capability comparison The core =sucoder= workflow (mirror, sync, permissions, branch management) is fully agent-agnostic. The table below summarises where agents differ in features that =sucoder= can exploit. | Capability | Claude | Codex | Gemini | Aider | OpenCode | Goose | Kimi | |------------------------------+---------------------------------+--------------------------------+-------------------------+-----------------------+----------------+---------------------------+------------------------| | Native =--worktree= flag | Yes (=claude --worktree NAME=) | No | No | No | No | No | No | | Subagent worktree isolation | Yes (=isolation: worktree=) | No | No | No | No | No | No | | Session resume | Yes (=claude --resume=) | Yes | Yes | Chat history support | Yes | Yes (=--resume=) | Yes (=--session=) | | System prompt mechanism | =--system-prompt= | Trailing text | =--prompt-interactive= | Private =--read= file | =--prompt= | =run --interactive --text= | Private =--agent-file=| | Write-access flag | =--dangerously-skip-permissions=| =--sandbox danger-full-access= | =--yolo= | =--yes-always= | =--auto= | Native policy | =--auto= | | TTY requirement | Works with subprocess | Works with subprocess | Requires =exec= (TTY) | Works with subprocess | Subprocess | Subprocess | Subprocess | | Default harness | *Yes* | Supported | Supported | Supported | Supported | Supported | Supported | Agents without native =--worktree= support can still use git worktrees manually. The agent just needs to run =git worktree add= inside the mirror; =sucoder worktrees= will discover and display them regardless of which agent created them. Similarly, the remote execution support (=--target=, SSH tunnels, ControlMaster) is entirely agent-agnostic---any agent CLI can be launched over SSH. When adding a new harness, implement its profile in =AGENT_PROFILES= in =sucoder/config.py= and add a row to these tables. ** Which agent binary actually launches =sucoder= hands the agent name to =execvp= (or to =sudo -u <agent_user>=), so the binary is chosen by a PATH lookup---but the PATH that decides is the *agent user's login* PATH, not yours. Both launch paths resolve inside a shell that has sourced the agent's login profile, and a stock Debian =~/.profile= prepends =$HOME/.local/bin=. What such a shell does /not/ pick up is =~/.bashrc=, which returns early when non-interactive, so version managers wired up there---nvm and friends---never load. A shell you type =codex= into and a shell =sucoder= launches can therefore resolve to different installs, and the stale one wins silently. Before each local launch, =sucoder= asks that same shell (=bash -lc 'command -v <agent>'=, through the same =sudo= escalation the launch uses) and logs the path and =--version= it gets back: #+begin_example INFO Agent binary: codex -> /usr/local/bin/codex (codex-cli 0.147.0) #+end_example Resolving through the launch's own shell rather than =sucoder='s PATH is what makes the report trustworthy: it accounts for the agent user's login profile, and for =sudo='s =secure_path= when sudo is in play. If it finds a *newer* build of the same name elsewhere---anywhere on PATH, or under the agent user's =~/.local/bin= or =~/.nvm/versions/node/*/bin=---it warns that the resolved one is shadowing it. The usual fix is a symlink from a directory that comes earlier on the agent user's login PATH, or an =agent_launcher.nvm= block to pin resolution outright. The whole check is best-effort and contained: probes are bounded, and any failure inside it is logged at debug level and swallowed rather than allowed to abort a launch that would otherwise succeed. The check is skipped for remote launches (PATH belongs to the remote host), under =--dry-run= (it would spawn =--version= subprocesses), and when an =agent_launcher.nvm= block is configured, since that pins resolution deliberately and it happens inside the nvm shell. ** Tool versions on the target (launch preflight) The section above answers /which agent binary/ runs. This one answers a neighbouring question the shipped prompts and skills make unavoidable: the agent is told to use =gh=, =git=, =jq=, =rg= and =tmux=, and a session inherits whatever happens to be in the target's =$HOME=. Until now =sucoder= never said how old any of it was. Before every launch, =sucoder= runs one probe on the target (a generated bash script fed to =bash -l -s=, so it resolves under the same login PATH the agent will get) and logs what it found: #+begin_example INFO Tool versions on n0043.savio4: gh: gh version 2.67.0 (2025-02-11), floor 2.90.0 -- BELOW FLOOR [/global/home/users/u/bin/gh] git: git version 2.43.7, floor 2.34.0 [/usr/bin/git] jq: jq-1.6, floor 1.6 [/usr/bin/jq] rg: not found, floor 13.0.0 tmux: tmux 3.7c, floor 3.0 [/global/home/users/u/bin/tmux] WARNING Tool preflight: gh: gh version 2.67.0 (2025-02-11), floor 2.90.0 -- BELOW FLOOR [...] (this is a HOST tooling fact, not a property of the repository; sucoder does not manage these binaries and did not block the launch) #+end_example *Why this exists.* A target's =gh= was 2.67.0 (February 2025). Every =gh pr edit= and =gh issue view= failed with =GraphQL: Projects (classic) is being deprecated ... (repository.pullRequest.projectCards)=---an error naming a GitHub product sunset, emitted for a command carrying no project flags. An agent concluded /"gh pr edit is broken on this repo"/ and wrote that up as a repo-specific gotcha for others to inherit. The misattribution, not the outage, is what the preflight prevents: the version is in the log before the agent starts. It also makes visible, for the first time, the case where two targets carry different versions and a workflow verified on one silently fails on the other. *What it deliberately does not do.* It never gates---a stale tool must not stop a session starting, and nothing in the check can fail a launch, including its own bugs. It never touches the network: the floor is a constant in the config file, because a preflight that needs the internet fails on exactly the constrained targets it is meant to help. And it does not manage the tools; they live in the user's =$HOME=, shared across targets, and =sucoder= is not a package manager. Replacing a running, user-owned static binary is =mv= (which replaces the directory entry and leaves running processes on the old inode), never =cp= over the top---several targets may share one =$HOME= and another session may be mid-call. Auth and config are unaffected; for =gh= they live in =~/.config/gh/=. A version string it cannot read (=tmux 3.3a=, =jq-1.6= and =gh version 2.67.0 (2025-02-11)= are all handled, but the world is larger than five tools) is reported as unreadable---recorded, and neither a pass nor a warning. *** Floors and which tools are probed Both live in one place, =tool_preflight= in =config.yaml=: #+begin_src yaml tool_preflight: # enabled: false # skip the probe at launch; `sucoder doctor` still runs it floors: gh: "2.95.0" # override a shipped floor tmux: null # record the version, stop judging it shellcheck: "0.9.0" # add a tool (probed with --version) #+end_src The keys of =floors= are what gets probed, so one knob both adds a tool and sets its floor, and =null= silences a floor without losing the reading. Entries are merged over the shipped defaults (=gh 2.90.0=, =git 2.34.0=, =jq 1.6=, =rg 13.0.0=, =tmux 3.0=), so naming one tool does not quietly stop checking the rest. A floor =sucoder= cannot parse is rejected when the config loads, not at launch. *A floor is a staleness threshold, not a correctness boundary.* It says "old enough that a server-shaped failure should be suspected of being the client", which is the whole job; it does not certify that a passing version is free of any particular incompatibility---and for the incident above there is no version that would. [[https://github.com/cli/cli/issues/11983][cli/cli#11983]] is an /issue/, closed 2025-10-21 by its own reporter with no linked commit; duplicates kept arriving through March 2026 (#12476, #12640, #13069); the two PRs that would have fixed it (#13083, #13282) were both closed *unmerged*; and =projectCards= is still referenced across cli/cli today, =pkg/cmd/pr/edit/edit.go= included. What /is/ established is two readings: 2.67.0 failed on the affected target and 2.101.0 succeeded on it. The shipped =gh= floor sits between them---near enough to current to catch a genuinely stale client, far enough back that a fortnightly release does not re-fire it. All five defaults are conservative choices, not findings. *** =sucoder doctor= The same check on demand: #+begin_src sh sucoder doctor [<mirror>] #+end_src Unlike the launch-time preflight it *exits non-zero* when a tool is missing or below its floor (the same convention as =sucoder tunnel doctor=), and it prints the =mv=-not-=cp= note above. It refuses a non-confined SLURM target rather than quietly running =salloc=: a diagnostic must not bill a compute allocation. Run it without =-T= (or against a confined target) to probe the login node, or read the report the launcher already wrote to the session log. *Caveat worth knowing.* For a confined (=sbatch=) target the preflight runs on the login node at launch time, while the agent runs on a compute node. Tools under a shared =$HOME= (=gh=, =rg=, =jq=) are the same binary; a system =git= or =tmux= need not be. The probed hostname is printed in the report so that difference is visible where it is read. ** Pin agent Node version with nvm Some agents (e.g., Codex) bundle their own Node.js runtime. =sucoder= can wrap the launch so that the agent runs under a newer nvm-managed runtime instead (for example Node 22). 1. Install/activate the desired version for the agent user: #+begin_src shell sudo -u coder bash -lc 'export NVM_DIR=/home/coder/.nvm; . "$NVM_DIR/nvm.sh"; nvm install 22.11.0' sudo -u coder bash -lc 'export NVM_DIR=/home/coder/.nvm; . "$NVM_DIR/nvm.sh"; nvm alias default 22.11.0' #+end_src Replace the version string with the release you need (direct versions, =lts/*= aliases, etc.). 2. Update the mirror configuration so launches always source nvm before running the agent: #+begin_src yaml mirrors: project: canonical_repo: ~/src/project.git agent_launcher: command: - codex nvm: version: "22.11.0" # Any version/alias `nvm use` understands dir: /home/coder/.nvm # Optional; defaults to <agent home>/.nvm #+end_src When the =nvm= block is present, the launcher injects: - =export NVM_DIR=…; source "$NVM_DIR/nvm.sh"= - =nvm use <version>= (fails fast if the version is missing) - =exec <agent> …= so the rest of the CLI arguments (including any injected flags and context prelude) are preserved. If =dir= is omitted the helper assumes the agent’s home directory contains =~.nvm=. Any existing =agent_launcher.command= and =agent_launcher.env= settings continue to work. Workspace skills follow Anthropic’s Agent Skills Spec: - Each skill directory must contain a =SKILL.md= file with YAML frontmatter (`name`, `description`, and `license`; add `allowed-tools`/`metadata` as needed). If you copy a third-party skill, keep its license value and include the upstream notice in the directory. ** Unified skills catalog (local + curated) - Generate a tool-agnostic catalog of local skills (from the configured skills directory) plus the openai/skills curated list: #+begin_src shell python scripts/generate_unified_skills_catalog.py \ --output ../Skills/UNIFIED_SKILLS_CATALOG.md # choose any writable path #+end_src The script reads =../Skills/trusted_skill_sources.yaml= to determine which sources to include (local catalogs plus curated repos; defaults include =openai/skills= and =anthropics/skills=). If a target is not writable (for example =/home/ligon/.sucoder/skills=), use =--output= to pick a writable path. - The generated catalog shows entries per source and flags conflicts when the same skill name appears in multiple sources. - Any LLM can read the catalog and follow linked =SKILL.md= files. Curated entries are documentation-first; helper scripts are optional. - To install a curated skill manually, download =skills/.curated/<name>= from https://github.com/openai/skills and place it under your skills directory (e.g., =~/.sucoder/skills=). - The frontmatter `name` must match the directory name exactly. - Long-form references, scripts, and assets live under =references/=, =scripts/=, and =assets/= to keep the entrypoint concise. - The shared catalog (=~/.sucoder/skills/SKILLS.md=) in the skills repository lists every bundled skill so agents discover them automatically. - A reusable system prompt template (`default_system_prompt.org`) mirrors these expectations, including a reminder to confirm today's date at session start. - The upstream workflow is available in the =sucoder-skills= repository under =skill-creator/=; use it alongside =document-skill= when you need initialization or packaging scripts. - A starter configuration (`default_config.yaml`) points at the bundled skills directory and default prompt; copy or merge it into =~/.sucoder/config.yaml= as needed. Resource directories are summarized automatically when present: - =references/= — documentation to load on demand (each file includes a suggested load command). - =scripts/= — helper executables (never run automatically; humans can execute after review). - =assets/= — supporting files such as templates or images that may be referenced in outputs. * Skills Repository Skills are maintained in a separate repository (=sucoder-skills=) for security isolation. This separation prevents agents from modifying skills and tool code in the same commit, reducing the attack surface and simplifying security reviews. ** Why Skills Are Separate 1. *Security Isolation* — An agent working on tool code cannot simultaneously modify skills that influence agent behavior. 2. *Review Clarity* — Code reviews and content reviews use different security mindsets and can be performed separately. 3. *Blast Radius Reduction* — Malicious skill instructions cannot be hidden among tool code changes. ** Version Compatibility The tool enforces semantic versioning compatibility between itself and the skills repository: - Tool validates skills repository =VERSION= file on startup - Compatible: Tool requires =1.0.0=, skills has =1.x.x= ✓ - Incompatible: Tool requires =1.x.x=, skills has =2.0.0= ✗ Skills repository: https://github.com/ligon/sucoder-skills To update skills: #+begin_src shell cd ~/.sucoder/skills # Or ~/Projects/sucoder-skills git pull #+end_src To skip version checking (not recommended): #+begin_src shell export SUCODER_SKIP_SKILLS_VERSION=1 #+end_src ** Version Bumping Strategy *PATCH (1.0.0 → 1.0.1)*: Typo fixes, clarifications, examples *MINOR (1.0.0 → 1.1.0)*: New skills, new optional features, backward-compatible improvements *MAJOR (1.0.0 → 2.0.0)*: Breaking changes, requires tool update * Agent-Agnostic Project Instructions Projects can provide agent instructions and skills using agent-agnostic names. During mirror setup, =sucoder= creates symlinks so that Claude discovers them natively while other agents receive the content via prompt injection. | Canonical file | Symlink created | Purpose | |--------------------+--------------------------+--------------------------------| | =AGENT.md= / =.org= | =CLAUDE.md -> AGENT.md= | Project-level instructions | | =.skills/= | =.claude/skills -> .skills= | Project-level skill files | - Claude discovers =CLAUDE.md= and =.claude/skills= natively via the symlinks. - Non-Claude agents receive =AGENT.md= content in the system prompt. - One source of truth; no duplication across agent types. Symlinks are created during =ensure_clone= (both fresh clones and subsequent launches). If =CLAUDE.md= or =.claude/skills= already exist independently, =sucoder= leaves them in place and does not overwrite. * Agent Skills Tracking Agents may write skill files to =~coder/.claude/skills/= during sessions. These persist across sessions and influence future agent behavior. =sucoder= git-tracks this directory to maintain an audit trail: - The directory is initialized as a git repo on first use. - After each subprocess-mode session, any changes are auto-committed with a message referencing the mirror that produced them. - The commit history provides forensics: when a skill was added, which session produced it, and what it looked like before modification. * Auditing A dedicated compliance agent reviews changes made by the working agent---both skill files and code. The auditor runs as a separate Unix user (=auditor=) that is intentionally /not/ in the =coder= group, so write access to agent files is prevented by Unix permissions rather than prompt instructions. Read access is granted via world-readable bits (=o+r=). ** Setup #+begin_src shell make create-auditor-user # create auditor user (requires sudo) make auditor-perms # set o+r on paths the auditor needs # Install Claude CLI and authenticate (one-time): sudo -u auditor bash -c 'curl -fsSL https://claude.ai/install.sh | bash' sudo -u auditor claude login # Copy auditor prompts: cp default_auditor_prompt.org ~/.sucoder/auditor_prompt.org cp default_code_auditor_prompt.org ~/.sucoder/code_auditor_prompt.org #+end_src ** Usage The =--scope= flag controls what is audited: =skills= (default), =code=, or =all=. #+begin_src shell # Skills audit (default) — reviews ~coder/.claude/skills/ sucoder audit # diff review since last approved baseline sucoder audit --full # review all skills from scratch sucoder audit --approve # advance the baseline after review # Code audit — reviews mirror repo changes sucoder audit MIRROR --scope code # diff since last code baseline sucoder audit MIRROR --scope code --full # review all tracked code sucoder audit MIRROR --scope code --approve # Both at once sucoder audit MIRROR --scope all #+end_src ** Audit scopes *** Skills audit (=--scope skills=) Reviews agent-written skill files in =~coder/.claude/skills/=. Uses =~/.sucoder/auditor_prompt.org= as the system prompt. Tracks its baseline via =refs/audited= in the skills git repo. The auditor checks for instruction injection, scope creep, contradictions, sensitive data, and self-modification loops. *** Code audit (=--scope code=) Reviews code changes in a mirror repository. Requires a mirror name argument. Uses =~/.sucoder/code_auditor_prompt.org= as the system prompt. Tracks its baseline via =refs/audited-code= in the mirror. The auditor checks for: 1. Dependency injection (malicious or unexpected packages) 2. Credential leakage (hardcoded tokens, API keys, passwords) 3. Unsafe subprocess calls (=shell=True=, =eval=, =exec=) 4. Permission escalation (=chmod 777=, setuid, sudoers edits) 5. Unexpected network calls (outbound connections to unknown hosts) 6. Overly broad file operations (recursive deletes, out-of-tree writes) 7. Supply-chain risks (typosquatting, unusual version pinning) 8. Obfuscated code (base64-encoded strings, minified inline scripts) ** How it works 1. *Permissions check* --- Before examining content, the auditor verifies that all files are world-readable. Files the working agent has made unreadable (=chmod o-r=) are flagged immediately as suspicious. 2. *Audit mode* --- If no baseline exists (=refs/audited= for skills, =refs/audited-code= for code), a full review runs. Otherwise, only the diff since the last approved commit is examined. 3. *Auditor invocation* --- =claude -p= is run as the =auditor= user with a scope-specific system prompt. The skills auditor and code auditor use different prompts tuned to their respective concerns. 4. *Report* --- The auditor prints a summary. If concerns are found, each is listed with filename and severity. Pass =--approve= to advance the baseline; omit it to leave the baseline unchanged for human review. ** Auto-trigger after each session Auditing on demand only catches issues when the human remembers to run =sucoder audit=. To close that loop, opt in to a post-session auto trigger via =~/.sucoder/config.yaml=: #+begin_src yaml audit: auto_after_session: true # default false; explicit opt-in scope: all # skills | code | all #+end_src When enabled, =sucoder collaborate= (and other launch paths) call the audit subsystem after =_auto_commit_agent_skills= runs. The auditor operates exactly as if you had typed =sucoder audit MIRROR --scope <X>=, with three differences: - Reports are saved to =<log_dir>/audits/<mirror>-<kind>-<timestamp>.log= rather than printed to stdout (the session has ended; nobody is watching). When =log_dir= is unset, =~/.sucoder/logs/audits/= is used. - A one-line summary is logged: =INFO: Post-session code audit: no concerns (<path>)= when the report says so, =WARNING: Post-session code audit produced findings: <path>= otherwise. Watch =log_dir= or pipe sessions through =tee= if you want a durable record. - Failures are non-blocking. An expired auditor token, missing auditor user, network blip --- all log a warning and let session teardown succeed. Audit infrastructure must never make a successful agent session look failed. *Pre-flight*: the first auto-audit against a never-audited mirror has no =refs/audited= / =refs/audited-code= baseline, so it runs in *full* mode. That's the most expensive variant in LLM tokens. Run =sucoder audit MIRROR --scope all --approve= once after you're happy with the initial review; subsequent auto-audits then run in cheap diff mode (or skip entirely when there are no changes). ** Relationship with cq The =cq= knowledge commons and agent skills serve complementary roles: | Aspect | Skills | cq | |-----------+---------------------------------+---------------------------------| | Content | Curated instructions/procedures | Discovered learnings/pitfalls | | Lifecycle | Human-authored or promoted | Agent-proposed, human-confirmed | | Scope | Per-project or shared | Cross-project | The intended flow: agents propose learnings to =cq=; knowledge that proves itself over time gets promoted to a skill by a human. Agent-written skills in =~coder/.claude/skills/= are a fast path for immediate session context, subject to compliance audit. * Testing Run the automated tests with pytest. #+begin_src shell pytest #+end_src