Skip to content

Latest commit

 

History

History

README.md

@ai-sdlc/pipeline-cli

Shared core library for the AI-SDLC pipeline. Implements RFC-0012 — extracts Step 0-13 logic from ai-sdlc-plugin/agents/execute-orchestrator.md (now superseded by the inline commands/execute.md slash command body — AISDLC-98) and orchestrator/src/ into pure step functions exposed three ways:

  1. TypeScript library — import { executePipeline, ... } from '@ai-sdlc/pipeline-cli'
  2. CLI subcommands — ai-sdlc-pipeline <command> (yargs-driven, JSON on stdout)
  3. MCP tools — Phase 3 (AISDLC-100.3) wraps each step as an MCP tool from the plugin's MCP server

The package is published publicly as @ai-sdlc/pipeline-cli on npm (AISDLC-245.1 makes it publishable; the actual runtimeDependencies declaration in ai-sdlc-plugin/plugin.json lands in AISDLC-248.1's release PR once the first version is published, so adopters don't fail install on an unpublished version reference).

Why two tiers?

The pipeline ships in two tiers (RFC-0012 §2):

  • Tier 1 — slash command body. /ai-sdlc execute <task-id> runs in the main Claude Code session. The slash command body interleaves CLI subcommands (ai-sdlc-pipeline validate-task ..., ... compute-branch ...) with Agent tool calls for the LLM dispatch boundaries (Step 5b developer, Step 7b three reviewers in parallel). Subscription billing via Claude Code Max-20x. Operator-driven and interactive.

  • Tier 2 — executePipeline() composite. A single import + one async call drives Step 0-13 end-to-end. The two LLM dispatch boundaries go through an injected SubagentSpawner (subscription via claude --print, API key via @anthropic-ai/claude-code SDK, or MockSpawner for tests). Designed for unattended programmatic use: CLI invocation, GitHub Actions, webhooks, cron, and the existing pnpm watch flow once Phase 5 (AISDLC-100.5) migrates dogfood/src/watch.ts to call it.

Both tiers run the same Step 0-13 functions from this package, so behaviour is identical — only the LLM dispatch boundary differs. See docs/spawner.md for the SubagentSpawner deep-dive and docs/steps.md for the per-step contract reference.

Install

As part of the plugin (recommended)

If you've installed ai-sdlc-plugin in Claude Code, @ai-sdlc/pipeline-cli is pulled in automatically via the plugin's runtimeDependencies declaration. The Claude Code plugin runtime installs it under the plugin's own node_modules/:

<CLAUDE_PLUGIN_ROOT>/node_modules/@ai-sdlc/pipeline-cli/

All 16 bin/ entries are then resolvable as ${CLAUDE_PLUGIN_ROOT}/node_modules/@ai-sdlc/pipeline-cli/bin/<bin>.mjs:

ai-sdlc-pipeline.mjs            # Step 0-13 umbrella (RFC-0012)
cli-backlog-verify.mjs          # Backlog drift gate (CI)
cli-classify-budget.mjs         # AISDLC-141 classifier-budget filter
cli-classify-pr.mjs             # AISDLC-141 reviewer subset classifier
cli-deps.mjs                    # AISDLC-117 dependency-graph CLI
cli-deps-corpus.mjs             # AISDLC-167.5 deps-composition aggregator
cli-dor-corpus.mjs              # AISDLC-161 DoR calibration aggregator
cli-dor-digest.mjs              # AISDLC-162 DoR Slack digest
cli-dor-stats.mjs               # AISDLC-162 DoR analytics
cli-incremental-decide.mjs      # AISDLC-142 incremental review gate
cli-judgment.mjs                # judgment layer: doctor, list, ask, eval, export-corpus, replay
cli-orchestrator.mjs            # RFC-0015 autonomous orchestrator
cli-orchestrator-corpus.mjs     # AISDLC-178.7 orchestrator soak aggregator
cli-pr-unstick.mjs              # PR queue rescue helper
cli-task-complete.mjs           # AISDLC-203 atomic task lifecycle
cli-tui.mjs                     # RFC-0023 operator TUI
cli-tui-corpus.mjs              # AISDLC-178.7 TUI soak aggregator

AISDLC-245.4 will update the slash command bodies to use this resolution path so adopters don't need the full monorepo checked out.

As a workspace dependency (monorepo / dogfood)

The package lives at pipeline-cli/ in the ai-sdlc-framework/ai-sdlc monorepo:

git clone https://github.com/ai-sdlc-framework/ai-sdlc.git
cd ai-sdlc
pnpm install
pnpm --filter @ai-sdlc/pipeline-cli build
node ./pipeline-cli/bin/ai-sdlc-pipeline.mjs --help

Add it to a sibling workspace package's package.json:

{
  "dependencies": {
    "@ai-sdlc/pipeline-cli": "workspace:*"
  }
}

Invoking from CI / GitHub Actions (AISDLC-156)

Always invoke pipeline-cli CLIs via node ./pipeline-cli/bin/<bin>.mjs, never via pnpm --filter @ai-sdlc/pipeline-cli exec <bin>.

pnpm exec resolves binaries via the package's node_modules/.bin/ symlink directory, but a workspace package's OWN bin entries are NOT symlinked into its own node_modules/.bin/ — only its DEPENDENCIES' bins are. Invoking via pnpm exec from the workspace itself therefore returns:

ERR_PNPM_RECURSIVE_EXEC_FIRST_FAIL  Command "cli-classify-budget" not found

…on every invocation. In the AISDLC-156 incident, this silent failure caused the || echo '<fallback-json>' safety net in .github/workflows/ai-sdlc-review.yml to fire on every PR for the three cost-saver CLIs (cli-classify-pr, cli-incremental-decide, cli-classify-budget), defeating the AISDLC-141/142/147/149/154 optimizations entirely — every PR ran full-budget reviewers, blowing through Anthropic credits and posting CHANGES_REQUESTED whenever the key was exhausted.

The workflow now invokes each CLI as:

RESULT=$(node pipeline-cli/bin/cli-classify-pr.mjs classify --paths-file ... \
  || echo '{"reviewers":["testing","critic","security"],"fellOpen":true,...}')

pipeline-cli/src/cli/bin-invocation.test.ts is the regression guard: it spawns each bin/cli-*.mjs via node and asserts --help exits 0, AND asserts that the broken pnpm --filter ... exec form still fails (so a future operator who reverts the workflow trips a loud test failure instead of re-introducing the silent regression). When pnpm eventually fixes own-bin resolution — or we move to a different package manager — that test will fail and force a deliberate re-evaluation of whether the simpler form can be reintroduced.

From npm (standalone / non-plugin usage)

@ai-sdlc/pipeline-cli is published publicly on npmjs.org (AISDLC-245.1):

pnpm add @ai-sdlc/pipeline-cli
# Optional: only when using the API-key-billed ClaudeCodeSDKSpawner.
pnpm add @anthropic-ai/claude-code

The @anthropic-ai/claude-code SDK is a lazy import (NOT a hard dependency) so subscription-only consumers don't pay for ~50MB of SDK code they'll never use. See docs/spawner.md for the lazy-import rationale and how the failure surfaces when the SDK isn't installed.

Choosing an entry point — execute (CLI), /ai-sdlc execute (slash), pnpm dogfood watch (API key)

The Step 0-13 pipeline can be invoked three ways. They differ on who can call them, what drives the LLM dispatch, and how the work is billed. Pick the row that matches your situation:

Entry point Invoker Spawner Billing When to use
/ai-sdlc execute <task-id> (slash command body, ai-sdlc-plugin/commands/execute.md) Operator typing in their Claude Code session Agent tool calls in the SAME session Subscription (Claude Code Max) The default for internal dogfood. Operator drives, sees progress in real-time, decisions surface inline.
ai-sdlc-pipeline execute <task-id> (this CLI subcommand, AISDLC-182) Anything that can shell out — AI assistant in operator session, cron, webhook, GitHub Action Safe default plan with mock; real runs require --run with a real spawner (--spawner api-key, --spawner claude, --spawner codex, or --spawner copilot) Depends on --spawner A bare invocation is non-mutating: it validates the task and prints the planned branch/worktree without calling executePipeline(). An AI assistant working alongside the operator (or any non-slash-command context) that needs the FULL pipeline including reviewers + verdict-file write must pass explicit run intent and a real spawner: --run --spawner claude (subscription) or --run --spawner api-key (API token). --spawner mock is dry-run/plumbing only and refuses with --run. Wires AISDLC-177 rollback on every real-run outcome in the orchestrator's ROLLBACK_OUTCOMES set — developer-failed, developer-json-contract-violated, aborted, unknown-failure (the constant is imported from orchestrator/loop.ts so both surfaces stay in lockstep, AISDLC-191) — the slash command body does NOT yet wire rollback, so the umbrella is the consistency-over-parity win when an unattended dispatch fails mid-flight.
pnpm --filter @ai-sdlc/dogfood watch --issue <id> Cron / GitHub Action / unattended ClaudeCodeSDKSpawner (resolved internally) API key (paid Anthropic API) GitHub-issue-driven flow. Designed for unattended use where no operator session is available.

Why the execute umbrella subcommand exists (AISDLC-182)

Before this subcommand existed, an AI assistant working alongside the operator (e.g. Claude in the main conversation, NOT a slash command) had no clean way to invoke the full pipeline. The two existing surfaces both had gaps:

  • /ai-sdlc execute is a slash command body. Only the operator can type slash commands; an assistant cannot invoke them.
  • pnpm dogfood watch is API-key-billed. Acceptable for the GitHub-issue path; not appropriate for backlog-task internal dogfood per the dual-workflow architecture (subscription billing).

The per-step subcommands (validate-task, compute-branch, …) were exposed but no umbrella composed them into the Step 0-13 sequence. The 2026-05-04 dogfood incident — ~10 PRs shipped to main without reviewer verdicts because the assistant skipped Steps 7 (reviewers), 8 (aggregate), and 10 (verdict-file write that triggers DSSE auto-sign in the pre-push hook) — happened precisely because manually composing those steps was error-prone.

The execute subcommand is a thin wrapper around the existing executePipeline() library function. It does NOT re-implement Step 0-13; it composes them via the same in-package composite. The wrapper's only real responsibilities:

  1. Resolve a SubagentSpawner from the --spawner flag.
  2. Hook into onProgress so the per-iteration aggregated verdict lands at <worktree>/.ai-sdlc/verdicts/<task-id-lower>.json — the husky pre-push hook (scripts/check-attestation-sign.sh) reads from this exact path to auto-sign the DSSE envelope.
  3. Emit [ai-sdlc-progress] execute: <stage> lines so the dispatching session can surface progress.
# Safe default / plumbing check — validates and prints the planned branch/worktree.
# Does not call executePipeline(), create worktrees, flip task status, or push.
node ./pipeline-cli/bin/ai-sdlc-pipeline.mjs execute AISDLC-182

# Equivalent explicit dry-run form.
node ./pipeline-cli/bin/ai-sdlc-pipeline.mjs execute AISDLC-182 --dry-run

# Real run with API-key billing (requires ANTHROPIC_API_KEY in env)
node ./pipeline-cli/bin/ai-sdlc-pipeline.mjs execute AISDLC-182 --run --spawner api-key

# Mock spawner (default) is dry-run/plumbing only. This plans safely:
node ./pipeline-cli/bin/ai-sdlc-pipeline.mjs execute AISDLC-182 --spawner mock

# This refuses before validation/worktree setup because mock cannot do real work:
node ./pipeline-cli/bin/ai-sdlc-pipeline.mjs execute AISDLC-182 --run --spawner mock

--spawner options

Value Status Behaviour
mock shipped (default) Dry-run/plumbing only. A no---run invocation validates and prints the plan without resolving the spawner or mutating files. --run --spawner mock refuses before filesystem mutation.
api-key shipped Constructs the ClaudeCodeSDKSpawner (lazy SDK import). Requires ANTHROPIC_API_KEY in env and explicit --run. Same billing model as pnpm dogfood watch.
claude shipped (AISDLC-349; default for cli-orchestrator tick since AISDLC-352) Constructs the ShellClaudePSpawner — shells out to the operator's installed claude -p for each dispatch. Uses subscription auth (Agent SDK credit pool post-2026-06-15). Recommended for cron / daemon / sidecar dispatch from a plain shell.
codex shipped (AISDLC-202.2) Constructs the CodexHarnessAdapter (callback-driven Codex spawn_agent bridge). The CLI resolver wires a subprocess bridge whose path is read from CODEX_SPAWN_AGENT_BIN; when that env var is unset the resolver fails with a configuration message before any pipeline mutation. Programmatic callers can construct CodexHarnessAdapter directly with their own CodexSpawnAgentFn. Design map: docs/operations/codex-execution-path.md. Billing: Codex plan.
copilot shipped (AISDLC-429.2 + AISDLC-429.3) Constructs the CopilotHarnessAdapter (callback-driven GitHub Copilot CLI bridge). The CLI resolver wires a subprocess bridge whose path is read from COPILOT_SPAWN_AGENT_BIN; when that env var is unset the resolver fails with a configuration message before any pipeline mutation — refuses to silently fall back to ANTHROPIC_API_KEY billing. Programmatic callers can construct CopilotHarnessAdapter directly with their own CopilotSpawnAgentFn. Design map: docs/operations/copilot-execution-path.md. Operator runbook: docs/operations/copilot-spawner.md. Billing: GitHub Copilot subscription.
claude-cli removed (RFC-0041 Phase 3.3 / AISDLC-377.6) Was the ClaudeCliInlineSpawner inline-manifest path (AISDLC-198). Deleted after the AISDLC-377.4 deprecation-warning window. Yargs --spawner claude-cli is rejected at parse time; programmatic callers receive CLAUDE_CLI_SPAWNER_REMOVED_MESSAGE. Migration: docs/operations/claude-cli-spawner-removed.md.
--spawner codex — Codex CLI host-bridge dispatch (AISDLC-202.2 + AISDLC-251)

Phase 2 of the Codex execution path ships the CodexHarnessAdapter, a host-agnostic SubagentSpawner over Codex's spawn_agent host tool. Codex CLI does not expose Claude Code's plugin Agent system, so the adapter centralises the developer + reviewer dispatch contract that the AISDLC-201 run had to reconstruct by hand:

  • Per-SubagentType system prompts derived from ai-sdlc-plugin/agents/<type>.md (overridable via constructor option).
  • A single injected spawnAgent callback that wraps the operator's host bridge — no Codex CLI version coupling lives in pipeline-cli.
  • Reviewer envelopes returned by the adapter pass through Step 8 aggregation without manual reshaping (coerceReviewerVerdict consumes them directly; verdicts are tagged harness: 'codex').

Canonical bridge script (AISDLC-251): The repo ships a ready-to-use bridge at scripts/codex-spawn-agent-bridge.mjs. Set CODEX_SPAWN_AGENT_BIN to this script — no hand-written bridge needed:

# Canonical setup — set CODEX_SPAWN_AGENT_BIN to the canonical bridge script,
# then run ai-sdlc-pipeline execute with --spawner codex.
export CODEX_SPAWN_AGENT_BIN="$(pwd)/scripts/codex-spawn-agent-bridge.mjs"
node ./pipeline-cli/bin/ai-sdlc-pipeline.mjs execute AISDLC-NNN --run --spawner codex

The canonical bridge uses the verified flag set for codex-cli 0.128.0: codex exec -s workspace-write --skip-git-repo-check --color never for developer dispatch, and codex exec -s read-only --skip-git-repo-check --color never for reviewer dispatch. A developer agent must be able to edit the task worktree; reviewers remain read-only. The composed prompt is passed on stdin with - as the prompt argument. DO NOT use --quiet (errors with "unexpected argument") or --model o4-mini (HTTP 400 on ChatGPT-account auth) — see AISDLC-249/247 smoke test notes. If the bridge or codex exec exits 0 with empty stdout, the adapter treats that as a bridge error with diagnostics instead of returning an empty developer JSON envelope. The bridge only forwards safe model/provider extraArgs (--model/-m, --oss, --local-provider) and strips config, sandbox, cwd, and writable-root overrides.

Two ways to use it:

# CLI form — requires CODEX_SPAWN_AGENT_BIN (use canonical bridge above or your own).
export CODEX_SPAWN_AGENT_BIN="$(pwd)/scripts/codex-spawn-agent-bridge.mjs"
node ./pipeline-cli/bin/ai-sdlc-pipeline.mjs execute AISDLC-202 --run --spawner codex
// Programmatic form — inject any CodexSpawnAgentFn (host-tool wrapper,
// in-process bridge, etc.). Tests use a deterministic mock.
import { CodexHarnessAdapter } from '@ai-sdlc/pipeline-cli';

const adapter = new CodexHarnessAdapter({
  spawnAgent: async ({ agentType, systemPrompt, userPrompt, cwd, timeoutMs }) => {
    // Wrap Codex's spawn_agent host tool here.
    return { output: '<agent JSON return>', parsed: { /* optional pre-parse */ } };
  },
});

Bridge JSON-line protocol (the wire format subprocessCodexSpawnAgent produces / consumes):

Direction Stream Shape
Adapter → bridge stdin (single JSON line) { agentType, systemPrompt, userPrompt, cwd, timeoutMs }
Bridge → adapter stdout (single JSON envelope) { output: string, parsed?: unknown }
Bridge → adapter exit code 0 for success; non-zero surfaces stderr as the error

The adapter normalises reviewer responses into the canonical ReviewerVerdict envelope ({ approved, findings, summary, harness: 'codex' }) before returning, so Step 8 aggregation runs unchanged.

cli-judgment - measure and inspect the judgment layer

cli-judgment is the operator tool for the judgment layer: check the setup, run one evaluation, measure a judgment over a labelled corpus before promoting it, and replay logged answers under other thresholds. Invoke it directly:

node pipeline-cli/bin/cli-judgment.mjs <command> [options]

Shared options: --cwd <dir> (working directory), --config <file> (use this judgment config file instead of the one on the trusted base branch), --artifacts-dir <dir> (default $ARTIFACTS_DIR, else .ai-sdlc/artifacts). The API key is read from the provider's environment variable and is never printed.

Command What it does
doctor [--live] Reports whether the layer is enabled, the provider, whether its key is present, and whether the model is pinned to an exact version. --live sends one minimal request and reports the returned model version and latency (exit 1 if it fails).
list Lists every registered judgment with id, version, riskClass, direction, egressClass, the configured mode and the effective mode (with the reason when the runtime would downgrade or disable it).
ask <judgment-id> --input <json-file> Runs one forced evaluation and prints the answers with probabilities, the outcome and the thresholds used, as JSON. Refuses, naming the config key (spec.egress.allow), when the judgment's egress class is not allowed.
eval <judgment-id> --corpus <jsonl> Runs the judgment over a corpus (one {"input": ..., "label": ...} object per line), with the answer cache on, composes at the configured thresholds and compares with the judgment's agrees. Prints and writes n, the share of items in the act, escalate and abstain bands, act-band precision, a confusion table of decision against label, latency p50 and p95, total input tokens and cost.
export-corpus <task-type> [--corpus-dir <dir>] [--out <file>] Turns a classifier calibration corpus (.ai-sdlc/classifier-corpus/<task-type>.yaml) into eval JSONL. Only entries carrying an operator override become rows; the override is the label. Task types: capture-triage, capture-severity, pr-comment-is-capture, dor-answer-is-new-concern, decision-recommendation. Prints to stdout unless --out is given, so eval capture.triage --corpus <file> can measure the classifier judgments (capture.triage, capture.severity, capture.pr-comment, dor.answer-segment, decision.recommendation).
replay --since <date> [--judgment <id>] Reads the judgment log and recomputes outcomes from the logged answers, with no provider calls. Without --threshold it uses the logged thresholds and reports how many logged outcomes it reproduced; it also reports agreement with the logged incumbent where present. The log keeps a hash of the input, not the input, so a judgment whose compose reads its input cannot be replayed.

ask, eval and replay accept --threshold <name>=<number> (repeatable) to override the configured thresholds, and --source-kind <kind> (default backlog; only backlog items may be decided permissively).

eval extras:

  • --sweep <name>=<from>:<to>:<step> recomputes the report for each threshold value from the answers already collected, so a sweep makes no additional provider calls. step must be above zero and a sweep is limited to 1000 steps.
  • The report is written to .ai-sdlc/judgment-evals/<id>-<provider>-<model>-<date>.json under the working directory (path components are sanitised).
  • It prints the promotion snippet for .ai-sdlc/judgment-config.yaml, with n and actBandPrecision equal to the report, and states whether the result is MET or NOT MET for the judgment's riskClass bar (corpus path: n of at least 50 and act-band precision of at least 90%, or 95% for relax). Promotion itself stays an operator decision made in a pull request.
  • Exit code: non-zero only when the corpus is unreadable or malformed (the message names the line), the judgment is unknown or has no agrees, or a flag is invalid. It exits zero whether or not the bar is met.

A live check of the provider contract (one request with a choice, a score and a yes/no question) sits next to the Jev adapter tests and runs only when TYPESAFE_API_KEY is set and AI_SDLC_LIVE_CONTRACT=1; otherwise it is reported as skipped.

cli-usage ingest — usage ledger from Claude Code transcripts

cli-usage ingest reads the Claude Code session and subagent transcripts on this machine and appends one record per model call to the machine-level usage ledger (~/.ai-sdlc/usage/ledger-YYYY-MM.jsonl, or the directory in AI_SDLC_USAGE_DIR). The ledger is never committed.

node pipeline-cli/bin/cli-usage.mjs ingest [--backfill] [--projects-dir <path>] [--max-seconds <n>] [--json]
Option Meaning
--backfill Ignore stored cursors and read every transcript from the start (already-seen calls are skipped, so this is safe to repeat).
--projects-dir Transcript projects directory. Default: $CLAUDE_CONFIG_DIR/projects, else ~/.claude/projects.
--max-seconds Stop starting new work after this many seconds (default 30). Progress is saved; the next run continues.
--json Print the result as JSON.

The result reports files scanned, calls written, repeats skipped and errors.

Repeated lines for one message. The same message id appears on several transcript lines. Within one batch (up to 2000 records, flushed at the end of each transcript) the line with the largest output count is kept; across batches and runs the first record written wins and later repeats are skipped. A message whose lines straddle a batch boundary or two runs can therefore keep a slightly lower output count than its final line reports. This is a known limitation.

What is stored. Counts, ids, model, timestamp, harness, billing pool and attribution only. No prompt, response, file content, tool output or subagent description is ever read into a record, logged or printed. A usage or rate-limit notice found in a transcript becomes a line in limit-events.jsonl holding its timestamp, session id and a short fixed category, never the message text.

Scope. A call made inside a repository that has an .ai-sdlc/ directory is framework scope and keeps its repository, task and source file. Every other call is other scope: tokens, model, timestamp and harness are kept; repository, task, working directory, branch and source are not written. Set AI_SDLC_USAGE_SCOPE=framework-only to skip transcripts of other projects entirely.

Attribution. The agent role is the subagent sidecar's agentType (main-session for a session transcript, subagent-unknown when the sidecar is missing). The task comes from a .worktrees/<task-id> path segment, then a task id in the branch name that has a backlog file, then the worktree's .active-task file. The billing pool is set only when the transcript states its entrypoint; otherwise it is unknown.

Safety. Transcripts are treated as untrusted input: symlinks are not followed, file and directory names are validated, lines over 4 MiB are skipped and counted as errors, and a truncated last line is retried on the next run. Ingestion does nothing in a remote sandbox (CLAUDE_CODE_ENV=ccr or CLAUDE_REMOTE_EXECUTION=1). Set AI_SDLC_USAGE_INGEST=off to switch it off.

Triggers. The plugin's Stop and SessionStart hooks and each orchestrator tick launch ingestion as a detached background process with a time limit, so a session or tick is never delayed and any failure is swallowed.

cli-usage reports — where the allotment is going

Reports read the usage ledger built by cli-usage ingest. They print counts, ids and attribution only, never prompt, response, file content or tool output.

cli-usage report --group-by model [--group-by role ...] [--since <iso>] [--until <iso>] [--scope framework|other|all] [--format text|json|csv]
cli-usage window                       # units used in the current session and weekly windows, implied allotment, time to the limit
cli-usage task <id>                    # tokens and units for one task, split by role
cli-usage context                      # per session: first-call tokens, turns, total cache read
cli-usage snapshot --window weekly --used-pct 42
cli-usage allotment [--window weekly]  # implied allotment per snapshot, with probable changes marked

--group-by takes model, role, task, repo, pool, day or window and can be repeated. Each row shows calls, input, cache write (5 minute and 1 hour), cache read and output tokens, weighted units and API-equivalent cost. A model with no price shows unpriced, is left out of the cost totals, and the total is labelled partial. --until is exclusive.

Units are a proxy. The provider reports subscription use as a percentage and publishes no conversion from tokens, so reports weigh each call in units. By default one input token of the reference model is one unit and every other token class and model is weighed against it using the current price history, so the weights follow the price feed. Weights set in the usage config override the derived ones. Every report that shows units says that the weights are a proxy.

Calibration. cli-usage snapshot records the percentage the provider shows for a window, together with the units the ledger counted in that window at that moment (written to snapshots.jsonl in the usage directory). The implied allotment is units divided by the fraction used. Window observations that a harness wrote to limit-events.jsonl are used as snapshots too. When two consecutive snapshots of one window imply allotments that differ by more than the tolerance while their model mix is similar, the row is marked as a probable allotment change and AllotmentChangeSuspected is emitted. Recording a snapshot emits UsageLimitObserved.

Config. Plan name, monthly price, windows, unit weights and the tolerance live in .ai-sdlc/usage-config.yaml (kind UsageConfig). It is read from the base branch (origin/main), never from the working tree. A usage-config.yaml in the usage directory takes precedence on that machine. With no file the defaults apply: a 5 hour session window that opens at the first call, a trailing 168 hour weekly window, a 25% tolerance and a model-mix overlap of at least 0.8. ai-sdlc init ships a commented template at .ai-sdlc/templates/usage-config.yaml.

Window modes. first-use opens a window at the first call after the previous one ended; fixed counts repeating cycles from an anchor (your plan's reset time); trailing is the last lengthHours. The projected time to the limit divides the implied allotment still unused by the rate since the window's first call.

Context overhead. cli-usage context shows the size of the context at the first call of each session (the fixed prefix every later turn re-reads), the number of turns and the total cache read, largest first. Sessions outside a framework repository (other scope) show no path.

Cost report. cli-cost-report --usage-ledger (or --usage-dir <path>) builds the unified view from the usage ledger and prefers it over --cost-ledger-jsonl and --ledger-dir when it has records. The existing inputs keep working without it.

Health. A successful ingest reports the usage.ingest capability as live and a failed one as degraded with a reason. ai-sdlc doctor shows the time of the last successful ingest.

Quickstart — Tier 1 (slash command body)

The /ai-sdlc execute slash command body (in ai-sdlc-plugin/commands/execute.md) calls the CLI directly. To rebuild a fragment of that flow by hand:

# Always passes JSON on stdout.
node ./pipeline-cli/bin/ai-sdlc-pipeline.mjs --help

# Step 0 — sweep merged worktrees.
node ./pipeline-cli/bin/ai-sdlc-pipeline.mjs sweep-worktrees

# Step 1 — validate the task is ready to execute.
node ./pipeline-cli/bin/ai-sdlc-pipeline.mjs validate-task AISDLC-100.7

# Step 2 — compute the branch name + worktree path.
node ./pipeline-cli/bin/ai-sdlc-pipeline.mjs compute-branch AISDLC-100.7

# Step 3 — create the worktree.
node ./pipeline-cli/bin/ai-sdlc-pipeline.mjs setup-worktree AISDLC-100.7

# Step 4 — flip status + write .active-task sentinel.
node ./pipeline-cli/bin/ai-sdlc-pipeline.mjs begin-task AISDLC-100.7

# Step 5 — render the developer subagent prompt (caller dispatches separately).
node ./pipeline-cli/bin/ai-sdlc-pipeline.mjs build-dev-prompt AISDLC-100.7

Tier 1's distinctive trait is that the LLM dispatch boundaries (Step 5b — spawn developer, Step 7b — spawn 3 reviewers) are NOT calls into pipeline-cli — they are direct Agent(developer, code-reviewer, test-reviewer, security-reviewer) tool calls in the main Claude Code session. The slash command body parses the JSON each pipeline-cli subcommand emits and feeds the next step.

Quickstart — Tier 2 (executePipeline())

For unattended use — webhooks, cron, GitHub Actions, custom dashboards — import the composite and pass a spawner:

import { executePipeline, defaultSpawner } from '@ai-sdlc/pipeline-cli';

const spawner = await defaultSpawner();
//   ↳ resolves to ShellClaudePSpawner if `claude` CLI is on PATH
//     (subscription, no tokens spent), otherwise to ClaudeCodeSDKSpawner
//     if ANTHROPIC_API_KEY is set (API key billing), otherwise throws.

const result = await executePipeline({
  taskId: 'AISDLC-100.7',
  workDir: process.cwd(),
  spawner,
  // optional: cap on TOTAL review iterations (default 2 — initial + 1 retry)
  maxReviewIterations: 2,
  // optional: progress callback per iteration
  onProgress: (iteration, verdict) => {
    console.log(`iteration ${iteration}: ${verdict.decision}`);
  },
});

console.log(result.outcome);
//   ↳ 'approved' | 'needs-human-attention' | 'developer-failed' | 'aborted'
console.log(result.prUrl);          // string | null
console.log(result.siblingPrUrls);  // string[]

This is the canonical Tier 2 pattern. Phase 5 (AISDLC-100.5) migrates the existing dogfood/src/watch.ts to call executePipeline() directly instead of re-implementing the orchestration in TypeScript prose. New unattended consumers (webhooks, cron, GitHub Actions) should follow the same shape: build a SubagentSpawner once, then await executePipeline({ taskId, workDir, spawner }) per task.

For tests, swap defaultSpawner() for MockSpawner:

import { executePipeline, MockSpawner } from '@ai-sdlc/pipeline-cli';

const spawner = new MockSpawner({
  developer: {
    type: 'developer',
    output: '...',
    parsed: { /* DeveloperReturn shape */ },
    status: 'success',
    durationMs: 0,
  },
  'code-reviewer':     { /* ReviewerVerdict shape */ },
  'test-reviewer':     { /* ... */ },
  'security-reviewer': { /* ... */ },
});

const result = await executePipeline({
  taskId: 'AISDLC-EXAMPLE',
  workDir: tmpProjectRoot,
  spawner,
  skipFinalizeCommit: true, // tests usually don't have a real git repo
});

See docs/spawner.md for the full SubagentSpawner catalogue (ShellClaudePSpawner, ClaudeCodeSDKSpawner, defaultSpawner(), MockSpawner, custom spawner howto).

Layout

Tests live next to the code they exercise (one *.test.ts per source file). The integration test for the Tier 2 composite is colocated with execute-pipeline.ts.

pipeline-cli/
├── package.json
├── README.md                       (this file)
├── docs/
│   ├── spawner.md                  # SubagentSpawner deep-dive (when to use which, custom-spawner howto)
│   └── steps.md                    # Per-step contract / inputs / outputs / side effects
├── tsconfig.json
├── vitest.config.ts
├── bin/
│   └── ai-sdlc-pipeline.mjs        # shebang wrapper around dist/cli/index.js
└── src/
    ├── index.ts + index.test.ts    # public barrel + barrel re-export coverage
    ├── types.ts + types.test.ts    # PipelineOptions, StepResult, SubagentSpawner, etc.
    ├── execute-pipeline.ts         # Tier 2 composite entry point
    ├── execute-pipeline.test.ts    # full Step 0-13 integration with MockSpawner
    ├── __test-helpers/             # FakeRunner + tmp backlog task fixture builder
    │   ├── fake-runner.ts
    │   └── make-task.ts
    ├── runtime/
    │   ├── index.ts                # barrel — exports SubagentSpawner + Runner surface
    │   ├── exec.ts                                                   # Runner abstraction over child_process.execFile
    │   ├── subagent-spawner.ts                                       # SubagentSpawner interface + MockSpawner
    │   ├── shell-claude-p-spawner.ts                                 # Tier 2 default — `claude --print --agent <type>` shell-out (subscription)
    │   ├── claude-code-sdk-spawner.ts                                # Tier 2 alternative — @anthropic-ai/claude-code SDK (API key)
    │   └── default-spawner.ts                                        # `defaultSpawner()` resolver: which→shell, env→sdk, else throw
    ├── steps/                      # each step.ts has a colocated step.test.ts
    │   ├── index.ts                # barrel
    │   ├── 00-sweep.ts             # Step 0 — sweep merged worktrees
    │   ├── 01-validate.ts          # Step 1 — validate backlog task spec
    │   ├── 02-compute-branch.ts    # Step 2 — branch name + worktree path
    │   ├── 03-setup-worktree.ts    # Step 3 — git worktree add
    │   ├── 04-flip-status.ts       # Step 4 — status flip + .active-task sentinel
    │   ├── 05-build-dev-prompt.ts  # Step 5 — developer prompt template
    │   ├── 06-parse-dev-return.ts  # Step 6 — parse + gate developer JSON
    │   ├── 07-build-review-prompts.ts # Step 7 — 3 reviewer prompts
    │   ├── 08-aggregate-verdicts.ts   # Step 8 — verdict aggregation
    │   ├── 09-iterate.ts           # Step 9 — review iteration loop
    │   ├── 10-finalize.ts          # Step 10 — Done + completed/ + attestation + chore commit
    │   ├── 11-push-and-pr.ts       # Step 11 — push + gh pr create
    │   ├── 12-sibling-prs.ts       # Step 12 — cross-repo sibling PRs
    │   └── 13-cleanup.ts           # Step 13 — sentinel cleanup
    └── cli/
        └── index.ts                # yargs subcommand router

Step contracts

Every step exports a pure async function. The return shape is documented in src/types.ts + the per-step JSDoc and consolidated in docs/steps.md. The JSON returned by the CLI subcommands matches the TypeScript return shape exactly.

# Step Function CLI command
0 Sweep merged worktrees sweepMergedWorktrees sweep-worktrees
1 Validate task validateTask validate-task <id>
2 Compute branch computeBranchName compute-branch <id>
3 Setup worktree setupWorktree setup-worktree <id>
4 Begin task (flip status + sentinel) beginTask begin-task <id>
5 Build developer prompt buildDeveloperPrompt build-dev-prompt <id>
6 Parse developer return parseDeveloperReturn parse-dev-return --return <json>
7 Build review prompts buildReviewPrompts build-review-prompts <id>
8 Aggregate verdicts aggregateVerdicts aggregate-verdicts --verdicts <json>
9 Iterate review loop iterateReviewLoop (Tier 2 composite only)
10 Finalize task finalizeTask finalize-task <id> --developer-return <json> --verdict <json>
11 Push + open PR pushAndPr push-and-pr <id> --developer-return <json> --verdict <json>
12 Sibling PRs siblingPrs sibling-prs <id> --developer-return <json> --main-pr-url <url>
13 Cleanup sentinel cleanupTask cleanup-task <id>

SubagentSpawner contract (LLM dispatch boundary)

The pipeline is purely deterministic except for two LLM dispatch points:

  • Step 5b — spawn the developer subagent with the prompt rendered in Step 5
  • Step 7b — spawn code-reviewer, test-reviewer, security-reviewer in parallel

Both go through the SubagentSpawner interface (RFC-0012 §8). That's the only piece of the pipeline that varies between Tier 1 (Agent tool from the main session), Tier 2 subscription (claude --print), Tier 2 API key (Claude Code SDK), and tests (MockSpawner).

interface SubagentSpawner {
  spawn(opts: SpawnOpts): Promise<SubagentResult>;
  spawnParallel(opts: SpawnOpts[]): Promise<SubagentResult[]>;
}

Production spawners (Phase 2 — AISDLC-100.2):

  • ShellClaudePSpawner (subscription billing) — shells out to the operator's installed claude CLI with claude --print --output-format json --permission-mode bypassPermissions --agent <type> <prompt>. No API tokens consumed; reuses the operator's logged-in Claude Code session. RFC §8.2's sketch said --subagent <type> but the actual flag is --agent <type> (verified empirically against the installed CLI on 2026-04-30). See docs/spawner.md for the full Q5 (RFC §15) resolution.
  • ClaudeCodeSDKSpawner (API-key billing) — uses @anthropic-ai/claude-code programmatically. The SDK is lazy-imported (NOT a hard dependency of pipeline-cli) so subscription-only consumers don't have to install ~50MB of SDK code they'll never use; install it with pnpm add @anthropic-ai/claude-code only when you need API-key billing.
  • defaultSpawner() — convenience resolver: prefers claude CLI on PATH (subscription), falls back to ANTHROPIC_API_KEY (API key), throws with an instructional error if neither is available.

MockSpawner (shipped here for tests) accepts either fixed results per subagent type or a callback per type so iteration N>1 can return different fixtures than iteration 1.

Full deep-dive — selection guide, custom spawner howto, lazy-import mechanics, Q5 resolution: docs/spawner.md.

Hard rules (NEVER violated by any step)

These come from RFC-0012 §3.1 + the AI-SDLC governance hooks:

  1. Never gh pr merge. Step 11 only opens PRs.
  2. Never git push --force / -f. Step 11 aborts cleanly on non-fast-forward.
  3. Never delete branches (no git branch -D / -d).
  4. Never edit .ai-sdlc/**. Pre-tool-use hook blocks anyway. .github/workflows/** is only refused when the project's .ai-sdlc/agent-role.yaml lists it under blockedPaths — not blocked by default.
  5. Never run destructive git ops (no reset --hard, checkout -- ., restore .).
  6. Step 13 is mandatory — the sentinel is removed in a finally block from executePipeline.

Testing

  • Unit tests are colocated with the source they exercise — src/steps/<step>.ts ↔ src/steps/<step>.test.ts. Each step has happy-path + error-path coverage.
  • Integration test lives at src/execute-pipeline.test.ts and runs the full Step 0-13 against MockSpawner + FakeRunner in a tmp project root.
  • Test helpers (FakeRunner, makeTmpProject, writeTaskFile) live in src/__test-helpers/ so they're picked up by Vitest's default include glob alongside the colocated *.test.ts files.
  • Coverage gate is 80% lines/functions, enforced by vitest.config.ts and the workspace-level scripts/check-coverage.sh.
pnpm test                  # vitest run
pnpm test:coverage         # with v8 coverage + thresholds
pnpm test:watch            # iteration mode

Documentation

  • docs/spawner.md — SubagentSpawner selection guide (when to use ShellClaudeP / ClaudeCodeSDK / Mock / custom), lazy SDK import, Q5 resolution.
  • docs/steps.md — per-step contract, inputs, outputs, side effects, when each step runs.
  • docs/decisions-config.md — adopter-facing schema reference for .ai-sdlc/decisions-config.yaml (RFC-0035 Phase 7 / AISDLC-291): pillar owners, capacity tiers, fatigue config, cli-decisions fatigue {set, clear, status} CLI surface.
  • spec/rfcs/RFC-0012-two-tier-pipeline-architecture.md — the full design.

Phases

Phase Task What changes Status
1 AISDLC-100.1 Create pipeline-cli/ — Step 0-13 pure functions + CLI router + executePipeline() composite shipped
2 AISDLC-100.2 ShellClaudePSpawner + ClaudeCodeSDKSpawner + defaultSpawner() shipped
3 AISDLC-100.3 Wrap each step function as an MCP tool in ai-sdlc-plugin/mcp-server/ in flight
4 AISDLC-100.4 Refactor commands/execute.md to use the CLI; delete agents/execute-orchestrator.md in flight
5 AISDLC-100.5 Migrate dogfood/src/watch.ts to call executePipeline() in flight
6 AISDLC-100.6 Add pipelineVersion to attestation envelope shipped
7 AISDLC-100.7 Documentation pass — README + spawner doc + per-step docs shipped
8 AISDLC-100.8 / AISDLC-245.1 Publish as @ai-sdlc/pipeline-cli to npm; ai-sdlc-plugin declares it as runtimeDependencies shipped

See spec/rfcs/RFC-0012-two-tier-pipeline-architecture.md for the full design.