Skip to content

Multi-harness support: Codex + agents-md targets, target-matrix contract suite, library workflow, flag-drift canary - #9

Open
ccevans wants to merge 14 commits into
mainfrom
feat/multi-harness
Open

ccevans wants to merge 14 commits into
mainfrom
feat/multi-harness

Conversation

@ccevans

@ccevans ccevans commented Aug 2, 2026

Copy link
Copy Markdown
Owner

What's in here

Bobby now supports five harnesses, and — more importantly — adding the sixth is cheap and can't repeat the class of bug that shipped in Cline for its entire life.

New targets

target Rules Skills Subagents Dashboard executor
codex AGENTS.md .codex/skills/ No codex (codex exec --json)
agents-md AGENTS.md .agents/skills/ No — (generic tier)

agents-md is the generic tier for the AGENTS.md ecosystem (Copilot, Windsurf, Zed, opencode, Jules, Amp, Factory…). Its limits are enforced by tests, not just documented: no subagent claim, no executor derivation, nothing tool-specific written.

The foundation: one contract, every target

test/lib/target-matrix.test.js runs ~19 shared invariants against every registered target via describe.each(TARGETS). Registering a target inherits the whole suite — no hand-written per-target tests.

It exists because per-target tests let the same bug live in one target and not another. Proof it works: reintroducing the literal CLAUDE.md reference in templates/agents/bobby-build.md.ejs fails the cline and cursor legs while claude-code (whose rules file that is) correctly passes. Both new targets passed the suite with zero test edits beyond registration.

library workflow — CLI and library projects no longer stall

Every built-in workflow except design ended in test, and bobby-test verifies by exercising a running app and is forbidden from running the spec suite. On a project with nothing to serve, every test case came back BLOCKED and tickets silently stopped advancing — no error, no guidance.

New library / library-secure workflows end at review (which runs the suite independently). bobby init picks them automatically for stacks with no health checks and no dev command, writing the choice visibly into .bobbyrc.yml. Live-app agents invoked with nothing to observe now block with an actionable reason instead of stalling.

Found by dogfooding: bobbycode itself hit this on day one, and its local workaround is deleted in this PR in favour of the built-in.

Weekly flag-drift canary

Per-PR verification proves argv is correct at merge time and nothing after; these CLIs ship weekly. A scheduled workflow (Mondays + manual dispatch, never on push/PR) runs Bobby's real argv against each CLI unauthenticated: reaching auth is the passing state, unknown option is drift. A control probe adds a bogus flag that must be rejected, so a CLI that ignores unknown flags can't produce a vacuous pass.

argv comes from the real buildArgs — an AC asserts zero flag strings appear in the workflow file, so the canary can't drift from the code it guards.

Verification

Every path, flag, and convention here was checked against a real CLI run or the tool's shipped code, never documentation. That rule is in the README's contributing section because every wrong claim Bobby has shipped came from a docs page or a --help example. It caught three real problems while building this:

  • -p means --profile in Codex, and the prompt is positional. Reusing the claude -p / cursor-agent -p shape would have silently passed the entire prompt as a config profile name — no crash, no unknown-option error, just an agent with no instructions.
  • .codex/skills/ is user-level ($CODEX_HOME/skills), not project-level as this epic's ticket originally assumed. The loop doesn't depend on project auto-loading — AGENTS.md names the directory and prompts reference skills by path.
  • The canary reported drift in a CLI that was behaving perfectly. resolveExecutor treats an unknown flavor as a custom binary with claude-style flags — correct for dashboard.executor: /path/to/x, catastrophic in a probe. Now guarded against EXECUTOR_NAMES.

Live results: Bobby's exact argv starts a real turn against @openai/codex 0.146.0; cursor-agent clears argument parsing. The README support matrix records per-row verification status (Live / Live argv / Paths verified / Unverified) rather than implying uniform confidence — cline is explicitly labelled best-effort.

Test plan

  • 956 tests across 49 suites, lint clean (0 errors; 41 pre-existing warnings)
  • A script cross-checks every documented path, subagent flag, and executor derivation against the actual adapters — all five README rows match shipped behaviour
  • Matrix bug-catch verified by deliberately reintroducing the CLAUDE.md bug (locally, not committed)
  • Canary run against both installed CLIs, plus its missing-binary and unknown-flavor paths

Not in scope

Tier-2 targets (Copilot, OpenCode, Windsurf, Zed) are tracked but unbuilt — Windsurf/Zed/opencode aren't installed here, and shipping adapters from documentation is the exact failure mode above. Gemini/Antigravity is a dated hold until its rename settles.

🤖 Generated with Claude Code

ccevans and others added 14 commits July 31, 2026 01:39
One app UI for the whole solo-dev loop, served by `bobby app` (alias:
`bobby dashboard`; classic UI frozen at /classic/ for one release).
Home = brief + one "Do this next" button + the Needs-you queue; Board
with ticket create/detail; workspace detail with live SSE logs, diff,
and Approve / Send back / Merge. Vanilla ESM, no build step, PWA
manifest, HQ design tokens (stage lamps, amber = needs-you), desktop
rail + phone tab bar. Transport is one seam (LocalTransport now) so
the same frontend later runs over the encrypted relay.

Backend: GET /api/brief, POST /api/go (executes nextAction.argv via
new lib/dashboard/actions.js — server-side twin of `bobby go`), ticket
write routes (create/move-with-aliases/patch/comments), GET
/api/workflows, GET /api/config; buildServer gains appDir/sprintsDir.

Fixes:
- per-workspace pipeline was recorded but not honored — approve()
  always advanced through the server default workflow (_pipelineFor)
- an EventSource opened during page load kept the window load event
  (and the tab spinner) hanging forever

Verified: 855 tests green; Playwright walked the real app end to end
against a stub executor — Home → Do it → confirm → live log →
Needs-you → Approve — on desktop and phone viewports, zero console
errors; /classic/ serves the old UI with rewritten asset paths.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
One command, two tiers, no hostages. Without a Pro key, `bobby app`
(and its alias `bobby dashboard`) serves the classic dashboard exactly
as published — free forever. With Bobby Pro, the same command serves
the App, with classic kept at /classic/.

The App UI moves out of this MIT package into @bobbycode/pro-dashboard
(private), delivered via `bobby pro install` — a license check on
MIT-licensed files would be theatre, so the paid code simply never
ships here. commands/app.js resolves the UI: BOBBY_APP_DIR (dev) →
Pro key + installed package → classic fallback, and prints which tier
it chose and why.

The loop API (brief/go/ticket writes/workflows/config) stays MIT — it
serves the classic UI and the relay too.

Verified all three tiers live: no key → classic + upsell note; dev
override → App; real signed key + `bobby pro install <tarball>` from
the actual buyer zip → App at /, classic at /classic/. 855 tests green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Approach A (foundation-first) selected 36/29/19: matrix invariant suite
first, then Codex target + executor, then generic agents-md target, docs
last. Scoring rationale, spec, and the no-unverified-harness-claims rule
in the epic's plan.md.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The Pro package now carries both the Bobby App (app/) and the classic
dashboard add-on assets (ui/), so the App lookup moves to app/.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…Zed, Gemini hold

Six more children under the multi-harness epic, all demand-driven and
gated on the matrix suite + agents-md generic tier landing first.
Windsurf/Zed are spikes where "covered by generic" is a valid close;
Gemini/Antigravity is an explicit dated hold, not a build.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The proxy test slept a fixed 50ms for a real HTTP round trip, which
lost the race under full-suite parallel load and failed intermittently.
It now resolves on the frame itself.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
bobbycode has no live app, so the built-in default workflow's final test
stage (live-app verification by contract) would strand every ticket at
testing. Override default/secure to end at review locally, define real
feature areas, and file TKT-013 (critical) for the upstream fix: built-in
library workflow, stack-aware default, and a loud block instead of a
silent stall. Also notes the workflow-list display bug for overridden
workflows.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Per-PR verification proves argv correct at merge time only; harness CLIs
ship weekly. A scheduled workflow probes each executor's real buildArgs
output against the real binary: auth failure = pass, unknown-option =
drift, plus a control probe so a CLI that ignores unknown flags can't
make the canary vacuous. TKT-004/TKT-009 now require registering their
flavors in it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
One contract, run against every target in the registry via
describe.each(TARGETS). Adding a target now inherits ~all coverage from
registration alone instead of needing a hand-written suite.

Invariants: adapter interface shape, all declared paths written, extras
created, no writes into another target's directories, no scaffolded file
referencing another target's rules file, rules names its own harness,
skills/agents reference own paths, hooks scaffolded only where supported,
idempotent rescaffold, pre-existing rules backed up and merged, command
frontmatter matches the adapter's own transform, prompts reference only
this target's agents path.

Verified the suite catches the bug it exists for: reintroducing the
literal `CLAUDE.md` reference in templates/agents/bobby-build.md.ejs fails
the cline and cursor legs while claude-code (whose rules file that is)
correctly passes. That bug shipped in Cline for its entire life because
only cursor ever got a hand-written sweep test for it.

Removes 262 lines of per-target assertions the matrix now covers; keeps
genuine quirks (cursor transformCommand edge cases, ignore-file contents).

Refs TKT-002

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Target: rules -> AGENTS.md, skills -> .codex/skills, agents ->
.codex/agents, prompts -> .codex/prompts.

Executor: `codex exec --json`, derived from target: codex.

Everything verified against the real @openai/codex 0.146.0 binary
(strings + `codex exec --help`), never documentation:

- AGENTS.md is Codex's native project-instruction mechanism — the binary
  carries a full "AGENTS.md spec" in its base instructions and ships
  /init to create one.
- The prompt is POSITIONAL. `-p` is `--profile` in Codex, so reusing the
  claude/cursor `-p <prompt>` shape would have silently passed the whole
  prompt as a config profile name with no error. The prompt is emitted
  last so it can never be read as a flag value.
- JSONL comes from `--json`; there is no --output-format and no --verbose.
- Permission modes map to verified flags:
  bypassPermissions -> --dangerously-bypass-approvals-and-sandbox
  acceptEdits       -> --sandbox workspace-write --ask-for-approval never
  plan              -> --sandbox read-only

Correction to the ticket's premise: Codex skills are USER-level
($CODEX_HOME/skills, default ~/.codex/skills). Project-level auto-loading
of .codex/skills/ is NOT confirmed, so the loop does not depend on it —
AGENTS.md names the directory and every prompt references its agent and
skill by path. supportsSubagents() is false for the same reason: the
binary mentions subagents but the definition format is unverified.

Live-verified: Bobby's exact argv run against the real binary starts a
turn and streams JSONL (thread.started / item.completed / turn.started),
with no unknown-option error.

Codex passes all 19 target-matrix invariants with zero test edits beyond
registration — the matrix suite (TKT-002, cherry-picked here to prove it)
went 47 -> 66 tests on target count alone. 914 total green.

Refs TKT-003, TKT-004

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Rules -> AGENTS.md, skills -> .agents/skills, agents -> .agents/agents,
commands -> .agents/commands. Lets Bobby work in any tool that reads the
AGENTS.md convention without shipping a dedicated adapter for each.

Verified against shipped code, not documentation:
- Cursor 3.13's skill-root array literally contains ".agents/skills/",
  alongside .cursor/, .claude/ and .codex/ skill roots.
- @openai/codex 0.146.0 carries a full "AGENTS.md spec" in its base
  instructions and uses the adjacent ".agents/plugins" namespace.

Scope is deliberately narrow and asserted by tests: rules + skills only.
No subagent claim (no cross-tool convention exists; prompts reference
agents by path). No dashboard executor derivation — unlike cursor/codex
this target does not imply a specific CLI, so the dashboard stays on
claude unless dashboard.executor says otherwise.

Passes all 19 target-matrix invariants with zero test edits beyond
registration. 906 tests green, lint clean.

Refs TKT-005

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Documents all five targets — claude-code, cursor, codex, cline, agents-md
— with rules/skills/commands/agents paths, subagent support, and which
dashboard executor each derives.

The second table is the point: it records HOW each row was checked, not
just what it claims. Bobby has shipped wrong harness claims before (a
model name copied from a CLI's own stale --help, a "no subagent registry"
line a newer build contradicted), and both came from trusting docs. So
rows are marked Live / Live argv / Paths verified / Unverified, and cline
is explicitly labelled best-effort rather than quietly implying parity.

Also states the agents-md tier's limits plainly (rules + skills; no
subagent dispatch, no executor derivation) and adds a contributing
section pointing at cursor.js as the template and the matrix suite as the
acceptance bar, with the no-unverified-claims rule.

Verified: a script cross-checks every documented path, subagent flag, and
executor derivation against the actual adapters — all five rows match
shipped behaviour. 936 tests green across 48 suites.

Refs TKT-006

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Per-PR verification proves an executor's argv is correct at merge time and
nothing after; these CLIs ship weekly. A scheduled workflow (Mondays, plus
manual dispatch — never on push/PR) probes the real binaries with no
account required.

Mechanism: agent CLIs parse argv before checking auth, so Bobby's exact
command line run unauthenticated must fail on auth, never "unknown
option". A second control probe adds a bogus flag that MUST be rejected —
without it, a CLI that silently ignores unknown flags would produce a
vacuous pass.

argv comes from the real buildArgs functions. Verified by the AC: zero
flag strings appear in the workflow file, so the canary cannot drift from
the code it guards.

Building this caught a bug in the canary itself: resolveExecutor treats an
unrecognized name as a custom binary path with claude-style flags — correct
for `dashboard.executor: /path/to/x`, catastrophic here, since probing an
unregistered flavor built claude argv and reported it as codex drift.
Now guarded against EXECUTOR_NAMES with a distinct unknown-flavor outcome.

Live-verified against both installed CLIs: codex stops at authentication
as expected; cursor-agent clears argument parsing. A missing binary
reports install-failed, distinct from drift.

Failures file one idempotent GitHub issue per flavor (commenting on the
existing one rather than piling up weekly duplicates).

Refs TKT-014

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
CLIs, libraries, and npm packages silently stalled at the `test` stage.
bobby-test verifies by exercising a running application and its skill
forbids running the spec suite as a substitute, so on a project with
nothing to serve every test case came back BLOCKED and the ticket simply
stopped advancing — no error, no guidance.

- Built-in `library` (plan/build/review) and `library-secure` workflows.
  Review already runs the suite independently, which is the right
  verification for a library.
- `default_workflow` config key, and `stackDefaultWorkflow()` which honors
  an explicit `default_workflow` in stack JSON and otherwise infers
  `library` when a stack has no health checks and no dev command. `bobby
  init` writes the choice into .bobbyrc.yml with a comment explaining why,
  so it is visible rather than implicit.
- Live-app agents (test, qe, ux, pm, watchdog) invoked where there is
  nothing to observe now block the ticket with an actionable reason and
  name the fix, instead of emitting BLOCKED cases one at a time.
- `bobby workflow list` shows effective stages for overridden workflows
  (it printed the shadowed built-in ones) and marks the project default.

Dogfood proof: bobbycode's own .bobbyrc.yml drops its manual override for
`default_workflow: library`, and the loop still ends at review.

Refs TKT-013

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant