Flow-Next

Agents generate. Flow-Next proves.

Roadrunner: faster on the actual work

Faster than your agent alone.
And better.

A workflow plugin for your coding agent. It takes an idea, a ticket, a bug report or a slow page all the way to a pull request.

  • Faster delivery. The change comes back in about the time plain Claude Code, or your harness, takes, often less; large features in about half the time.
  • Better work. The way Flow-Next works already beats a plain agent on its own; cross-model review and live QA widen the gap to up to 25% better outcomes, especially on large and long-horizon work.
  • Fewer misses, less slop. Optional cross-model and cross-harness reviews catch what your agent misses and cut the slop. QA drives the real app, and every “done” comes with proof.
  • Your harnesses, working together. One routing block sends each job to the model or harness that fits it (reviews, workers and scouts across Claude Code, Codex, Cursor, Copilot, Grok and OpenCode), automatically.
  • Hands-off when you want it. Run it unattended: it decides, records every decision in the PR, and opens it ready for review, or as a draft when a call is left for you.
$ /flow-next:flow "the /reports page takes four seconds, it should take under one"
fn-1 · full CSV export illustrative run · review + QA on

How flow-next turns a request into a pull request you decide on

Illustrative run with review and QA gates on. A messy support message asking for a full CSV export goes into one command, /flow-next:flow, which shapes it into a spec with three acceptance criteria. An agent writes the code. A reviewer from a different model family finds a bug, the last page of the export is never written, and sends the work back with NEEDS_WORK. The fix adds a test for the last page and passes re-review with SHIP. An automated QA agent then clicks through the running app, like a user would, and all 1,000 rows export. A pull request opens with its evidence and a DONE receipt. Flow-next stops there; merging is your call.

In the field 3x speed-up across the entire R&D organisation, from discovery to merged code, at a higher quality of output: the reported average from teams running the engineering methodology built on Flow-Next, 1,000+ developers across a private-equity portfolio. How teams run it →
One paragraph of policy, nobody in the loop 38 PRs landed from one unattended run: each one planned or worked directly by judgment, reviewed by another model family, QA'd in the running app, CI-green, and merged with a receipt. The field case →
Against plain Claude Code, or your harness Up to 25% better outcomes on large and long-horizon work, with the change back in about the same time, often less. Same model, more than 170 full end-to-end runs, each case run several times. Better even before review; review and QA, added where the risk calls for them, widen the gap. The 7.0 results →
Claude CodeOpenAI CodexFactory DroidGrok BuildCursorRepoPromptGitHub CopilotOpenCode

What changed

Roadrunner: faster on the actual work, and better outcomes.

The latest update, and the three preceding notable releases. The changelog carries the whole history.
7.0.0 Roadrunner

Faster, better outcomes, one unattended mode.

A small fix is changed directly with no spec, a single-task spec is built in your conversation, and review goes by risk: three reviewers from another model family for changes to persisted or shared state, concurrency, security, data layout or several files, one for a small local fix, none for a wording or display change. Unattended, flow --auto never asks, keeps fixing until review says SHIP, and writes every decision into the PR. It is now the one unattended mode: Ralph and the HTML render lenses are gone.

Route taken
  1. change directly
  2. check it like a user
  3. hand it back
  4. review in the background
no spec · 1 reviewer · 1 re-review of the fixes
  1. 6.7.0 Codex reviews on a newer model Codex reviews and agents start on GPT-6.1 Sol and step down when your account cannot serve it; Claude reviews gain two more fallbacks.
  2. 6.6.0 Faster scouts, and gates that see every repo Plan and prime run their scouts on a lower-cost model close to the top tier. When your code lives in sibling repos, a sibling code change always runs the full gates.
  3. 6.5.0 Fewer questions before the build Refine asks only what would change what gets built and only you can settle, and flow sends you there only when such a decision is open. Review, QA and the pull request evidence are unchanged.
  4. All release notes

One command, or the stages beneath it

Say what you have. Flow runs the stages that fit.

The first command reads what you hand it and runs the right ones from the four below, stopping at the decision that is yours. Each of the four is also a command you can run yourself when you want that control. Walk it once in 30 minutes →
  1. The conductor

    /flow-next:flow ("this endpoint is slow" / "what should I do next")

    Hand it anything: an idea, a spec id, a tracker issue, a pasted bug report, a slow page, a why question. It reads the content and the context and takes the lightest route that fits: a small fix is changed directly, with no spec and no ceremony; a feature runs through the stages below, capture to work to the checks to the PR. It stops only at a decision that is yours: a product choice the spec leaves open, or the pull request itself. Run it with nothing after it and it continues from the most recent item. Add --explain to read the route without running it.

  2. The stages it chooses from, which you can also run by hand if you prefer living in 2025

    /flow-next:capture ("capture what we discussed as a spec")

    Flow's first stop for an idea, or a ticket that describes one. Turns what you gave it into a spec with source-tagged acceptance criteria, saves it and offers the file in your editor, then names the next step and why.

  3. /flow-next:work fn-12 --no-plan ("work fn-12 on a new branch, review with codex")

    Flow's default for a ready spec. A single-task spec is built right in your conversation (add --worktree for an isolated one), reviewed by risk, and every done carries commits and test evidence.

  4. /flow-next:qa fn-12 ("QA fn-12 against the running app")

    Drives the running app from the spec and files findings with screenshots. It cannot pass by reading the source. With pipeline.qa set to auto, Flow runs it when the criteria describe UI behaviour, the surface is drivable, and a target can start; otherwise it records the skip.

  5. /flow-next:make-pr fn-12 ("open the PR")

    Flow ends its run here. A pull request that explains itself, with merge left to you. When you want a loop to carry it the rest of the way, land is a separate driver, described in the factory below.

Whatever you have

One command, many starting points.

A few rows from the route matrix, and the first move for each: Every route, with a diagram →
  1. A small fix

    Changed directly, checked the way a user would, and handed back fast. No spec needed.

  2. An idea, or a ticket that describes one

    Capture writes a source-tagged spec and shows it to you.

  3. A bug report or a stack trace

    Check for prior fixes, reproduce and confirm the cause, bisect when it can. The failing test becomes the acceptance criterion; the fix must pass on head where it failed on base.

  4. A page that takes four seconds

    Measure first on the real page, endpoint, or command. The baseline and the target become the criterion.

  5. A benchmark to push

    Repeated attempts against a frozen harness, each attempt's measurement recorded as evidence.

  6. A rename that must keep behaviour

    Pin the current behaviour with a characterization test, then change the structure under it.

  7. A why question about the code

    A cited answer from the repo, its history, and memory. Nothing is written.

  8. A spec with an open PR

    Resolve the review threads and CI, then stop when merge is the only step left.

What you get

What changes when you run it.

Decide what to build, build it, and verify the result. Shape the process around your work.

Everything reaches your queue checked as hard as its risk calls for.

Review goes by risk. Changes that touch persisted or shared state, concurrency, security or data layout, and multi-file features, get three reviewers from another model family; a small change in one area gets one, and a small wording or display fix skips review and records why. Turn review off and the run says so in its report.

Attended, you get the result first, then one review and one re-review of the fixes. Unattended, it keeps fixing and re-reviewing until SHIP. Every done carries commits and test evidence.

review (by risk)  3 reviewers
first pass        NEEDS_WORK
fix + re-review   SHIP
flowctl done      evidence recorded

Open a PR that already makes its argument.

The reviewer's first screen is the reasoning behind the change and the criteria it claims to satisfy.

The pull request arrives explaining itself: which acceptance criterion each change satisfies, what is still open for a human, what deliberately did not change. An unattended run adds every decision it made on your behalf, with the evidence.

GitHub PR
  Why · User impact · Scope
  Blast radius · Verification
  Tradeoffs + Decisions · Open items

Decide what to build before anyone builds it.

An idea too big to write down gets charted one decision at a time; a conversation becomes a spec; a product owner and an engineer refine it in their own sessions on one file, when a decision is open.

Chart resolves decisions against evidence, capture writes source-tagged criteria with a read-back, refine (interview) runs one question pass on the spec, focused by whoever runs it, or a read-only research pass, prospect ranks what to build next.

/flow-next:chart      resolve D-IDs
/flow-next:capture    spec + read-back
/flow-next:refine     any scope | research

Your team's context lives in the repo.

The reasons behind the code sit next to the code. A new teammate, or the next agent run, starts from what the last one learned.

A review correction becomes a lesson the next task can read. Specs, decisions, glossary, and memory stay in your repository.

.flow/specs/    intent + criteria
.flow/memory/   lessons that stuck
GLOSSARY.md     your domain words

Prove it in the running app, not by reading the source.

Live QA drives the app the way a user would, from the spec's own criteria, and files what it finds with screenshots and a verdict you can audit.

The QA skill is forbidden from passing by reading code; a committed feature map records how each surface is reached so every run starts warm.

/flow-next:qa fn-12
  P0 0 · P1 1 · P2 2
  6/6 R-IDs exercised · SHIP

Hand over as much as the receipts have earned.

Start by watching a single task run. Move up a rung when the receipts have earned it, and step back down whenever you want.

One dial from a supervised pair to a loop draining the backlog overnight. The same gates apply at every rung, and each records its verdict or its skip. Unattended, it never asks: it decides from evidence and writes each decision into the PR.

/flow-next:flow fn-12        you watch
/flow-next:flow --auto       it runs unattended
a bot or a cron              it repeats

Choose the model for each job.

Name a model per role once in your CLAUDE.md, or say it in the prompt for a single run. Pick another family for review and the model that wrote the diff never reviews it.

A tier line in your instruction file, a --review= flag, or a sentence in the moment. Leave it unset and everything runs on the session model.

spec    -> the model you trust for judgment
work    -> a capable model for this spec
review  -> a different family

One way of working, from a solo Sunday to a rollout.

The same rails carry a solo developer on a Sunday and a fifty-person organisation on a rollout. The spec is the handover object, and it reads the same to product, engineering, and the next agent run.

Product and engineering refine the same file, each with their own focus, tracker projection to Linear, GitHub, GitLab or Jira, and a one-hour discovery interview that yields eight to eleven implementation-ready specs.

product  -> Goal, Boundaries
engineer -> Architecture, Edge cases
both     -> Acceptance Criteria

Your process outlives your agent.

Switch harness, model, or vendor and nothing has to move. The specs, the gates, and the record of what happened are files in your repository. In a harness that can dispatch subagents, the same routing runs across models in-host with no bridge at all.

The same specs, gates, receipts, and task state across harnesses. Everything in your repository, and nothing outside it.

Claude Code  -> write the spec
Codex        -> implement
Cursor       -> review

same .flow/ state, every role
uninstall:   rm -rf .flow/

The factory

Bless it. Loop it. Ship it.

Your judgment lives in the spec and in one paragraph of policy. Two loops read both, choose the shape per spec and the model per job, and print the reason every time. Nothing merges that the gates did not pass.

Flow --auto, the build loop

One invocation carries one blessed spec hop after hop to a pull request, ready unless something is left for a human, and ends with a verdict your host can act on; --tick runs one hop for hosts that loop.

/flow-next:flow --auto
  • Classifies each hop from the same routing references the attended conductor reads: plan, plan review, work, QA, make-pr.
  • Decides per spec whether to plan first, and prints why.
  • Never asks: it decides from evidence, keeps fixing until review says SHIP, and writes every decision into the PR. With --until=merge it merges after a 10-minute window for review bots.
  • Backlog mode triages the whole board in dependency order, one item per run.

Land, the ship loop

Babysits the one pull request you name until it merges, on the repo's own rules and with your authorization.

/loop 15m /flow-next:land 123
  • Makes one CI fix or one flake rerun per failure, then stops and says so.
  • Resolves review threads through resolver agents, replies with the fix and the line.
  • Squash-merges with a head-commit match only when you authorized it, and writes nothing to the repository afterwards.

An always-on factory is one process

Something long-running that listens for a blessed spec, or a merged PR, or a cron tick, and calls flow --auto; land ships what flow --auto opens. A host loop, a goal, or any scheduler with a shell is enough. Every verdict is typed, so the process reads NEEDS_HUMAN and stops instead of spinning. Forge, below, is one built that way.

/goal keep running /flow-next:flow --auto until PILOT_VERDICT=NO_WORK

Two routing axes, both printing their reason

The pipeline shape per spec and the model per job are decided separately. Override either with a flag, a field, or a sentence; leave both alone and it still runs.

Shape per spec

Recommended next: plan
Scheduling: rolling
Triage-skip: docs-only
PILOT_VERDICT=ADVANCED

Model per job

reviewer: family B
implementer: fast tier
1 or 3 reviewers by risk
receipt: model recorded
# your prompt, or a paragraph in your instruction file
Work the ready specs. Decide per spec, from its Boundaries
and the size of its Touches, whether to plan first or work
directly; a named open product decision goes to refine. Auth and
the migration you implement yourself; plain CRUD goes to the
implementer tier over a bridge. Reviews from a different
family either way; NEEDS_WORK twice, stop bridging.
hop 1 fn-61 · plan Recommended next: plan - staged across two PRs
hop 2 fn-62 · work no_plan - fully known, low blast radius
hop 3 fn-61 · work Scheduling: rolling · 3 lanes
hop 4 fn-63 · work review inside: NEEDS_WORK → fix pass → SHIP
land #172 · CI green 5/5 threads · merged
hop 5 fn-64 · qa 6/6 R-IDs exercised · SHIP
hop 6 fn-64 · make-pr ✓ PR #173 (ready)
hop 7 PILOT_VERDICT=NO_WORK spec=- stage=- reason="no ready spec with satisfied deps"
hop 1 fn-61 · plan Recommended next: plan - staged across two PRs
hop 2 fn-62 · work no_plan - fully known, low blast radius
hop 3 fn-61 · work Scheduling: rolling · 3 lanes
hop 4 fn-63 · work review inside: NEEDS_WORK → fix pass → SHIP
land #172 · CI green 5/5 threads · merged
hop 5 fn-64 · qa 6/6 R-IDs exercised · SHIP
hop 6 fn-64 · make-pr ✓ PR #173 (ready)
hop 7 PILOT_VERDICT=NO_WORK spec=- stage=- reason="no ready spec with satisfied deps"
Sculpt it into the factory you need A bug-fixing factory that reads the issue tracker, a feature factory that drains a blessed backlog, a docs factory that runs the lint-only tier. Same loops, a different paragraph.
You set the gates a change must pass Green CI, every review thread resolved, the review decision your branch protection requires, a live QA verdict, a merge command of your own. Land merges when all of them hold and you authorized the merge, and never before.
It knows when to stop and ask A design conflict, a red build it cannot fix within its budget, or a question only you can answer ends the run with a named verdict and waits for you. It never spins, and it never guesses on your behalf.
Built on it: Forge, a factory-manager bot A Grok bot that wakes on a push, a tracker change or a schedule, classifies each ready spec from its git state, hands the named flow-next job to a cloud agent or a CLI, and stays with it until land ships. Forge ↗ · by Daniel Killenberger ↗

The checks

Context makes the work right. The checks make it provable.

Each configured check records its verdict, and a skipped stage records its reason. The bar is yours. The checks hold it at every rung of autonomy.
Plan review Optional spec and design review before tasks exist. A different family checks the approach; findings are fixed once and re-reviewed once. /flow-next:plan-review
Implementation review Three reviewers when a change touches persisted or shared state, concurrency, security or data layout, or spans several files; one for a small local fix; none for a wording or display change; then a re-review of the fixes only. Unattended runs repeat until SHIP; the verdict and model go in a receipt. /flow-next:impl-review
Live QA Drives the running app from the spec; P0/P1/P2 with evidence; cannot pass by reading source. On auto, it runs when the criteria describe UI behaviour, the surface is drivable, and a target can start. /flow-next:qa
Completion review A multi-task spec against its criteria before a PR opens; every R-ID accounted for or named as uncovered. /flow-next:spec-completion-review

Why the checks exist

What happens on the twentieth change

SlopCodeBench puts 11 models through 93 sequential checkpoints. Not one solved a problem end to end; the best strict pass rate fell from 17.2% to 0.5%, and quality-aware prompts changed the rate of decay not at all. The authors name one thing as untested: structural discipline enforced across checkpoints, through tooling. That is what Flow-Next's gates are built to hold. The study did not test them.
0.68 agent 0.31 human
Code-quality erosion across 93 checkpoints · 48 maintained human repositories · arXiv 2603.24755 ↗

Where it already runs

From CAD suites to thirty-year-old legacy stacks.

Enterprise engineering organisations worldwide: CAD and construction software, proptech, education. Modern monorepos and hundred-repo microservice estates beside legacy systems three decades old, one of them migrated with these rails. GitHub Enterprise, GitLab, Jira. Optimised for Windows too, because the field runs Windows.
CAD and construction softwareProptechEducationHundred-repo microservice estatesA thirty-year-old legacy stack, migratedGitHub Enterprise · GitLab · JiraWindows in the field
“I am enjoying your version of all these cool new plugins. So far yours has worked the best.”
@patrickmichalina · issue #5 ↗
“Hello, really enjoying this project, thanks for making it and making it public (also huge compliments on your website!)”
@possibilities · PR #95 ↗
“it’s been really useful in my workflow.”
@raydocs · issue #4 ↗

The official engineering framework and methodology across the portfolio of one of the largest private-equity funds: 1,000+ developers, 100+ product owners and product managers, and 30+ QA engineers onboarded. Evidence and evals →

The capabilities, one page each

Everything it does, on one screen. Flow chooses among them for you.

Most cells are stages Flow routes to on its own; flow --auto and land are the drivers that run without you. Each is a command you can run yourself, and each cell opens its page.
Flow Hand it anything. It picks the route from the cells below, runs it, and stops at your decision. The one command a first run needs. /flow-next:flow Chart Resolve one decision at a time until an idea is briefable. /flow-next:chart Capture Conversation to spec, source-tagged, with a read-back. /flow-next:capture Refine (interview) One question pass on the spec, optionally focused by a scope, or a read-only research pass. /flow-next:refine Prospect & Strategy Ranked candidates grounded in the repo; the strategy anchor they answer to. /flow-next:prospect Plan Optional task graphs when you ask for one, tasks have separate owners, or delivery comes in staged PRs. /flow-next:plan Work One task inline, several in parallel lanes, reviewed by risk. /flow-next:work Reviews Plan, implementation, completion; cross-family, receipted. /flow-next:impl-review Live QA & Drive Run the app like a user; file what breaks with evidence. Off, on, or auto. /flow-next:qa Make PR & Resolve PR A PR that argues its case; review threads resolved by agents. /flow-next:make-pr Flow --auto & Land The build loop that carries one ready spec to a PR without asking, and the ship loop that merges it. /flow-next:flow --auto Model routing Four tiers, a routing block, bridges to other CLIs. reviewer: <model> Tracker sync Project specs to Linear, GitHub, GitLab or Jira. Projection, never coordination. /flow-next:tracker-sync Memory & Audit Lessons that stuck, audited against the code that moved. /flow-next:audit Prime & Features Agent-readiness assessment; a user-POV feature map QA reuses. /flow-next:prime Visual & Prose See any spec, plan or diff visually; one prose contract for every artifact. /flow-next:visual Map A semantic feature index scouts and prime anchor to. /flow-next:map Plan-sync Downstream task specs updated when the code moved under them. /flow-next:sync Deps & Worktrees Execution order across specs; isolated worktrees for parallel work. flow-next-deps Export for external review Hand a plan or diff to any model outside the harness. flow-next-export-context Setup & Uninstall A copy-less install; removal is one directory. /flow-next:setup
All skills, with what each prints →

Pick your path

Start where you are.

Run one change to a pull request, preview a route with your team, or deploy the plugin across an organisation. One conductor, one set of gates.

You are solo and want to feel it today.

/flow-next:flow "add a full CSV export to the reports page"

Flow captures it, works the spec, gets it reviewed, and opens the PR, stopping only where a decision is yours. The first-run guide shows the output to expect at every step.

Your first 30 minutes →

Your team is adopting it together.

/flow-next:flow --explain fn-3

Product and engineering read the same route and the same reason. The spec stays the shared artifact: product fills intent, engineering fills constraints, reviewers read one handover.

How teams run it →

Your organisation is rolling it out.

managed-settings.json

Deploy it once through Claude Code managed settings and every developer has flow on next launch, with the same gates in every repo.

Org-wide deployment →
New to Flow-Next? Get your first spec worked, reviewed, and handed off. Everything lives in your repo; your specs and evidence stay readable when you stop using the plugin.