A testable design-production loop for AI coding agents. It turns "make it look good" into a build → verify → fix loop with a pass/fail exit, so design quality is checked, not vibed.
Built as a Claude Code skill, but the method is agent-agnostic — see Using it with other agents.
The loop freezes a versioned bar from a real reference, protects an immutable house system, optionally shops curated component inventory for proven raw material, builds a piece, runs deterministic gates (task-success + tokens + a11y + visual-regression) before any model judges it, then runs fresh-context critics on screenshots and recorded interactions — repeating until every required gate and critic passes against the frozen bar.
Two rules run through everything:
- Deterministic gates run before model judges. Don't spend the strongest model judging a failure a linter or test can prove. A mechanical failure is a hard fail.
- Critics judge evidence only — rendered screenshots and recorded interactions, never code, prompts, model identity, or prior verdicts. Verdicts are binary.
Most "AI design" workflows generate confidently and evaluate poorly — the same model that built the thing also grades it, from the code, on a bar it invented. This loop separates the roles: an immutable bar it can't move, deterministic gates it can't argue with, and independent critics with fresh context that see only the rendered result. Anything that passes becomes a candidate asset; it's promoted to a reusable building block only after clean repeat runs and a human spot-check.
EXAMPLE.md walks a full condensed run — a pricing section looped against Linear — from frozen bar through two critic-caught defects to a pass. Read that first if you want to know what using this actually feels like.
Three landing pages, three invented products, one basic prompt each ("make a beautiful site for a made-up product"). Left is the first pass, before the loop. Right is after the critic panel refused it — as many rounds as it took — until an independent, harsh eye called it genuinely great. Same prompt every time; the only variable is the loop.
Borealis — an aurora forecast. The first-pass aurora was a faint, structureless smear; the panel refused it four times (twice "genuinely great, but doesn't beat the bar") before the hero became real drifting curtains and a ribbon of aurora light threaded the whole page.
Volta — a modular synthesizer in a browser. Four rounds took it from a generic dark-SaaS hero to a committed instrument: an editorial serif headline, brushed-metal modules, and a live signal threading every section.
Meridian — custom star maps. Already strong on the first pass, so the loop's catches were subtler — here, the night-sky atmosphere failing to carry past the hero.
Clone straight into your skills directory:
git clone https://github.com/devonbweaver/design-loop ~/.claude/skills/design-loopOr clone anywhere and symlink (keeps git pull easy):
git clone https://github.com/devonbweaver/design-loop
./design-loop/install.shRestart Claude Code, then:
/design-loop
or just ask: "run the design loop against <a reference URL>".
The loop is a methodology, not Claude-Code plumbing. SKILL.md is the playbook;
only the packaging differs:
- Claude Code — clone into
~/.claude/skills/, invoke/design-loop. Its subagents give each critic genuinely fresh context and let you route the critics to different models (Brief → a mid model, System → a small one, Craft → the strongest). - Cursor / Windsurf / Cline / Continue — drop
SKILL.mdinto your rules or system prompt (e.g..cursorrules, a project rule file). Run each critic as a separate fresh chat, given only the screenshot + the frozen bar. The fresh context is the point, however you get it. - Aider / Codex / plain API — paste
SKILL.mdas the system prompt and script the fan-out: one call per critic, each with a clean context.
The only hard requirements are agent-agnostic:
- a way to render the output (a screenshot; for motion, a frame filmstrip + a
prefers-reduced-motionrun) — no render means a blind critic, and - a way to give each critic a fresh context (a subagent, or just a new chat).
The deterministic gate (Playwright + a linter) needs no agent at all.
- A specific reference — one page/screen that does the thing brilliantly. A vague bar ("good SaaS design") is the #1 reason the method fails.
- The ability to render your output and, for interaction work, to record it.
- Optional: Playwright for the visual-defect + conformance gate and interaction capture.
SKILL.md # the loop: 8 phases, two governing rules
EXAMPLE.md # a full condensed run, start to pass
references/
freeze-and-versioning.md # freeze the bar; immutable house vs derived bar; provenance
reference-and-brief.md # set a bar the critic can enforce: region-vs-point, binding brief
craft-standards.md # what makes UI genuinely beautiful: the taste the Craft critic holds
visual-defect-gate.md # the floor: visual brokenness only (overflow, contrast, jank), not the goal
inventory.md # the real DESIGN.md / component catalog to shop from before generating
critic-escalation.md # match verdict cost to change risk (tiers 0–4) + cost controls
motion-bar.md # motion token baseline + binary motion rules
interaction-loop.md # recorded-interaction evidence + scenario matrix
build-from-inventory.md # shop proven component patterns, then gauntlet them
judge-reliability.md # the Motion critic + keeping the panel honest
MIT — see LICENSE.


