A benchmark for the one thing LLMs provably can't do in 3D — wire a node graph. Every socket is checked against a real Blender node catalog, every link type‑checks, cycles are caught, and the wrong answers are labeled.
· ▶ Live dashboard · code MIT · data CC BY 4.0
Results at a glance (3-task slice) — the
expertgraphs are 100% valid, thenaive-llmcorruptions 0%, and every verdict resolves sockets/types against the real Blender catalog. The dashboard renders the graph and highlights the exact offending socket/link.
Node-based tools (Geometry Nodes, shader graphs) require graph reasoning that text‑only models fail at: they invent sockets that don't exist, plug Color into Vector, and can't even detect a cycle in a text‑described graph (GPT‑4 measurably can't). Studies show models read graphs at ~90% but generate valid ones far worse — which is why the Blender‑MCP ecosystem exists at all.
node-forge-eval turns that into a machine‑graded benchmark. Each task ships a correct graph and a set of corruptions — each reproducing a documented LLM mistake — and a validator that checks every candidate against the real Blender node catalog (catalog.json).
Why this is hard to fake: the validator resolves every socket against the actual node definitions and type‑checks every link with Blender's real conversion rules (scalars↔vectors convert; Geometry is strict). An invented socket or a Geometry→Float link fails deterministically — no rendering, no guessing.
nodeforge (pure stdlib)
catalog ── the Blender node-definition catalog (sockets + data types)
validator── socket existence · type compatibility · cycles · disconnected outputs
corrupt ── auto-generate wrong graphs (invented socket / type mismatch / cycle / disconnect)
runner ── grade a candidate graph
score ── aggregate -> results/results.json (leaderboard, corruption gallery)
export ── AI-training data (records + chosen/rejected graph pairs)
dashboard (Next.js + React Flow)
interactive node-graph viewer: pick a candidate, see the graph, with the offending
socket/link highlighted and the validator's verdict.
| Task | Corruptions it ships |
|---|---|
01-scatter-on-surface |
invented socket · type mismatch (Geometry→Float) |
02-transform-array |
cycle · disconnected output |
03-noise-displace |
type mismatch (Geometry→Vector) · invented socket |
make build # worker image (python)
make run # grade every graph -> results/results.json (valid passes, corrupt fails)
make test # validator unit tests + invariants
make export # dataset/{records,preference_pairs}.jsonl
make dashboard # static site: React Flow node-graph explorerEverything is pure Python + JSON — no Blender needed to grade or view. The
catalog is grounded in real Blender 4.x node definitions and regenerable via a
Blender export script (tools/, roadmap).
- Proven, not asserted. Every verdict comes from resolving sockets/types against the real node catalog and analyzing the graph topology.
- Contrastive + self-mining. Wrong candidates are auto-generated by corrupting the correct graph — clean chosen/rejected training pairs.
- Roadmap: a Blender render‑diff stage to catch graphs that are structurally valid but semantically wrong; auto‑export of the catalog from a live Blender.
No credentials are needed or committed; the validator only parses JSON (no code execution, no network). See SECURITY.md.
Code: MIT. Benchmark data (tasks/, catalog.json, dataset/): CC BY 4.0.
