Skip to content

Repository files navigation

node-forge-eval

A benchmark for the one thing LLMs provably can't do in 3D — wire a node graph. Every socket is checked against a real Blender node catalog, every link type‑checks, cycles are caught, and the wrong answers are labeled.

ci deploy-dashboard  ·  ▶ Live dashboard  ·  code MIT · data CC BY 4.0

node-forge-eval dashboard — an interactive React Flow node graph with the offending node/link highlighted red and the validator's INVALID verdict

Results at a glance (3-task slice) — the expert graphs are 100% valid, the naive-llm corruptions 0%, and every verdict resolves sockets/types against the real Blender catalog. The dashboard renders the graph and highlights the exact offending socket/link.

Node-based tools (Geometry Nodes, shader graphs) require graph reasoning that text‑only models fail at: they invent sockets that don't exist, plug Color into Vector, and can't even detect a cycle in a text‑described graph (GPT‑4 measurably can't). Studies show models read graphs at ~90% but generate valid ones far worse — which is why the Blender‑MCP ecosystem exists at all.

node-forge-eval turns that into a machine‑graded benchmark. Each task ships a correct graph and a set of corruptions — each reproducing a documented LLM mistake — and a validator that checks every candidate against the real Blender node catalog (catalog.json).

Why this is hard to fake: the validator resolves every socket against the actual node definitions and type‑checks every link with Blender's real conversion rules (scalars↔vectors convert; Geometry is strict). An invented socket or a Geometry→Float link fails deterministically — no rendering, no guessing.

What's in the box

nodeforge (pure stdlib)
  catalog  ── the Blender node-definition catalog (sockets + data types)
  validator── socket existence · type compatibility · cycles · disconnected outputs
  corrupt  ── auto-generate wrong graphs (invented socket / type mismatch / cycle / disconnect)
  runner   ── grade a candidate graph
  score    ── aggregate -> results/results.json (leaderboard, corruption gallery)
  export   ── AI-training data (records + chosen/rejected graph pairs)

dashboard (Next.js + React Flow)
  interactive node-graph viewer: pick a candidate, see the graph, with the offending
  socket/link highlighted and the validator's verdict.
Task Corruptions it ships
01-scatter-on-surface invented socket · type mismatch (Geometry→Float)
02-transform-array cycle · disconnected output
03-noise-displace type mismatch (Geometry→Vector) · invented socket

Quickstart

make build        # worker image (python)
make run          # grade every graph -> results/results.json  (valid passes, corrupt fails)
make test         # validator unit tests + invariants
make export       # dataset/{records,preference_pairs}.jsonl
make dashboard    # static site: React Flow node-graph explorer

Everything is pure Python + JSON — no Blender needed to grade or view. The catalog is grounded in real Blender 4.x node definitions and regenerable via a Blender export script (tools/, roadmap).

Design principles

  • Proven, not asserted. Every verdict comes from resolving sockets/types against the real node catalog and analyzing the graph topology.
  • Contrastive + self-mining. Wrong candidates are auto-generated by corrupting the correct graph — clean chosen/rejected training pairs.
  • Roadmap: a Blender render‑diff stage to catch graphs that are structurally valid but semantically wrong; auto‑export of the catalog from a live Blender.

Security

No credentials are needed or committed; the validator only parses JSON (no code execution, no network). See SECURITY.md.

License

Code: MIT. Benchmark data (tasks/, catalog.json, dataset/): CC BY 4.0.

About

A benchmark for the one thing LLMs can't do in 3D — wire a node graph. Validates Geometry-Nodes graphs (sockets/types/cycles) against a real Blender catalog, with a React Flow viewer.

Topics

Resources

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages