Four hand-written WAT modules. No compiler is involved, deliberately: the
question is about the design in doc/design/0004-dispatch-design.md, and a
compiler between the design and the number would only add doubt.
Predictions were recorded in doc/design/0002-measure-first.md before the
first run. Do not read them until you have your own expectation.
| Module | Measures | Decides | |
|---|---|---|---|
| B1 | b1_protocol.wat |
vtable-slot dispatch: 3 loads + call_ref, monomorphic site |
Whether the design is viable at all |
| B1L/B1i/B1c | (same module) | controls: the same walk at 1 load, 0 loads, and no dispatch | Where in B1 the cost actually is |
| B2 | b2_megamorphic.wat |
the same site with 10 receiver types, plus a depth-matched one-type control | Whether we beat the JVM where its per-call-site cache thrashes |
| B3 | b3_arith.wat |
i31 fast-path add vs a boxed slow path |
Whether boxed arithmetic can be cheap |
| B4 | b4_cast.wat |
ref.cast by target depth and by input variety, against a no-cast floor |
How to shape the type graph |
| B5/B5x/B7 | (in b2_megamorphic.wat) |
guarded call-site specialisation, and the hit rate at which it starts paying | Whether the server lane is reachable at all |
| B6 | b6_boundary.wat |
lowering a 4 KB aggregate across the component boundary, against memory.copy |
What 0007's finding costs, and the representation lever (0008) |
| B8 | b8_boxed.wat |
the boxed-i64 lane: allocation per op, mixed dispatch, canonicalization probe, and what a real slow path costs B3's fast path | The numeric representation past i31 (0022 C, 0025) |
Each runs on both node (V8: speculative inlining) and wasmtime
(no adaptive tier). Both numbers are reported; neither is "the" answer.
B1–B3 are compared against JVM Clojure equivalents in bench/s0/jvm/, run on
the same machine in the same session. A cross-machine comparison is not one.
bb bench-s0 # everything, at the sizes the numbers were taken at
bb bench-s0 B1 # one benchmark
bb bench-s0 --n 2000000 --reps 5 # a quicker, noisier pass while editingNeeds wasmtime, wasm-opt, node and wasm-tools — nix develop has them
at the pinned versions. The driver prints the machine, the tool versions and the
command to reproduce, which is what gets pasted into
doc/design/0002-measure-first.md alongside the numbers.
Every lane's result is checked against the value the benchmark is supposed to
compute, and a mismatch stops the run. The driver also refuses an n whose
expected answer is the one an empty loop would return — for a ring walk that is
any multiple of the ring length, and without the check a benchmark doing no work
passes. Both guards were confirmed by breaking the benchmark on purpose.
Controls are not optional here. A benchmark and the variant you subtract
from it usually differ in more than one way, and the difference is then not an
attribution — see .claude/rules/measurement.md, and the B1 incident in
doc/status.md that put the rule there.
bench plus whatever controls it needs, all with the signature
(i32) -> i32 — take an iteration count, return a value that depends on every
iteration. No imports, so the same file runs unchanged under node and
wasmtime run --invoke. The JVM counterpart lives in jvm/, takes
<variant> <n> <reps> <warmup>, and prints EDN.
wasmtime has no clock in the guest and no way to loop across invocations, so
its lane is timed by process slope: wall time at n and at 2n, difference
over n. Startup, compilation and instantiation are identical in both and cancel.
Read .claude/rules/measurement.md. The failure modes that matter — the
optimizer deleting unconsumed work, process-spawn overhead swamping short runs,
warm-up effects — all produce plausible-looking wrong numbers rather than
errors.
All measured — B1–B4 as contracted, B5 and B7 which B1's and B2's findings
forced, B6 (the component boundary, measured 2026-07-30), and B8 (the
boxed-i64 lane, S3's numeric-representation benchmark). Numbers, controls and
what each means are in doc/design/0002-measure-first.md; the S0 verdict is
doc/design/0010-*, the B8 decision doc/design/0025-*.
B5 and B7 live inside b2_megamorphic.wat rather than their own files so that
their type graph and rings are provably identical to the baselines they are
compared against. That is deliberate: comparing across files is how three of
this project's controls went wrong.