Write Python or JavaScript. Run natively or in WebAssembly. Give supported compute to the GPU.
Run the GPU lab · Write GPU code · Read the docs
JavaScript engine · Python frontend · WebAssembly runtime · GPU compute · Test262 conformance · Benchmarks · Architecture · Engineering journal
Quick start · What runs where · Performance · Language support · Security · Contribute
Zipp brings Python and JavaScript into the same Rust register-bytecode VM. Build an embedded scripting runtime, run a folder of code in a browser Worker, or turn a familiar Torch model into a browser GPU compute graph.
The experimental Python frontend compiles directly to Zipp bytecode. It does not ship CPython or translate your Python program into browser JavaScript. Both frontends use the same engine, with native and WebAssembly builds.
- Bring a project, not just a snippet. The playground loads folders, modules and data, with an editor, virtual files, console and graphics in one place.
- Start with familiar ML code. The bundled Torch subset supports eager CPU
tensors, autograd and training. Experimental
torch.compile(model)records dense inference and training steps for WebGPU, WebGL2, WebAssembly SIMD kernels or an explicit CPU fallback: relu, gelu, sigmoid or tanh layers, softmax, MSE or a fused cross-entropy, and SGD, momentum or Adam updates in one graph. Thezipp_gpugraph protocol underneath (version 4, with masks, dropout, slicing and gathers) also carries rank-4 broadcasting and batched matmul, so an MNIST-scale step written as a graph runs in one submission (7.8 ms on WebGPU, 3.9 ms on the WebAssembly kernels, for 784-256-10 at batch 64) - and as a prepared session, with the weights and optimizer state resident on the device and eight steps per submission, 0.79 ms per step (1.8 ms per step when driven from Python throughcompiled.prepare). Natively,zipp pyruns the same graphs on a hardware GPU through wgpu: on an RTX 5090 that small model's prepared step takes 0.27 ms (0.12 ms at eight steps per run), against about 0.78 ms for PyTorch's eager CUDA. - Run a language model from its GGUF file. Graph inputs can stay in a checkpoint's own Q4_K and Q6_K blocks, decoded inside the matmul on every backend, so Qwen3-0.6B holds at 373 MB instead of 2,274. The experimental, source-only local-model plugins run it from the file, in the browser or split across several machines, and Hugging Face's own Qwen3, Qwen2, Llama, Mistral and GPT-Neo modelling files run on the Python frontend unmodified.
- Keep the host in control. Execution budgets and explicit host capabilities let embedders decide which resources a program can use.
- Explore one engine across languages. An optional trusted-code build adds Python-to-JavaScript evaluation inside the very same VM instance.
Python support and Torch compatibility are experimental. The Torch compatibility guide and Python frontend guide explain the supported surface and remaining differences.
| What you want to do | Where your code runs | Where the numerical or drawing work runs |
|---|---|---|
| Run JavaScript scripts or embed Zipp | Native Rust VM/JIT, or Zipp's WebAssembly build | CPU; host APIs are supplied by the embedding application |
| Run Python projects in Zipp | Experimental Python frontend on the same VM | CPU, including the bundled torch subset |
Compile a supported Torch model or submit a zipp_gpu graph with zipp py |
Python in the native engine; gpu-lab's runtime in a second engine state | A hardware GPU through wgpu (Vulkan, Direct3D 12, Metal), synchronously; the CPU evaluator when there is none |
Compile a supported Torch model or submit a Python zipp_gpu graph from the playground |
Python in the WASM Worker; JavaScript handles the graph | WebGPU compute shaders or WebGL2 fragment shaders; visible CPU fallback in auto mode |
| Use GPU graphs from browser JavaScript | An ordinary browser ES module or Worker | The same GPU runtime, without requiring Python or the Zipp VM |
| Run a local language model from a GGUF file (experimental, source only) | A model plugin's Python in the WASM engine; JavaScript reads and binds the checkpoint | The same GPU runtime, with the weights kept as the file's quantized blocks |
| Draw a custom browser visualization | Browser JavaScript with a canvas | WebGL/WebGL2 through the browser-selected adapter |
GPU support does not automatically move all Python or JavaScript onto a GPU.
The graph API executes its supported operations on the selected backend.
The playground's ui drawing API uses a 2D canvas; the Life computation runs
on the selected graph backend.
import torch
from torch import nn
model = nn.Sequential(nn.Linear(2, 4), nn.ReLU(), nn.Linear(4, 1))
x = torch.tensor([[1.0, 2.0], [3.0, 4.0]])
inference = torch.compile(model)(x) # no zipp_gpu import or backend name
inference.submit(lambda y: print(y.tolist()))Your model, tensor creation and forward pass use the supported Torch API.
Zipp records the inference graph and the host executes it on the selected
backend. submit(callback) is a Zipp extension: browser GPU completion is
asynchronous; this is not a drop-in implementation of PyTorch's torch.compile.
It supports float32 inference and an opt-in dense-model GPU training path; it does not run CUDA scripts.
The callback receives a regular CPU Torch-compatible tensor. Eager CPU
nn.Conv2d also supports forward/backward passes and optimizer updates;
GPU convolution remains future work.
Try Samples → Python: Torch ML inference (GPU) in the playground. The complete example draws its predictions and checks them against eager inference with the same weights. Explicit GPU selection fails visibly if unavailable; automatic selection reports the backend it used.
The ordinary training step stays familiar:
import torch.nn.functional as F
optimizer = torch.optim.SGD(model.parameters(), lr=0.01)
target = torch.tensor([[1.0], [2.0]])
def train_step(x, target):
optimizer.zero_grad()
loss = F.mse_loss(model(x), target)
loss.backward()
optimizer.step()
return loss
compiled_step = torch.compile(train_step, training=True)
compiled_step(x, target).submit(lambda loss: print(loss.item()))training=True and .submit(...) are experimental Zipp extensions, not
PyTorch's synchronous torch.compile API. Forward computation, first-order
gradients and SGD updates execute on the selected backend. The success callback
receives the loss after the CPU model's weights and gradients have been updated.
Wait for that callback before recording the next step.
Try Samples → Python: Torch ML training (GPU) to watch a small network learn
y = x², with predictions and a live loss curve. Its
model and training step also run in
native PyTorch; the playground driver
provides asynchronous scheduling and graphics.
What torch.compile captures today is float32 dense layers with relu, gelu,
sigmoid or tanh, softmax/log_softmax over the last dimension, sum/mean
over all elements or one dimension, MSE-style losses and F.cross_entropy
against integer class targets (the protocol's fused operation), trained with
SGD (momentum, dampening, Nesterov, weight decay), Adam or AdamW. Momentum
buffers and Adam moments are read from optimizer.state, updated in the graph
and written back with the weights, so the next step continues where this one
left off.
Each call captures a new graph, uploads inputs and weights, and reads back the
loss, gradients and updated weights. To keep them on the device, prepare the
step once from an example call and then feed only each step's batch:
prepared = compiled_step.prepare(x, target) # records once; weights and
prepared.step(lambda loss: print(loss.item()), x, target) # optimizer state stay resident
prepared.steps(lambda losses: print(len(losses)), [(x, target)] * 8) # 8 steps, one submission
prepared.sync(lambda p: print("model and optimizer.state are current"))
prepared.dispose()prepare, step, steps, sync and dispose are Zipp extensions too. While a
prepared session is live its parameters are refused to compiled calls and eager
optimizer.step(); sync() leaves the model and optimizer.state exactly
where the same number of eager steps would (checked against PyTorch 2.11 to
about 1e-7). On an RTX 5090 in Chrome, a 784-256-10 batch-64 Adam step driven
from Python costs (warm medians) 74 ms per compiled() call on WebGPU, 5.1 ms
per prepared.step, and 1.8 ms per step with prepared.steps at eight steps
per run (WebGL2: 89 / 3.3 / 2.2 ms; WebAssembly: 74 / 4.9 / 3.6 ms). Small
examples demonstrate correctness, not a GPU speedup over native PyTorch. There is
no graph cache or multi-GPU training. See the
Torch compatibility guide for supported operations,
limits and failure behavior. NCA experiments remain in their separate repository.
Build with python-js-interop for trusted mixed-language projects:
import javascript
print(javascript.eval("[1, 2, 3].map(x => x * 2)")) # [2, 4, 6]This uses Zipp's own JavaScript evaluator in the same VM instance, with copied
lists, dictionaries and scalar results. from js import eval is also available.
It is opt-in because JavaScript shares VM globals with the Python runtime.
There are no live cross-language object proxies or JavaScript .py imports yet.
See the build instructions and contract.
Python running on Zipp WASM, with Game of Life on the GPU. This is a recording
of the actual folder playground: the Python source is compiled and executed by
Zipp's WebAssembly VM. The displayed life.py uses ordinary Torch matrix
operations and ReLU, with no zipp_gpu imports. These same rules run in regular
PyTorch. A separate playground driver compiles each update for WebGL2 and
receives its result asynchronously for drawing.
Open the full project playground · Static screenshot · Portable Torch Life rules · Playground driver
The recording shows a glider gun, pulsar, growing patterns and interacting
debris in one toroidal Life grid. GIFs autoplay and loop; the landing page offers
a pause button and honors reduced-motion settings. Capture provenance
records the engine, adapter, source hashes and duration. Regenerate with
python crates/zipp-wasm/playground/capture-demos.py (Chrome, Python Playwright,
ffmpeg and the Python WASM build required).
| Fast to start | Modern JavaScript | Ready to embed |
|---|---|---|
| 7.4 ms median process launch in the canonical capture. No snapshot to load. | 100% of the documented corrected core Test262 suite: 95,680 / 95,680, zero failures or skips. Original and corrected results are reported separately. | A native CLI, a Rust embedding API, and a browser WebAssembly runtime. |
- Explore the whole engine. The lexer, parser, register VM, GC, inline caches and JITs live together in this repository.
- Choose how to run it. Use the native JIT for trusted programs, a browser Worker for WebAssembly, or the separately built hardened native runner.
- Use Python and JavaScript. Explore the experimental Python frontend or call the browser's GPU graph runtime directly from JavaScript.
- Watch code at work. Explore Game of Life, Langton's ant, bouncing balls, and GPU graph examples with visible source, canvas and console output.
- See the evidence. Benchmarks include exact-output checks, raw results, confidence intervals and the workloads that still need work.
The performance results and language coverage explain the measurements and their scope.
The landing page embeds the complete project playground and exposes it at
/playground. It supports folders, loose files,
Python/JavaScript samples, entry files, arguments, editing, console/canvas output
and GPU selection. Files stay in the browser's virtual filesystem.
For a local source checkout:
cd crates/zipp-wasm
./build-variants.sh all
node playground/serve.cjsOpen the printed playground URL, choose Samples → Python: Game of Life, select
WebGL2 or WebGPU, and press Run. The console identifies the actual backend.
auto visibly falls back to CPU/WASM when hardware is unavailable; explicit GPU
selection reports an error instead. Browser settings decide which adapter is used.
The native NCA research project now lives independently at neuralautomata.com. Zipp does not include its PyTorch runner, models, checkpoints or native endpoints.
The 0.0.21 release provides x86-64 CLI binaries for Windows and Linux
(complete: JavaScript, the experimental Python frontend as zipp py, torch and
the native GPU path) and four browser WebAssembly packages:
| Download | Use it for |
|---|---|
zipp-wasm-0.0.21-web.zip |
JavaScript applications and embedding (about 1.33 MB Brotli) |
zipp-wasm-0.0.21-web-python-base.zip |
JavaScript plus experimental Python, without torch (about 1.70 MB) |
zipp-wasm-0.0.21-web-python.zip |
JavaScript plus experimental Python projects with torch built in, and the browser GPU adapters (about 2.06 MB) |
zipp-wasm-0.0.21-web-torch.zip |
The torch package (zipp_torch.wasm and its zipp_torch.js loader, about 0.38 MB) that adds torch to web-python-base at run time through addPythonPackage |
See GitHub Releases for published
assets and 0.0.21 release notes for scope and limits.
0.0.21 is the latest published release; the download commands below use it.
Every WASM archive carries the exact source revision, language profile and
checksums. The sizes are each module's Brotli-11 transfer size, as measured in
crates/zipp-wasm/README.md.
Save this as app.js, then choose your platform below:
const greet = name => `Hello, ${name}!`;
console.log(greet("Zipp"));Building an application? Start with the Rust embedding guide, the browser example, or the execution profiles.
Download and run with PowerShell
Download, extract, and run the native Windows executable from PowerShell:
$version = '0.0.21'
$archive = "zipp-$version-x86_64-pc-windows-msvc.zip"
Invoke-WebRequest "https://github.com/f2i-com/zipp.org/releases/download/v$version/$archive" -OutFile $archive
Expand-Archive -LiteralPath $archive -DestinationPath .
& ".\zipp-$version-x86_64-pc-windows-msvc\zipp.exe" js .\app.jsUse mjs instead of js for an ES module entry, including top-level await.
Download and run from your shell
Download, extract, and run the native Linux binary:
version=0.0.21
archive="zipp-$version-x86_64-unknown-linux-gnu.tar.gz"
curl -fLO "https://github.com/f2i-com/zipp.org/releases/download/v$version/$archive"
tar -xzf "$archive"
"./zipp-$version-x86_64-unknown-linux-gnu/zipp" js ./app.jsThe archive preserves the executable bit. If another tool removes it, restore it
with chmod +x zipp-0.0.21-x86_64-unknown-linux-gnu/zipp.
Clone the repository and build with Cargo
Install stable Rust and its platform toolchain (MSVC Build Tools on Windows, or a C compiler and linker on Linux). On Windows, run this in PowerShell:
git clone https://github.com/f2i-com/zipp.org.git zipp
Set-Location zipp
cargo build --locked --release
.\target\release\zipp.exe js .\app.jsOn Linux:
git clone https://github.com/f2i-com/zipp.org.git zipp
cd zipp
cargo build --locked --release
./target/release/zipp js app.js
./target/release/zipp mjs app.mjs # ES module entry, including top-level awaitA release build uses fat LTO and one codegen unit, so the final link is deliberately slower than a development build. The resulting executable has no runtime data-file dependency.
The CLI also runs Python: Zipp's own Python 3 implementation, parsed by its
own front end (crates/zipp-pyparse) and lowered
straight to the engine's register bytecode (no transpilation to JavaScript and
no second interpreter), so a .py file runs on the same VM. On the command
line hot Python loops are JIT-compiled (x86-64), and over the 36 programs of
the Python benchmark suite Zipp runs at about 1.75x CPython 3.13's time, with
range_loop, int_arith, float_arith, tuple_swap and global_read faster
than CPython; WebAssembly and budgeted embedders run the same code
interpreted. Compiled code is cached per user, so a hello-world starts in
about 20 ms. Save this as fib.py:
def fib(n):
a = 0
b = 1
for i in range(n):
a, b = b, a + b
return a
print(fib(30))zipp py fib.py # 832040
zipp run fib.py # frontend chosen by extension, shebang or directive
zipp run tool # an extensionless `#!/usr/bin/env python3` script
zipp run --lang=python - # from standard input (no project folder)Classes (including metaclasses, descriptors and __slots__), exceptions
with full tracebacks, generators, closures, comprehensions, match
statements, f-strings, the builtin types and a set of standard-library
modules (math, json, re, collections, itertools, functools,
dataclasses, enum, contextlib, typing, struct, hashlib, ...)
all work; async does not yet. Semantics are checked
differentially against CPython: the 155 programs of tests/python_corpus/*.py
must print exactly what CPython 3.13 prints, in the default, no-fast-path and
no-JIT modes.
A folder runs as a project: zipp py examples/python/project runs its
main.py, and zipp py lab/train.py --steps 20 runs one script of a
folder with arguments. Every file of the folder (subfolders included, up to
8 MiB each and 64 MiB in total) is loaded into the program's virtual
filesystem, so open(), os, os.path, pathlib and json.load see the
project's data; .py files are modules and packages by folder
(legacy/fast_memory.py is legacy.fast_memory, with or without an
__init__.py), and a package's own modules import each other relatively
(from .graph import Graph) as CPython resolves them; sys.argv carries the
arguments; and files the program
writes are copied back under the folder when it finishes (never over a file
it could not see, such as one in dist/, a dot-folder or over the limits).
Any script name runs, extensionless shebang scripts included, and output
appears as the program prints it. A test_*.py
entry runs its tests through the bundled pytest subset. The bundled
library also includes a torch subset (tensors over typed arrays with
reverse-mode autograd, nn, nn.functional, optim and its schedulers,
torch.utils.data, linalg, fft, distributions, half-precision, complex,
sparse and quantized tensors, save/load in supported PyTorch checkpoint
layouts) that runs on the engine's native CPU kernels, so
supported ML code can train and evaluate inside Zipp. Eager execution is CPU;
torch.compile records supported inference and training steps for the GPU
(dense layers with relu, gelu, sigmoid or tanh, softmax, MSE or fused
cross-entropy, SGD with momentum, Adam or AdamW), and compiled.prepare()
keeps the weights and optimizer state on the device between steps (see
Train a small model on the browser GPU).
Natively, zipp py runs the same recorded graphs on a hardware GPU through wgpu when one is present, synchronously and with the same results, and on the engine's tensor kernels otherwise (Native GPU). The
subset covers enough of what real model code reaches for -- torch.autocast
among it, a no-op here since every tensor is float32 -- that Hugging Face's
modeling_qwen3.py, modeling_qwen2.py, modeling_llama.py and
modeling_mistral.py compile and run exactly as published, within 1.8e-7 of
transformers (interop). The scope
matrix, limits and the bytecode design are in
docs/PYTHON_FRONTEND_EXPERIMENT.md. The
feature is on by default in the CLI (--no-default-features builds the
JavaScript-only binary) and off by default in the zipp-vm library and the
WebAssembly package, which offers it as a
separate build variant,
with torch either built in or added at run time as a separate package.
Python programs can also compute on the GPU in the browser: the bundled
zipp_gpu library records a float32 graph (broadcasting arithmetic, @,
relu/gelu/sigmoid/tanh, softmax, reductions, fused cross-entropy, SGD and
Adam update steps, a Conway-life step) and submits it, and the host runs it
through WebGPU, WebGL2, compiled WebAssembly kernels or a JavaScript
reference, calling the program back with the outputs. Graph.prepare turns
a graph into a session: inputs are fed per step, carried tensors such as
weights and optimizer state stay on the device, and several steps go in one
submission (0.79 ms per step on an RTX 5090 at eight steps per run, against
3.6 ms for one step per submission). A run that fails after device work
began poisons its session - only dispose() remains - rather than letting a
retry run on state of no particular step. Natively the same code runs on a
GPU when one is present, and otherwise evaluates on the CPU, bit for bit the
same as the reference.
A weight input can also be a checkpoint's own Q4_K or Q6_K blocks, decoded
inside the matmul on every backend, and matmul_fixed computes a product in
integers, the same bits on every backend that runs it
(FIXED-POINT.md). The
WebAssembly kernels are no_std Rust, and the committed binary is rebuilt and
compared in CI.
See crates/zipp-wasm/README.md.
There is also a local playground
that runs a folder of Python or JavaScript files on the WebAssembly engine:
open or drop a whole project folder (subfolders, data files and binary
checkpoints included), browse it in a file tree, pick any script as the
entry, give it arguments, and run it in the browser; files the program
writes show up in the tree. It has an editor, a console and a canvas the
program draws on through a small ui API (draw/update/on_click/
on_key hooks for animation and input), and a GPU sample that steps life
on the compute backend every frame:
cd crates/zipp-wasm && ./build-variants.sh all && node playground/serve.cjsYes: JavaScript can use WebGL directly. WebGL is a browser JavaScript API, and this project's WebGL2 graph backend, WebGPU graph backend, and CPU reference backend are written in JavaScript. Python is one way to author work for that runtime.
Save the following as gpu-example.html in the repository root, start the local
server above, and open http://127.0.0.1:8765/gpu-example.html (using the printed
port). It runs directly in the browser and requires no Python or Zipp build.
<!doctype html>
<meta charset="utf-8">
<title>JavaScript GPU graph</title>
<pre id="output">Running a WebGL2 graph…</pre>
<script type="module">
import { createRuntime } from "/crates/zipp-wasm/gpu-lab/src/runtime.mjs";
const output = document.getElementById("output");
let runtime;
try {
runtime = await createRuntime({ backend: "webgl2" });
const result = await runtime.execute({
version: 1,
nodes: [
{ id: 0, op: "input", shape: [3], data: [-2, 3, 4] },
{ id: 1, op: "input", shape: [3], data: [10, 20, 30] },
{ id: 2, op: "mul", a: 0, b: 1 },
{ id: 3, op: "relu", a: 2 }
],
outputs: [{ name: "values", id: 3 }]
});
output.textContent = JSON.stringify({
backend: result.backend,
adapter: runtime.info().adapter,
values: result.outputs.values.data // [0, 60, 120]
}, null, 2);
} catch (error) {
output.textContent = error.message;
} finally {
runtime?.dispose();
}
</script>Use backend: "webgpu" for WGSL compute shaders, or "auto" to try WebGPU,
WebGL2, compiled WASM and JavaScript in that order. Explicit GPU selections
fail if unavailable; auto reports any initialization fallback through
runtime.info().fallbackAttempts. Await each execution before submitting
another to the same runtime. Only named outputs are read back to JavaScript.
For your own graphics, browser JavaScript can create a separate canvas and call
canvas.getContext("webgl2", { powerPreference: "high-performance" }) to work
with WebGL directly. The graph API supplies a bounded set of float32 operations;
it does not turn arbitrary JavaScript into shaders. The WebGL2 backend shows
how float tensors are uploaded and processed with GLSL.
Paste this into a Python entry file in the WASM playground, select WebGL2 or WebGPU in its toolbar, and press Run:
from zipp_gpu import Graph
def show(result):
print(result["backend"], result["outputs"]["values"]["data"])
g = Graph()
a = g.tensor([-2, 3, 4])
b = g.tensor([10, 20, 30])
g.submit(show, values=(a * b).relu()) # callback receives [0, 60, 120]Python records the graph and submits it through a host request. The Worker's
JavaScript runtime validates it, allocates GPU buffers/textures, runs the shaders,
reads the requested outputs, and delivers the callback between VM calls.
Supported operations include elementwise arithmetic, ReLU, matrix multiplication,
sum and a wrapped Conway-Life update. These are float32 operations, with the
shape/work limits in the GPU Lab documentation.
Running this Python code through the native Zipp CLI runs the same graphs on a
hardware GPU through wgpu (Vulkan, Direct3D 12 or Metal) when one is present,
synchronously, and otherwise on its local CPU reference evaluator (--no-gpu or
ZIPP_GPU=0 forces it). It does not start PyTorch or CUDA.
The HTML example runs in the browser's JavaScript engine. JavaScript files
loaded into the Zipp playground editor run inside the Zipp VM and do not
automatically receive browser document, canvas or WebGL objects.
The current playground connects zipp_gpu requests for Python guest programs.
Its JavaScript guest GPU bridge is not yet wired into that playground. For a
custom Zipp embedder, the repository supplies
createZippGPUAdapter and
gpuExecute / gpuExecuteAsync:
the host must install the guest shim, grant gpu.execute, drain and dispatch
the host-call queue, and deliver callbacks. The adapter tests cover that contract;
they are not a claim that the stock JavaScript playground exposes it already.
Run the browser build in a dedicated Worker, with a deadline controlled by your page. The complete example includes setup, cleanup and resource limits.
Complete browser setup and Worker example
Download the browser bundle, then serve its JavaScript and WebAssembly files from the same origin as your app:
version=0.0.21
archive="zipp-wasm-$version-web.zip"
curl -fLO "https://github.com/f2i-com/zipp.org/releases/download/v$version/$archive"
unzip "$archive"
mkdir -p public/zipp-wasm
cp "zipp-wasm-$version-web/zipp_wasm.js" \
"zipp-wasm-$version-web/zipp_wasm_bg.wasm" \
public/zipp-wasm/For arbitrary code, do not run the synchronous engine on the page's main
thread. Add this dedicated module Worker as public/zipp-wasm/worker.js:
import init, { Engine } from "./zipp_wasm.js";
await init({
module_or_path: new URL("./zipp_wasm_bg.wasm", import.meta.url),
});
self.onmessage = ({ data }) => {
let engine;
try {
engine = new Engine();
engine.initScript(data.source);
self.postMessage({ type: "result", output: engine.takeOutput() });
} catch (error) {
self.postMessage({ type: "error", error: String(error) });
} finally {
engine?.dispose();
}
};
self.postMessage({ type: "ready" });Start one fresh Worker per run from responsive page code. The page owns both the load deadline and the execution deadline, so it can forcibly terminate a Worker even while guest JavaScript is blocking it:
export function runZipp(source, timeoutMs = 2_500) {
return new Promise((resolve, reject) => {
const worker = new Worker("/zipp-wasm/worker.js", { type: "module" });
let settled = false;
let timer = setTimeout(
() => finish(reject, new Error("Zipp WebAssembly failed to load")),
15_000,
);
function finish(callback, value) {
if (settled) return;
settled = true;
clearTimeout(timer);
worker.terminate();
callback(value);
}
worker.onmessage = ({ data }) => {
if (data.type === "ready") {
clearTimeout(timer);
timer = setTimeout(
() => finish(reject, new Error("JavaScript execution timed out")),
timeoutMs,
);
worker.postMessage({ source });
} else if (data.type === "result") {
finish(resolve, data.output);
} else if (data.type === "error") {
finish(reject, new Error(data.error));
}
};
worker.onerror = (event) =>
finish(reject, event.error ?? new Error(event.message));
});
}
const lines = await runZipp('console.log("Hello from Zipp");');
document.querySelector("#output").textContent = lines.join("\n");Serve the app over HTTP(S), not file://, configure .wasm as
application/wasm, and adjust /zipp-wasm/worker.js if the app is hosted below
a URL prefix. A Content Security Policy must allow 'wasm-unsafe-eval' in
script-src and the Worker URL in worker-src. Engine also enforces
instruction, heap, output, source, and WebAssembly-memory ceilings; Worker
termination supplies the separate wall-clock boundary. The browser build is
interpreter-only and grants no host capabilities by default. See the
zipp-wasm guide before exposing bridges or
accepting multi-tenant input.
| Input | Use | Boundary |
|---|---|---|
| Trusted programs and benchmarks | zipp js / zipp mjs |
Maximum-throughput native CLI with JITs enabled. |
| Arbitrary browser-hosted code | zipp-wasm in a dedicated Worker |
Interpreter-only safe-sandbox build; terminate and replace the Worker at the wall deadline and between tenants. |
| Hardened native execution | zipp-sandbox |
Separately resolved, no-JIT, unsafe-forbidden engine with instruction, heap, output, import, and wall-time limits. |
Build the hardened native runner separately so Cargo cannot unify its safety features with the ordinary JIT workspace:
cargo build --locked --release --manifest-path crates/zipp-sandbox/Cargo.toml
./crates/zipp-sandbox/target/release/zipp-sandbox script.jsImports are denied unless the host supplies one canonical root. The native
runner is language/process/resource containment, not a kernel sandbox; use a
restricted account, container, or OS sandbox when the threat model requires
one. See SECURITY.md for the full deployment checklist.
Call evaluation order and diagnostic switches
Every shipping profile evaluates a method call's reference before its
arguments, as EvaluateCall requires: receiver.m(input.value) runs the
getter on m before the getter on value, an argument's coercion cannot
replace the method already fetched, and when both sides throw the method's
exception is the one seen. The fused CallMethod lowering — the op the method
inline caches, intrinsic arms and inlining key on — is used only for
arguments that provably cannot observe the order (literals, register-resident
locals, arithmetic over literals, array/object/closure literals of such
parts); every other argument shape takes the captured GetProp +
CallWithThis path. That path is not the slow one: the interpreter serves
a captured boot intrinsic (arr.push(i % 13), s.charCodeAt(a[i]),
m.get(k + 1)) through the same inline and name-dispatched builtin lanes as
the fused form once the captured value is proven identical to the live
prototype intrinsic, and answers the read itself from the same proof; a
captured value that differs — an own shadow, an override installed before or
by the arguments, a subclass method — is invoked exactly as captured. Two
environment variables exist for diagnostics and benchmarking, read once per
process:
| Variable | Effect |
|---|---|
ZIPP_RELAXED_CALL_ORDER=1 |
Re-admits the pre-audit "primitive-operand" class (global, cell and property reads and arithmetic over them) to the fused lowering. Faster on some rows; observably wrong for a getter or proxy trap on either side of the call. Never a shipping profile. |
ZIPP_STRICT_CALL_ORDER=1 |
Forces the default, and wins when both are set. |
crates/zipp-vm/tests/call_order_default.rs runs the audit's probes under the
default, the interpreter, forced JIT, GC stress and both switches in clean
child processes.
Start with the recorded native results. These are stamped benchmark captures,
with the source revisions and raw evidence below; they are not a new measurement
of every change on main. Ratios are Zipp / competitor, so lower is faster.
| Native workload group | Zipp / Node geomean | Coverage |
|---|---|---|
| All 30 workloads | 0.728× | 13 normal + 17 hostile rows, weighted equally. |
| Normal workloads | 0.614× | All 13 normal rows. |
| Hostile workloads | 0.829× | All 17 stress-oriented rows. |
Median process launch in the canonical four-engine capture is 7.4 ms for Zipp, 30.4 ms for Node, 43.3 ms for Bun and 82.6 ms for Deno.
The goal is ambitious and literal: become faster than Node, Bun, and Deno on every maintained benchmark while preserving exact output and tier parity. Zipp is not there yet. The tables below show both the wins and the remaining gaps.
Full Node, Bun and Deno results, confidence intervals and methodology
The canonical table is the clean PGO capture at engine commit 8229b3fc, the
last before B280 changed native call lowering (see the note after the table):
real13_8229b3fc_pgo_2026-09-02.json
and head_clean_8229b3fc_pgo_2026-09-02.json.
Both artifacts record publishable:true, ALL_CORRECT=1, 15 complete
counterbalanced repetitions, 10,000 bootstrap samples, exact output, and no
source, engine, input, environment, process-health, or harness drift.
Node v24.12.0 · Bun 1.3.14 · Deno 2.6.10 · Zipp 0.0.13 canonical PGO SHA-256
bf9fddab…dc9986.
Cold medians include process launch; bold marks the lowest displayed median.
| Retained benchmark | Node | Bun | Deno | Zipp | Zipp / Node |
|---|---|---|---|---|---|
| async-promise-chain | 334 ms | 369 ms | 359 ms | 372 ms | 1.12× |
| class-prototype-hot | 297 ms | 332 ms | 329 ms | 226 ms | 0.77× |
| json-large | 270 ms | 192 ms | 322 ms | 271 ms | 1.01× |
| map-set-heavy | 784 ms | 855 ms | 1,264 ms | 672 ms | 0.84× |
| markdown-render | 268 ms | 207 ms | 316 ms | 209 ms | 0.77× |
| parse-large-js | 273 ms | 230 ms | 296 ms | 233 ms | 0.86× |
| polymorphic-objects | 328 ms | 331 ms | 340 ms | 309 ms | 0.94× |
| regex-log-scan | 478 ms | 564 ms | 460 ms | 448 ms | 0.94× |
| sparse-array | 81 ms | 113 ms | 129 ms | 73 ms | 0.91× |
| typedarray-math | 200 ms | 914 ms | 170 ms | 144 ms | 0.72× |
| Zipp / engine paired geomean | 0.878× [0.875, 0.884] | 0.752× [0.747, 0.755] | 0.765× [0.761, 0.774] | — | — |
The three architecture diagnostics remain outside the retained-ten headline:
| Diagnostic | Node | Bun | Deno | Zipp | Zipp / Node |
|---|---|---|---|---|---|
| polymorphic-objects-v2 | 81 ms | 87 ms | 131 ms | 24 ms | 0.30× |
| property-ic-shapes | 265 ms | 158 ms | 319 ms | 10 ms | 0.04× |
| sparse-array-v2 | 171 ms | 366 ms | 184 ms | 99 ms | 0.59× |
| Zipp / engine paired geomean | 0.186× [0.183, 0.188] | 0.166× [0.164, 0.169] | 0.145× [0.143, 0.149] | — | — |
Across all 13 normal rows, Zipp measures 0.614× Node [0.611, 0.617], 0.531× Bun [0.528, 0.533], and 0.521× Deno [0.519, 0.526]. It wins 33 of 39 point comparisons and 31 of 39 Bonferroni exact-sign comparisons.
The separately measured 17-case hostile corpus covers closures, mixed locals, shape churn, GC survival, async lifetimes, modules, a React-shaped kernel, a warm router, a JavaScript bytecode VM, and vendored NanoID:
| Hostile metric | vs Node | vs Bun | vs Deno |
|---|---|---|---|
| ordinary equal-row geomean | 0.829× [0.820, 0.833] | 0.647× [0.643, 0.655] | 0.419× [0.415, 0.423] |
| category-balanced geomean | 0.862× [0.852, 0.865] | 0.661× [0.656, 0.673] | 0.432× [0.429, 0.436] |
For the requested project-wide view, the explicit equal-row aggregate across
all 30 normal and hostile rows is 0.728× Node [0.723, 0.730], 0.594× Bun
[0.591, 0.598], and 0.460× Deno [0.458, 0.464]. It is calculated as
exp((13 × ln(G13) + 17 × ln(G17)) / 30); its descriptive bootstrap resamples
the two separately captured suites as independent strata.
The aggregate is ahead, but the literal every-row target is not met. Zipp has 21 of 30 Node point wins. The current Node point gaps are async promises and JSON in the normal set, plus closure calls, both shape stressors, allocation survival, long-lived async, React reconcile, and the warm router in the hostile set (the JSON and long-lived async intervals cross parity; both are 1.005×). The hostile guide reports each ratio rather than hiding these behind the geomean.
FASTER_THAN_NODE_ON_EVERY_ROW=0
FASTER_THAN_EVERY_ENGINE_ON_EVERY_ROW=0
The 8229b3fc engine keeps the c28781cf levers (the inline dense-Array
store lane, the fused | 0 add, B263's register classes) and adds B269-B273:
RegExp exec under a heap ceiling no longer walks the heap per exec, the recycle
pool's fallback sort is run-adaptive, a function that reaches itself through a
captured cell gets the native cross lane, bodies with for...of or try
receive a frame-backed cross entry instead of the interpreter trampoline, and
small holders take the holder-grain write barrier so an overwritten young value
no longer floats into old space. Each landed with a one-binary latch A/B; the
capture-to-capture row moves sit inside the intervals. B274-B278 then changed
only the interpreter (the wasm rows above), but B280 (v0.0.15) made
specification-order method calls the default and changed native call lowering.
The newer publishable capture at 14770703
(real13_14770703_pgo_2026-09-11.json,
head_clean_14770703_pgo_2026-09-11.json)
is slower against Node: headline ten 1.0025x [0.995, 1.009], all 13 0.677x
[0.671, 0.682] and hostile 17 cold 0.875x [0.869, 0.883]. HANDOFF B314 places
the change between 8229b3fc and e6e0f65d, with B280's strict default the
unmeasured suspect. No headline is claimed for 14770703. See the
bench guide, hostile suite, and
PERF_ROADMAP.md for exact methodology and remaining work.
These sections preserve earlier measurements and module snapshots. They describe the named captures, including their limitations, rather than the current release artifact.
Historical interpreter and WebAssembly comparisons
The v0.0.5 release was also measured against pinned interpreter builds of QuickJS-NG v0.16.2 and Boa v0.22.0. These are clean release-default builds on the same Windows x86-64 host, with identical generated source, exact-output validation, six counterbalanced repetitions, and 10,000 paired-bootstrap samples. Ratios are Zipp / competitor, so lower is faster.
| Native diagnostic | Zipp interpreter / competitor | 95% CI | point wins |
|---|---|---|---|
| frozen real13 vs QuickJS-NG | 0.6413× | 0.6386–0.6452 | 12 / 13 |
| micro5 vs QuickJS-NG | 0.8556× | 0.8405–0.8761 | 5 / 5 |
| micro5 vs Boa | 0.2539× | 0.2501–0.2590 | 5 / 5 |
micro5 vs Boa --optimize |
0.2522× | 0.2493–0.2593 | 5 / 5 |
The native result is an aggregate win, not a universal claim: QuickJS-NG was 1.0099× faster at the point median on the retained sparse-array row, while Zipp led the other twelve. In the historical v0.0.5 browser-WASM release capture, Zipp measured 0.2274× Boa but 2.1074× QuickJS-NG on adjusted execution across the five diagnostic workloads. That release's stripped module is 5,595,833 bytes raw (1,254,075 Brotli-11), between QuickJS-NG's 1,528,293-byte reactor (417,087 Brotli-11) and Boa's 21,296,176-byte module (5,484,164 Brotli-11).
The clean default-feature v0.0.6 release binary at engine-source commit
e3acee352074 reran all 13 current real13 inputs against QuickJS-NG v0.16.2,
with the runner selecting Zipp's interpreter through ZIPP_NOJIT=1. All 39
canonicalized validation outputs matched after the documented QuickJS CRLF-to-LF
normalization; raw output bytes and hashes remain recorded. Across six
counterbalanced rounds, Zipp won all 13 point medians: Zipp / QuickJS-NG was
0.6089665× (descriptive 95% interval 0.6072021–0.6122180) for cold
fresh-process time and 0.6058409× (0.6041440–0.6090422) after paired
empty-launch subtraction. This confirms the native interpreter result on the
final engine code; it does not predict WASM performance.
The production module recorded in this historical comparison, built from v0.0.13, is 5,558,860 bytes raw, 1,812,458 at
gzip-9, and 1,248,649 at Brotli-11 (SHA-256
bd8614fe5f3a3b8ef67f4b917cdefebb3fe69afa39a9804a0d3f6b0b6b267126). The
official QuickJS-NG v0.16.2 reactor is 1,528,293 bytes raw and 417,087 at
Brotli-11, so Zipp is 3.586× as large raw and 2.958× as large on the wire.
Main has moved past that module. On 2026-09-05 an external audit of the
WASM build was implemented as B274-B278 (see
PERF_ROADMAP.md): interpreter-side changes that close
four cliffs the native PGO capture never sees. Every unit-addressed read on a
non-ASCII string decoded from byte zero, so scanning loops were quadratic;
the string-part allocation preflight walked the whole heap on a window blind
to the heap's size; eval / new Function code owned no inline caches; and
an array with a named property lost its dense read path. The figures below
are interleaved A/B medians of the wasm artifact built from 400bcfe3
against the same artifact with these changes, on a shared developer machine
with other work running, so they are diagnostic, not a canonical capture;
the control kernels (ASCII scans, plain and fused calls, main-code property
loops) moved within ±4%.
| Wasm kernel | 400bcfe3 |
with B274-B278 |
|---|---|---|
sequential charCodeAt over 64K non-ASCII units |
4,462 ms | 3.2 ms |
| word tokenizer over 64K mostly-ASCII units with a few accents | 9,828 ms | 8.2 ms |
one-unit slice loop, 16K non-ASCII units |
449 ms | 4.3 ms |
join of 4,000 parts × 200, 300K objects retained |
8,423 ms | 138 ms |
join of 4,000 parts × 200, small heap |
548 ms | 130 ms |
monomorphic property loop installed through new Function |
20.9 ms | 14.8 ms |
a[i] loop on an array carrying a named property |
13.3 ms | 9.6 ms |
That recorded module predates these changes. The harness that produced the rows is
crates/zipp-wasm/tests/node/bench.cjs-style (persistent Engine, warmed,
interleaved builds).
We also attempted a direct, unscaled WASM run over the same v0.0.6
normal 13 and hostile 17 sources used by the v0.0.6 Node/Bun/Deno reruns in
target/bench-results/real13-v006-6650647a718c-pgo-15.json and
target/bench-results/hostile17-v006-6650647a718c-pgo-15.json. Those sources
are newer than the retained canonical public capture below. The WASM capture
preserves their exact bytes and Node output oracle, but it is explicitly
publishable:false: the production Zipp WASM API cannot load the two module
rows, QuickJS-NG's official reactor cannot drain pending jobs for three async
rows, and Zipp validated only 7 of the 28 script rows. Seventeen Zipp rows hit
the fixed production instruction or heap ceilings and four ended in other
engine errors. There were consequently no comparable normal-suite rows and
only five comparable hostile rows.
On those five available rows, Zipp / QuickJS-NG was 0.9604× for persistent
time and 0.9567× after paired-control subtraction, with Zipp ahead only on
warm-router (1 / 5 point wins). Those are incomplete row-level diagnostics,
not full-suite geomeans: this run does not establish that Zipp WASM is faster
than QuickJS-NG WASM. Zipp's separately sampled compile median was slower
(5.080 ms versus 1.795 ms), while its instantiation/start median was faster
(0.397 ms versus 1.676 ms). Their sums are not a measured end-to-end median.
The separate five-workload speed-kernel experiment remains useful attribution
evidence: it measured 0.0954663913× QuickJS-NG on persistent time. It is highly
specialization-sensitive, however; disabling the exact workload lanes measured
1.815× QuickJS-NG but 0.199× Boa, with Zipp ahead of Boa on all five rows.
That control used a dirty-tree diagnostic candidate, not the release artifact.
No current same-source normal-13-plus-hostile-17 Boa WASM run exists, so neither
micro result is a general interpreter ranking or a substitute for the
incomplete exact-suite result above.
The commands, exact revisions, all validation failures, per-row numbers, module
hashes, host-interface differences, and limitations are in
bench/comparison/README.md. These ecosystem
comparisons are deliberately separate from the canonical Node/Bun/Deno series
below.
Hosted CI at 1539eb4b
confirms 100% of the documented corrected core Test262 suite passes:
95,680 / 95,680 executions, zero failures and zero skips. This profile applies
five documented test corrections covering
nine inconsistent upstream executions. The original DateTimeFormat ECMA-402
shard also passes 488 / 488, without corrections or skips.
For comparison, the unmodified pinned core result is
99.991% of test262: 95,671 / 95,680 executions, with the same nine documented
upstream inconsistencies and zero skips. Validation evidence
retains both results; the corrected profile is not an unmodified-upstream claim.
The corpus is pinned to 4249661388e5d3f92a85186213da140a6481490f, including
staging and excluding the separate ECMA-402 suite.
| Validation profile | Passed | Failed | Skipped |
|---|---|---|---|
| Core with five documented test corrections | 95,680 | 0 | 0 |
| Original pinned core Test262 | 95,671 | 9 | 0 |
| Original DateTimeFormat ECMA-402 tests | 488 | 0 | 0 |
Eight original failures expose contradictions in the pinned Error/TypedArray harnesses; one Annex B test carries a superseded ES2017 expectation. The corrections retain both original and corrected reports, with exact file hashes and no skipped executions. The corrected profile permits no expected failures. Its 100% result is explicitly separate from unmodified upstream conformance. The German formatting failures are fixed in the VM without test changes. Passing the DateTimeFormat shard does not establish complete ECMA-402 coverage or universal ICU parity.
The 12 September 2026 correctness audit
recorded an earlier 95,939 / 95,942 result against corpus defaaf1571.
That older corpus has a different execution count. The audit tightened negative-test scoring and fixed 42
executions previously counted as passes despite reporting the wrong error type.
The runner's --expected-failures option now rejects unexpected failures, stale
expectations and skips; --json records the engine and corpus identities.
The former errored-module-cycle and deferred top-level-await failures are fixed.
The deep follow-up audit then exercises re-entrant buffer operations, strict numeric coercion, Promise capabilities, temporary GC roots, live collection iteration, weak references/finalization, compiler metadata boundaries and hostile browser-host exceptions. Its focused regressions run in both the ordinary and safe-sandbox profiles, and its weak reference/finalization Test262 shards pass 603 of 603 executions.
The standing correctness strategy compares default JIT, interpreter-only, forced-JIT, and majors-only-GC modes. A tier-differential fuzzer also generates self-checking programs and compares Node, the interpreter, and tier-forcing switches; benchmarks alone are never treated as proof of language correctness.
ES2015–ES2025 is essentially complete, including:
- classes, private elements, static blocks, and all eight decorator kinds;
- generators, async generators, promises, iterator helpers, and explicit resource management;
- 12 TypedArray kinds including
Float16Array,DataView, shared memory, and atomics; BigInt,Proxy/Reflect, modern regular expressions, andTemporalwith fifteen calendars;- ES modules, dynamic/typed/deferred/source-phase imports, and top-level await;
eval,Function,ShadowRealm, and browser-oriented embedding APIs (the WASM host SDK structured-clones data across its Worker boundary; guest code has nostructuredCloneglobal).
DateTimeFormat ships CLDR 48 data for en/en-US, de/de-DE, ja/ja-JP,
zh/zh-CN and ar-EG. Other Intl services
currently ship English locale data only. The detailed support notes and durable
architecture reference live in DOC.md.
Follow a program from source text to running code. Zipp's lexer, parser, bytecode compiler, NaN-boxed register VM, garbage collector, inline caches and native JITs are implemented here, so you can explore each stage in one codebase.
flowchart LR
A[JavaScript source] --> B[Lexer and parser]
B --> C[Register bytecode]
C --> D[Interpreter]
D --> E[Hot-loop OSR]
D --> F[Whole-function JIT]
E --> G[x86-64 / ARM64 native code]
F --> G
D <--> H[GC, shapes, inline caches]
G <--> H
x86-64 has the mature function, OSR, helper, inline-cache, integer, double, and guarded reducer tiers. ARM64 has a smaller guarded whole-function integer baseline. wasm32 and unsupported native targets use the pure interpreter.
Workspace map:
| Path | Purpose |
|---|---|
crates/zipp-vm |
Parser, compiler, VM, runtime, GC, JITs, and the experimental Python frontend (src/frontend). |
crates/zipp-cli |
zipp js / zipp mjs / zipp py command line. |
crates/regress-fork |
ECMAScript regex engine fork and conformance fixes. |
crates/zipp-wasm |
Browser/Worker embedding. |
crates/zipp-pyparse |
Zipp's own Python 3.13 lexer, arena syntax tree and parser, read directly by the Python compiler. |
crates/zipp-gpu |
Native GPU for zipp py: gpu-lab's runtime over wgpu (Vulkan, Direct3D 12, Metal). |
crates/zipp-sandbox |
Separately resolved hardened native runner. |
Bug reports, documentation improvements, small fixes and careful measurements are welcome. Browse the open issues or the roadmap to find a place to start.
For bug reports, include a small JavaScript example, the expected output, your Zipp version and execution profile. A clear reproducer makes it much easier to help.
Run the release tests before changing the engine:
cargo test --workspace --release
cargo check -p zipp-vm --no-default-features
cargo check -p zipp-vm --no-default-features --features safe-sandboxBuild the measured Windows PGO binary from an x64 Visual Studio Developer PowerShell with native Git Bash:
& 'C:\Program Files\Git\bin\bash.exe' tools/pgo.shRoutine benchmark artifacts belong under ignored target/bench-results/.
Only deliberately reviewed canonical evidence is promoted into bench/.
Commands, suite ownership, publication rules, and A/B examples are in
bench/README.md.
The project keeps negative results because a measured refutation is cheaper than repeating the same attractive mistake. Start with:
| Document | Use it for |
|---|---|
DOC.md |
Durable architecture, language, embedding, and development reference. |
PERF_ROADMAP.md |
Current performance evidence, open targets, and next gates. |
HANDOFF.md |
Exact current continuation state and commands. |
SECURITY.md |
Threat model and deployment requirements. |
docs/archive |
Dated historical handoffs, designs, and the full experiment ledger. |
Small, independently measured changes are preferred. Keep correctness and benchmark output exact, include an off-switch for risky optimizations, report neutral or negative evidence, and do not update the public engine table without a clean canonical capture.
If you find Zipp useful, give the project a star or share what you're building with it.



