Peer-to-peer LLM inference in the browser: pool your devices to run big open models, for chat and coding agents.
-
Updated
Oct 3, 2026 - JavaScript
Peer-to-peer LLM inference in the browser: pool your devices to run big open models, for chat and coding agents.
🐉 Hail Hydra — Multi-headed speculative execution framework for Claude Code. 10 AI agents, 3x faster, ~70% cheaper. Inspired by speculative decoding.
Minimal server-rendered dashboard framework with zero dependencies - layouts, widgets, and escape-safe rendering in plain Node.js
The Universal 245K Agent Microkernel for Local LLMs.
Playtest-graded benchmark for local LLM inference stacks — a coding agent builds a 10-file game, static + runtime gates + human playtest grade it. Validated on llama.cpp Vulkan / AMD Strix Halo.
Reproducible 16GB-VRAM Windows local-agent stack: DSH Desktop + KVMem (llama.cpp fork, KV cache in system RAM) + parameter panel. Flagship: Bonsai 2 CRACK 27B — 100+ tok/s, 262K ctx on RTX 4080. GSQ for intelligence at ~50 tok/s; QQZ default.
Proof-carrying, high-performance LLM inference in Rust on fe2o3
To associate your repository with the speculative-decoding topic, visit your repo's landing page and select "manage topics."