Blog
We share what we learn.
Engineering notes on local inference, heterogeneous computing, GPU kernels, speculative decoding, and hands-on benchmarks.
3 articles
Latest
Running DeepSeek V4.1 Flash on the Lucebox memory hierarchy: up to 26 tok/s from GPU, APU and SSD
DeepSeek V4.1 Flash is a 383 GB file. One Lucebox runs it by splitting its experts across the AMD Radeon AI PRO R9700, the unified memory of the Strix Halo and the SSD, placed by how often each is used: up to 26 tok/s writing code and about 120 tok/s reading the prompt, with a 128K context.
Vision LLM inference: Lucebox has 3.2x the throughput of NVIDIA DGX Spark
At the highest load each machine served, one Lucebox answered 58 image questions a minute (Qwen3.8-27B on its AMD Radeon AI PRO R9700 and the 284B DeepSeek V4 Flash Vision on its Strix Halo, at the same time) and an NVIDIA DGX Spark running llama.cpp 18. With the same Qwen model file and eight questions at once, the R9700 alone finishes 2.4x sooner.
Inference-aware cooling: 29% less fan speed on the AMD R9700, same temperatures
AI inference and cooling designed together on the R9700: the model forecasts the work ahead, a learned thermal model picks the slowest safe fan speed. Chat silent, agent loops 29% and batch work 25% quieter, same peak temperatures and throughput.
20 articles
More articles
ROCm beats Vulkan on Strix Halo
DeepSeek V4 Flash on one Strix Halo, both engines on the same box on the same day: Lucebox ROCm is 17 to 47% faster on prefill and 28 to 47% faster on speculative decode than llama.cpp Vulkan at 8K, 32K and 123K, with plain decode measured on the same box too.
Continuous batching: Qwen3.8-27B at 301 tok/s across five clients, up to 1.97× llama.cpp
One scheduler contract serves Qwen3.8-27B with DFlash2 and DeepSeek V4 Flash together: 106.6 tok/s at one client grows to 300.9 at five on the R9700, up to 1.97× llama.cpp at equal concurrency, on stable slots and paged KV.
Ling 3.0 Flash: up to 141.9 tok/s with adaptive DSpark and FlashKDA
Our result on one DGX Spark: up to 36.4% faster prompt reading, near-tied ordinary generation, and 141.9 tok/s in the best matched DSpark case.
Qwen3.8-27B on the AMD R9700: up to 227 tok/s
One Radeon AI PRO R9700 serves Qwen3.8-27B with the z-lab DFlash2 drafter: 208 tok/s HumanEval average, 3.8x llama.cpp decode with the same drafter, on a stock quant that matches an 8-bit reference.
Lucebox × Geometric: DeepSeek V4 Flash 0731 reaches 32.7 tok/s
On AMD Strix Halo, the 98.29 GB model scores 82/92 and reaches 32.7 tok/s with DSpark. Across our fixed HumanEval, GSM8K, and MATH evaluation, it averages 27.9 tok/s.
Tool prefix caching: 48× faster warm prefill for agent loops
Agent turns resend thousands of tool-definition tokens. Lucebox restores that stable prefix and processes only the new conversation: 1.04 seconds warm after a 50.35 second cold turn.
Lucebox beats DGX Spark by 3.63× on DeepSeek V4 Flash decode
51.1 tok/s on Lucebox versus our 14.09 tok/s average on one NVIDIA DGX Spark for the full 284B model. The complete $5,999 Lucebox costs 36% less than two DGX Sparks.
DeepSeek V4 Flash: 284B model, up to 32 tok/s on AMD Strix Halo
The full target runs locally with 128 GB unified memory, reaching up to 32 tok/s decode and roughly 250 tok/s indexed sparse prefill.
Laguna XS 2.1 on a RTX 3090: 296 tok/s peak, 152 tok/s at 256K
poolside's coding MoE with its official DFlash drafter: 296 tok/s peak, a flat 152 tok/s at 256K tokens, and prefill at 3,500 tok/s on one 24 GB card.
Luce KVFlash: 256K context with 72 MiB of KV on the GPU
KVFlash pages cold 64-token chunks to host RAM bit-exact, holding Qwen3.6-27B decode at 38.6 tok/s from 64K to 256K with unchanged accuracy.
Lucebox in a container: one image for every supported GPU
A prebuilt image spans the RTX 2080 Ti through RTX 5090. The fat-binary compile happens once in CI, with two host dependencies, self-tuning, and build provenance included.
Luce Spark: fit Qwen3.6 35B and Laguna XS.2 on a 16 GB GPU
Spark keeps only the experts traffic uses resident and swaps the rest: Qwen3.6 35B-A3B in 13.3 GiB and Laguna XS.2 in 14.6 GiB, self-tuning with one flag.
Gemma 4 26B edges out DeepSeek V4 Flash at 5× the speed
A ds4-eval-92 head-to-head: Gemma 4 26B on a 24 GB RTX 5090 Laptop ties DeepSeek V4 Flash on a 192 GB Mac at 78.3%, and decodes about five times faster.
Launch and tune Lucebox with real agent harnesses
Real-client profiles, launch scripts, and TQ3/DDTree results for OpenCode, Hermes, OpenClaw, Open WebUI, Codex, Claude Code, and Pi.
DFlash + PFlash on AMD Strix Halo: 2.5× end-to-end versus llama.cpp
Qwen3.6-27B on the Ryzen AI MAX+ 395 iGPU: 26.85 tok/s DFlash decode and a 2.51× end-to-end gain at 16K plus 1K generation.
Laguna XS.2 on a 3090: 111 tok/s and 5.4× prefill
Poolside Laguna XS.2 ported into DFlash and PFlash as the first MoE target supported by PFlash: about 107 tok/s decode and 5.4× faster 128K prefill than llama.cpp.
PFlash: 10× prefill speedup over llama.cpp at 128K on a RTX 3090
PFlash compresses 128K to 2.6K tokens before DFlash sees the prompt: 24.8 seconds to first token versus about 257 seconds for llama.cpp, with measured retrieval preserved.
DFlash on ggml: up to 207 tok/s Qwen3.5-27B on a RTX 3090
A standalone C++ and ggml speculative decoder with a DFlash block-diffusion draft and DDtree verifier: 3.43× AR and 128K context on 24 GB.
The eGPU myth: why a $300 dock will not make an AI workstation
tinygrad wrote an NVIDIA driver from scratch. We tested real models on an RTX 3090 over USB4: brilliant engineering, but the performance numbers are not there yet.
Megakernel: matching Apple Silicon efficiency at 2× the throughput
The first megakernel for hybrid DeltaNet and attention LLMs fuses all 24 layers into one CUDA dispatch, reaching 1.87 tok/J on an RTX 3090.