Blog

We share what we learn.

Engineering notes on local inference, heterogeneous computing, GPU kernels, speculative decoding, and hands-on benchmarks.

3 articles

Latest

The Lucebox tower on a wooden deck under a starry sky beside the DeepSeek whale, with the Lucebox logo above

Running DeepSeek V4.1 Flash on the Lucebox memory hierarchy: up to 26 tok/s from GPU, APU and SSD

DeepSeek V4.1 Flash is a 383 GB file. One Lucebox runs it by splitting its experts across the AMD Radeon AI PRO R9700, the unified memory of the Strix Halo and the SSD, placed by how often each is used: up to 26 tok/s writing code and about 120 tok/s reading the prompt, with a 128K context.

The Lucebox tower on a wooden deck under a starry sky, the Qwen bear holding a magnifying glass over a printed bar chart and the DeepSeek whale with a printed diagram

Vision LLM inference: Lucebox has 3.2x the throughput of NVIDIA DGX Spark

At the highest load each machine served, one Lucebox answered 58 image questions a minute (Qwen3.8-27B on its AMD Radeon AI PRO R9700 and the 284B DeepSeek V4 Flash Vision on its Strix Halo, at the same time) and an NVIDIA DGX Spark running llama.cpp 18. With the same Qwen model file and eight questions at once, the R9700 alone finishes 2.4x sooner.

An AMD Radeon AI PRO R9700 and a 120 mm case fan on a wooden deck under a starry sky

Inference-aware cooling: 29% less fan speed on the AMD R9700, same temperatures

AI inference and cooling designed together on the R9700: the model forecasts the work ahead, a learned thermal model picks the slowest safe fan speed. Chat silent, agent loops 29% and batch work 25% quieter, same peak temperatures and throughput.

20 articles

More articles

A Strix Halo mainboard on a wooden deck under a starry sky, with the ROCm and Vulkan logos above it

ROCm beats Vulkan on Strix Halo

DeepSeek V4 Flash on one Strix Halo, both engines on the same box on the same day: Lucebox ROCm is 17 to 47% faster on prefill and 28 to 47% faster on speculative decode than llama.cpp Vulkan at 8K, 32K and 123K, with plain decode measured on the same box too.

A Lucebox tower on a wooden deck under a starry sky, the Qwen bear holding its key on the left and the DeepSeek whale on the right

Continuous batching: Qwen3.8-27B at 301 tok/s across five clients, up to 1.97× llama.cpp

One scheduler contract serves Qwen3.8-27B with DFlash2 and DeepSeek V4 Flash together: 106.6 tok/s at one client grows to 300.9 at five on the R9700, up to 1.97× llama.cpp at equal concurrency, on stable slots and paged KV.

Ling 3.0 Flash running with Lucebox on an NVIDIA DGX Spark

Ling 3.0 Flash: up to 141.9 tok/s with adaptive DSpark and FlashKDA

Our result on one DGX Spark: up to 36.4% faster prompt reading, near-tied ordinary generation, and 141.9 tok/s in the best matched DSpark case.

Qwen3.8-27B running on an AMD Radeon AI PRO R9700 with the DFlash2 block-diffusion drafter

Qwen3.8-27B on the AMD R9700: up to 227 tok/s

One Radeon AI PRO R9700 serves Qwen3.8-27B with the z-lab DFlash2 drafter: 208 tok/s HumanEval average, 3.8x llama.cpp decode with the same drafter, on a stock quant that matches an 8-bit reference.

Lucebox and Geometric with AMD Strix Halo hardware and a blue DeepSeek whale under a starry sky

Lucebox × Geometric: DeepSeek V4 Flash 0731 reaches 32.7 tok/s

On AMD Strix Halo, the 98.29 GB model scores 82/92 and reaches 32.7 tok/s with DSpark. Across our fixed HumanEval, GSM8K, and MATH evaluation, it averages 27.9 tok/s.

An amber tool prefix feeding three progressively growing blue agent turns

Tool prefix caching: 48× faster warm prefill for agent loops

Agent turns resend thousands of tool-definition tokens. Lucebox restores that stable prefix and processes only the new conversation: 1.04 seconds warm after a 50.35 second cold turn.

An AMD-Powered Lucebox combining a Radeon AI PRO R9700 with Ryzen AI MAX+ 395

Lucebox beats DGX Spark by 3.63× on DeepSeek V4 Flash decode

51.1 tok/s on Lucebox versus our 14.09 tok/s average on one NVIDIA DGX Spark for the full 284B model. The complete $5,999 Lucebox costs 36% less than two DGX Sparks.

DeepSeek V4 Flash running locally on AMD Strix Halo unified-memory hardware

DeepSeek V4 Flash: 284B model, up to 32 tok/s on AMD Strix Halo

The full target runs locally with 128 GB unified memory, reaching up to 32 tok/s decode and roughly 250 tok/s indexed sparse prefill.

Laguna XS 2.1 on an RTX 3090: 296 tok/s peak, flat 152 tok/s at 256K context

Laguna XS 2.1 on a RTX 3090: 296 tok/s peak, 152 tok/s at 256K

poolside's coding MoE with its official DFlash drafter: 296 tok/s peak, a flat 152 tok/s at 256K tokens, and prefill at 3,500 tok/s on one 24 GB card.

Luce KVFlash keeping a small resident pool of KV on the GPU and paging the rest to host RAM

Luce KVFlash: 256K context with 72 MiB of KV on the GPU

KVFlash pages cold 64-token chunks to host RAM bit-exact, holding Qwen3.6-27B decode at 38.6 tok/s from 64K to 256K with unchanged accuracy.

Lucebox shipping as one Docker image that runs across the supported GPU range

Lucebox in a container: one image for every supported GPU

A prebuilt image spans the RTX 2080 Ti through RTX 5090. The fat-binary compile happens once in CI, with two host dependencies, self-tuning, and build provenance included.

Luce Spark serving a 33-35B MoE from a fraction of the experts on consumer memory

Luce Spark: fit Qwen3.6 35B and Laguna XS.2 on a 16 GB GPU

Spark keeps only the experts traffic uses resident and swaps the rest: Qwen3.6 35B-A3B in 13.3 GiB and Laguna XS.2 in 14.6 GiB, self-tuning with one flag.

Gemma 4 26B on an RTX 5090 Laptop next to DeepSeek V4 Flash on a MacBook

Gemma 4 26B edges out DeepSeek V4 Flash at 5× the speed

A ds4-eval-92 head-to-head: Gemma 4 26B on a 24 GB RTX 5090 Laptop ties DeepSeek V4 Flash on a 192 GB Mac at 78.3%, and decodes about five times faster.

Lucebox client harness experiments on an RTX 3090

Launch and tune Lucebox with real agent harnesses

Real-client profiles, launch scripts, and TQ3/DDTree results for OpenCode, Hermes, OpenClaw, Open WebUI, Codex, Claude Code, and Pi.

AMD Strix Halo running Qwen3.6-27B locally through Lucebox

DFlash + PFlash on AMD Strix Halo: 2.5× end-to-end versus llama.cpp

Qwen3.6-27B on the Ryzen AI MAX+ 395 iGPU: 26.85 tok/s DFlash decode and a 2.51× end-to-end gain at 16K plus 1K generation.

Laguna XS.2 running on a single RTX 3090 inside the DFlash daemon

Laguna XS.2 on a 3090: 111 tok/s and 5.4× prefill

Poolside Laguna XS.2 ported into DFlash and PFlash as the first MoE target supported by PFlash: about 107 tok/s decode and 5.4× faster 128K prefill than llama.cpp.

PFlash speculative prefill compression for DFlash

PFlash: 10× prefill speedup over llama.cpp at 128K on a RTX 3090

PFlash compresses 128K to 2.6K tokens before DFlash sees the prompt: 24.8 seconds to first token versus about 257 seconds for llama.cpp, with measured retrieval preserved.

Qwen3.5-27B DFlash on ggml

DFlash on ggml: up to 207 tok/s Qwen3.5-27B on a RTX 3090

A standalone C++ and ggml speculative decoder with a DFlash block-diffusion draft and DDtree verifier: 3.43× AR and 128K context on 24 GB.

RTX 3090, eGPU dock, and MacBook running NVIDIA on macOS over USB4

The eGPU myth: why a $300 dock will not make an AI workstation

tinygrad wrote an NVIDIA driver from scratch. We tested real models on an RTX 3090 over USB4: brilliant engineering, but the performance numbers are not there yet.

RTX 3090, the GPU behind the megakernel

Megakernel: matching Apple Silicon efficiency at 2× the throughput

The first megakernel for hybrid DeltaNet and attention LLMs fuses all 24 layers into one CUDA dispatch, reaching 1.87 tok/J on an RTX 3090.