Qwen3.8-2.4T (1-bit) on 4x NVIDIA DGX Spark: native MTP self-speculative decode via llama.cpp RPC. ~8 prose / 11 structured tok/s. The how, the traps, the config.
-
Updated
Aug 13, 2026 - Shell
Qwen3.8-2.4T (1-bit) on 4x NVIDIA DGX Spark: native MTP self-speculative decode via llama.cpp RPC. ~8 prose / 11 structured tok/s. The how, the traps, the config.
Run Qwen3.8-27B locally on Apple Silicon with speculative decoding, driving the Qwen Code CLI. 40.6 tok/s on an M4 Pro, ~18 GB, no cloud.
Event-driven benchmark of speculative decoding KV-cache overhead under memory pressure, comparing baseline, oracle, expected-reserve, and peak-reserve scheduler policies.
Accelerating LLM inference on Jetson Orin Nano 8GB with llama.cpp, CUDA, and MTP speculative decoding
长上下文推测解码的收益判据与增量回退调度(RAISE)参考实现 — Reference implementation of a break-even criterion and RAISE scheduling for long-context speculative decoding
a super fast llm response using small llm model to prefix large llm model
Suffix-tree speculative decoding for Hugging Face Transformers with cross-request caching and generation-pipeline integration.
vLLM serving stack for Gemma 4 31B on RTX PRO 6000 Blackwell, with FP8 KV cache, MTP speculative decoding, and an async FastAPI logging proxy in front.
LLM inference optimization experiments: speculative decoding, continuous batching, KV-cache, quantization
Simulates speculative decoding to find the optimal speculation length K across 576 configurations (3 draft models x 8 K values x 6 acceptance rates x 4 cost ratios). Key findings: 6.06x max speedup, breakeven at cost_ratio=0.25, optimal K grows from 1-3 at 50% acceptance to 7-15 at 95% acceptance.
Mini LLM inference server: paged KV cache, continuous batching, prefix caching, speculative decoding, OpenAI-compatible API
Real implementation of speculative decoding showing that speedup is strongly prompt- and target-dependent: GPT-2 → GPT-2-medium achieves mean best speedup 1.013×, while GPT-2 → GPT-2-large reaches 1.253× and up to 1.846× on high-agreement prompts.
Agent-Level Speculative Orchestration & Formal Dual-Engine Code Generation (Rust, FastMCP, DeepSeek)
Proof-carrying, high-performance LLM inference in Rust on fe2o3
Controlled EAGLE-3 draft adaptation for frozen Qwen3-8B: Rust latency, replication, scaling controls, and serving trade-offs.
Split-KV for MTP / Speculative Decoding verify (seqlen_q=2) in FlashAttention-2 + W8A8 kernel tuning on Ada (5.8x attention speedup at 105K context)
Inference Systems Lab: 1,536 controlled vLLM requests on one RTX 4090, cache/concurrency comparisons, raw traces and bounded streaming benchmark tools.
Benchmark open-weight LLMs on vLLM in Docker under one fixed protocol: decode at 0/4k/16k, cold TTFT, 4-way concurrency, power and VRAM, three runs with spread. Standard-library Python, charts and a results explorer.
To associate your repository with the speculative-decoding topic, visit your repo's landing page and select "manage topics."