Run Unsloth's GGUF build of Qwen3.6-27B locally with llama-server, then point
opencode at it with the Spinloop in this directory.
This is the dense 27B model: every parameter is active on each token. That makes it heavier to run than the 35B-A3B mixture-of-experts build despite the smaller total, but it fits in less memory and is a good fit for a single GPU.
- A recent build of llama.cpp that
includes
llama-server(e.g.brew install llama.cpp, or build from source). - A GPU is strongly recommended. The
UD-Q4_K_XLquant is roughly 17 GB on disk; for comfortable headroom plan for ~20 GB of VRAM (less if you offload fewer layers to the GPU).
llama-server can fetch GGUFs straight from Hugging Face. The quant is selected
with the :TAG suffix:
llama-server -hf unsloth/Qwen3.6-27B-GGUF:UD-Q4_K_XLOn first run this downloads the UD-Q4_K_XL weights into the llama.cpp cache
(~/.cache/llama.cpp) and then starts serving. Subsequent runs reuse the cache.
Prefer to download ahead of time? Use the Hugging Face CLI:
hf download unsloth/Qwen3.6-27B-GGUF --include "*UD-Q4_K_XL*"(huggingface-cli download ... works too on older installs.)
llama-server \
-hf unsloth/Qwen3.6-27B-GGUF:UD-Q4_K_XL \
--jinja \
-ngl 99 \
--ctx-size 32768 \
--host 127.0.0.1 --port 8080What the flags do:
-hf …:UD-Q4_K_XL— model repository and quant tag.--jinja— use the model's built-in chat template. Required for Qwen3 tool calling to work correctly.-ngl 99— offload all layers to the GPU. Lower it (or drop it) for CPU-only or limited VRAM.--ctx-size 32768— context window in tokens. Raise or lower to taste.--host/--port— the OpenAI-compatible API is served athttp://127.0.0.1:8080/v1.
Rather than remember those flags, this directory keeps them in a
preset.ini and lets spinloop build and run the command:
spinloop serve # from this directory; reads ./Spinloop and its PRESET
spinloop serve --dry-run # print the llama-server command without running itFor long contexts the K/V cache can dominate memory. Quantising it to q8_0
roughly halves that cost. KV-cache quantisation requires flash attention:
llama-server \
-hf unsloth/Qwen3.6-27B-GGUF:UD-Q4_K_XL \
--jinja -ngl 99 --ctx-size 32768 --host 127.0.0.1 --port 8080 \
-fa on \
--cache-type-k q8_0 \
--cache-type-v q8_0-fa on— enable flash attention (on older builds this is just-fa).--cache-type-k q8_0/--cache-type-v q8_0— 8-bit K and V caches.
curl http://127.0.0.1:8080/v1/modelsllama-server speaks the OpenAI-compatible API, which is exactly what the
llamacpp provider targets (default base URL http://localhost:8080/v1). Apply
the Spinloop in this directory:
spinloop harness apply examples/llamacpp/qwen3.6-27b/Spinloop
# or, from this directory:
spinloop harness applyThe Spinloop is:
PROVIDER llamacpp
ALIAS qwen3.6-27b
CONTEXT 32768 # match the server's --ctx-size
PRESET ./preset.iniALIAS is the name opencode shows for the model (and the section serve reads
from the preset). For a single-model server it's just a label — llama-server
serves whichever model it loaded regardless of what's requested — so call it
whatever you find readable. CONTEXT matches opencode's context window to the
context each request gets, so it doesn't overshoot what llama-server will
accept.
Serving more than one request at a time? Add a PARALLEL line (the file ships
one commented out):
PARALLEL 2llama.cpp's --ctx-size is a total KV-cache budget it divides across
--parallel slots, so spinloop scales it to compensate: with CONTEXT 32768
and PARALLEL 2, spinloop serve launches llama-server with --ctx-size 65536 --parallel 2 — each of the two slots still gets the 32768 tokens
CONTEXT asked for, rather than half of it. CONTEXT itself, and what
opencode sees as the context window, stay at 32768 either way. See
Parallelism for how vLLM and
oMLX differ.
Running on a non-default host or port? Add a BASEURL line to the Spinloop (the
file ships one commented out):
BASEURL http://127.0.0.1:9090/v1Now start opencode and select llamacpp/qwen3.6-27b.
- The bigger mixture-of-experts sibling:
examples/llamacpp/qwen3.6-35b-a3b - The next generation of this dense size, with
spinloop remotedeployment to AWS:examples/llamacpp/qwen3.8-27b