Skip to content

Latest commit

 

History

History
150 lines (113 loc) · 4.84 KB

File metadata and controls

150 lines (113 loc) · 4.84 KB

Qwen3.6-27B on llama.cpp

Run Unsloth's GGUF build of Qwen3.6-27B locally with llama-server, then point opencode at it with the Spinloop in this directory.

This is the dense 27B model: every parameter is active on each token. That makes it heavier to run than the 35B-A3B mixture-of-experts build despite the smaller total, but it fits in less memory and is a good fit for a single GPU.

Prerequisites

  • A recent build of llama.cpp that includes llama-server (e.g. brew install llama.cpp, or build from source).
  • A GPU is strongly recommended. The UD-Q4_K_XL quant is roughly 17 GB on disk; for comfortable headroom plan for ~20 GB of VRAM (less if you offload fewer layers to the GPU).

1. Pull the model

llama-server can fetch GGUFs straight from Hugging Face. The quant is selected with the :TAG suffix:

llama-server -hf unsloth/Qwen3.6-27B-GGUF:UD-Q4_K_XL

On first run this downloads the UD-Q4_K_XL weights into the llama.cpp cache (~/.cache/llama.cpp) and then starts serving. Subsequent runs reuse the cache.

Prefer to download ahead of time? Use the Hugging Face CLI:

hf download unsloth/Qwen3.6-27B-GGUF --include "*UD-Q4_K_XL*"

(huggingface-cli download ... works too on older installs.)

2. Start llama-server

llama-server \
  -hf unsloth/Qwen3.6-27B-GGUF:UD-Q4_K_XL \
  --jinja \
  -ngl 99 \
  --ctx-size 32768 \
  --host 127.0.0.1 --port 8080

What the flags do:

  • -hf …:UD-Q4_K_XL — model repository and quant tag.
  • --jinja — use the model's built-in chat template. Required for Qwen3 tool calling to work correctly.
  • -ngl 99 — offload all layers to the GPU. Lower it (or drop it) for CPU-only or limited VRAM.
  • --ctx-size 32768 — context window in tokens. Raise or lower to taste.
  • --host/--port — the OpenAI-compatible API is served at http://127.0.0.1:8080/v1.

Rather than remember those flags, this directory keeps them in a preset.ini and lets spinloop build and run the command:

spinloop serve              # from this directory; reads ./Spinloop and its PRESET
spinloop serve --dry-run    # print the llama-server command without running it

Optional: quantise the KV cache

For long contexts the K/V cache can dominate memory. Quantising it to q8_0 roughly halves that cost. KV-cache quantisation requires flash attention:

llama-server \
  -hf unsloth/Qwen3.6-27B-GGUF:UD-Q4_K_XL \
  --jinja -ngl 99 --ctx-size 32768 --host 127.0.0.1 --port 8080 \
  -fa on \
  --cache-type-k q8_0 \
  --cache-type-v q8_0
  • -fa on — enable flash attention (on older builds this is just -fa).
  • --cache-type-k q8_0 / --cache-type-v q8_0 — 8-bit K and V caches.

Check it's up

curl http://127.0.0.1:8080/v1/models

3. Point opencode at it

llama-server speaks the OpenAI-compatible API, which is exactly what the llamacpp provider targets (default base URL http://localhost:8080/v1). Apply the Spinloop in this directory:

spinloop harness apply examples/llamacpp/qwen3.6-27b/Spinloop
# or, from this directory:
spinloop harness apply

The Spinloop is:

PROVIDER llamacpp
ALIAS    qwen3.6-27b
CONTEXT  32768            # match the server's --ctx-size
PRESET   ./preset.ini

ALIAS is the name opencode shows for the model (and the section serve reads from the preset). For a single-model server it's just a label — llama-server serves whichever model it loaded regardless of what's requested — so call it whatever you find readable. CONTEXT matches opencode's context window to the context each request gets, so it doesn't overshoot what llama-server will accept.

Serving more than one request at a time? Add a PARALLEL line (the file ships one commented out):

PARALLEL 2

llama.cpp's --ctx-size is a total KV-cache budget it divides across --parallel slots, so spinloop scales it to compensate: with CONTEXT 32768 and PARALLEL 2, spinloop serve launches llama-server with --ctx-size 65536 --parallel 2 — each of the two slots still gets the 32768 tokens CONTEXT asked for, rather than half of it. CONTEXT itself, and what opencode sees as the context window, stay at 32768 either way. See Parallelism for how vLLM and oMLX differ.

Running on a non-default host or port? Add a BASEURL line to the Spinloop (the file ships one commented out):

BASEURL http://127.0.0.1:9090/v1

Now start opencode and select llamacpp/qwen3.6-27b.

See also