Created
April 26, 2026 08:35
-
-
Save J-GainSec/abc563d2bc0063530711e4342edf7537 to your computer and use it in GitHub Desktop.
Fix I used to get Qwen 3.6 working with tool calls on my DGX Spark, LM Studio and Openclaw.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| # openclaw-lmstudio-qwen36-fix | |
| Small compatibility shim that makes **OpenClaw** work with **LM Studio** serving **Qwen 3.6**. | |
| ## What this fixes | |
| In this setup, OpenClaw was failing against LM Studio even though: | |
| - LM Studio itself was healthy | |
| - the model loaded correctly | |
| - direct requests to the server worked | |
| The breakage came from two integration mismatches: | |
| 1. OpenClaw sent requests to: | |
| - `POST /chat/completions` | |
| but LM Studio expects: | |
| - `POST /v1/chat/completions` | |
| 2. OpenClaw did not send: | |
| - `reasoning_effort: "none"` | |
| and Qwen 3.6 would emit reasoning-heavy responses that OpenClaw often treated as empty or incomplete turns. | |
| This shim fixes both problems without modifying OpenClaw or LM Studio. | |
| ## What the shim does | |
| The proxy sits in front of LM Studio and: | |
| - rewrites `POST /chat/completions` to `POST /v1/chat/completions` | |
| - injects `reasoning_effort: "none"` if it is missing | |
| - forwards everything else unchanged | |
| ## Working topology | |
| ```text | |
| OpenClaw -> shim :4001 -> LM Studio :4000 | |
| ``` | |
| ## Files | |
| - `lmstudio_reasoning_shim.js` - the Node shim | |
| - `lmstudio-reasoning-shim.service` - optional `systemd --user` service | |
| ## Requirements | |
| - Node.js 18+ recommended | |
| - LM Studio already running and reachable locally | |
| - OpenClaw already configured to use an OpenAI-compatible provider | |
| ## LM Studio settings | |
| These are the settings used in the working setup: | |
| - Model: `qwen/qwen3.6-35b-a3b` | |
| - Context length: `65536` | |
| - Eval batch size: `8192` | |
| - Flash Attention: `On` | |
| - Offload KV cache to GPU: `On` | |
| - Serve on local network: `On` if OpenClaw is on another machine | |
| - LM Studio port: `4000` | |
| Important: | |
| - Simply setting `reasoning: false` inside OpenClaw was **not** sufficient | |
| - The effective fix was the request-level injection of `reasoning_effort: "none"` | |
| ## OpenClaw settings | |
| Point OpenClaw at the shim instead of LM Studio directly: | |
| ```json | |
| { | |
| "models": { | |
| "providers": { | |
| "litellm": { | |
| "baseUrl": "http://192.168.13.37:4001", | |
| "api": "openai-completions", | |
| "models": [ | |
| { | |
| "id": "qwen/qwen3.6-35b-a3b", | |
| "name": "Qwen 3.6 35B A3B (Spark)", | |
| "reasoning": false, | |
| "input": ["text"], | |
| "contextWindow": 65536, | |
| "maxTokens": 8192 | |
| } | |
| ] | |
| } | |
| } | |
| } | |
| } | |
| ``` | |
| The default model used in the working setup was: | |
| - `litellm/qwen/qwen3.6-35b-a3b` | |
| ## Install | |
| ### 1. Put the shim somewhere stable | |
| Example: | |
| ```bash | |
| mkdir -p ~/bin | |
| cp lmstudio_reasoning_shim.js ~/bin/ | |
| chmod +x ~/bin/lmstudio_reasoning_shim.js | |
| ``` | |
| ### 2. Run it manually | |
| ```bash | |
| LMSTUDIO_SHIM_HOST=0.0.0.0 \ | |
| LMSTUDIO_SHIM_PORT=4001 \ | |
| LMSTUDIO_UPSTREAM=http://127.0.0.1:4000 \ | |
| node ~/bin/lmstudio_reasoning_shim.js | |
| ``` | |
| ### 3. Or install the included systemd user service | |
| ```bash | |
| mkdir -p ~/.config/systemd/user | |
| cp lmstudio-reasoning-shim.service ~/.config/systemd/user/ | |
| systemctl --user daemon-reload | |
| systemctl --user enable --now lmstudio-reasoning-shim.service | |
| ``` | |
| ## Verify the shim directly | |
| ```bash | |
| curl http://127.0.0.1:4001/v1/chat/completions \ | |
| -H "Content-Type: application/json" \ | |
| -d '{"model":"qwen/qwen3.6-35b-a3b","messages":[{"role":"user","content":"Reply with exactly: OPENCLAW_OK"}],"max_tokens":64}' | |
| ``` | |
| Expected behavior: | |
| - assistant `content` is present | |
| - `reasoning_content` is empty | |
| - reasoning token count is `0` | |
| ## Verify OpenClaw end to end | |
| Example: | |
| ```bash | |
| openclaw infer model run \ | |
| --model litellm/qwen/qwen3.6-35b-a3b \ | |
| --prompt "Reply with exactly: OPENCLAW_OK" \ | |
| --gateway --json | |
| ``` | |
| Expected output: | |
| ```json | |
| { | |
| "outputs": [ | |
| { | |
| "text": "OPENCLAW_OK" | |
| } | |
| ] | |
| } | |
| ``` | |
| ## Why not just change OpenClaw config? | |
| Because in this setup OpenClaw: | |
| - did not send `reasoning_effort: "none"` | |
| - and sent the wrong path for LM Studio compatibility | |
| So the problem was not only model config. It was request shaping. | |
| -------------------------- | |
| lmstudio-reasoning-shim.service | |
| --------------------------- | |
| [Unit] | |
| Description=LM Studio reasoning-effort shim | |
| After=network.target | |
| [Service] | |
| ExecStart=/usr/bin/node /home/nigel/bin/lmstudio_reasoning_shim.js | |
| Restart=always | |
| RestartSec=2 | |
| Environment=LMSTUDIO_SHIM_HOST=0.0.0.0 | |
| Environment=LMSTUDIO_SHIM_PORT=4001 | |
| Environment=LMSTUDIO_UPSTREAM=http://127.0.0.1:4000 | |
| [Install] | |
| WantedBy=default.target | |
| ------------------------------ | |
| #!/usr/bin/env node | |
| "use strict"; | |
| /** | |
| * LM Studio compatibility shim for OpenClaw + Qwen 3.6. | |
| * | |
| * Root cause: | |
| * - OpenClaw sent POST /chat/completions | |
| * - LM Studio expects POST /v1/chat/completions for its OpenAI-compatible API | |
| * - OpenClaw also did not send reasoning_effort="none" | |
| * - Qwen 3.6 then emitted reasoning-heavy responses that OpenClaw often | |
| * interpreted as empty or incomplete turns | |
| * | |
| * This shim fixes both issues: | |
| * 1. rewrites /chat/completions -> /v1/chat/completions | |
| * 2. injects reasoning_effort="none" when it is missing | |
| * | |
| * Everything else is passed through unchanged. | |
| */ | |
| const http = require("node:http"); | |
| const { URL } = require("node:url"); | |
| const LISTEN_HOST = process.env.LMSTUDIO_SHIM_HOST || "0.0.0.0"; | |
| const LISTEN_PORT = Number(process.env.LMSTUDIO_SHIM_PORT || "4001"); | |
| const UPSTREAM = new URL(process.env.LMSTUDIO_UPSTREAM || "http://127.0.0.1:4000"); | |
| function upstreamPathFor(reqUrl) { | |
| if (reqUrl === "/chat/completions") { | |
| return "/v1/chat/completions"; | |
| } | |
| return reqUrl; | |
| } | |
| function rewriteJsonBody(req, bodyBuffer) { | |
| const contentType = String(req.headers["content-type"] || ""); | |
| if (!contentType.includes("application/json")) { | |
| return bodyBuffer; | |
| } | |
| let parsed; | |
| try { | |
| parsed = JSON.parse(bodyBuffer.toString("utf8")); | |
| } catch { | |
| return bodyBuffer; | |
| } | |
| if (req.method !== "POST") { | |
| return bodyBuffer; | |
| } | |
| if (req.url !== "/v1/chat/completions" && req.url !== "/chat/completions") { | |
| return bodyBuffer; | |
| } | |
| // OpenClaw did not send this, but LM Studio + Qwen 3.6 behaved correctly | |
| // once this was present on the OpenAI-compatible request. | |
| if (parsed.reasoning_effort == null) { | |
| parsed.reasoning_effort = "none"; | |
| return Buffer.from(JSON.stringify(parsed)); | |
| } | |
| return bodyBuffer; | |
| } | |
| const server = http.createServer((req, res) => { | |
| const chunks = []; | |
| req.on("data", (chunk) => chunks.push(chunk)); | |
| req.on("end", () => { | |
| const originalBody = Buffer.concat(chunks); | |
| const rewrittenBody = rewriteJsonBody(req, originalBody); | |
| const upstreamPath = upstreamPathFor(req.url); | |
| const headers = { ...req.headers }; | |
| headers.host = UPSTREAM.host; | |
| headers["content-length"] = String(rewrittenBody.length); | |
| const upstreamReq = http.request( | |
| { | |
| protocol: UPSTREAM.protocol, | |
| hostname: UPSTREAM.hostname, | |
| port: UPSTREAM.port, | |
| method: req.method, | |
| path: upstreamPath, | |
| headers, | |
| }, | |
| (upstreamRes) => { | |
| res.writeHead(upstreamRes.statusCode || 502, upstreamRes.headers); | |
| upstreamRes.pipe(res); | |
| }, | |
| ); | |
| upstreamReq.on("error", (error) => { | |
| res.writeHead(502, { "content-type": "application/json" }); | |
| res.end( | |
| JSON.stringify({ | |
| error: "upstream_request_failed", | |
| message: error.message, | |
| }), | |
| ); | |
| }); | |
| upstreamReq.end(rewrittenBody); | |
| }); | |
| }); | |
| server.listen(LISTEN_PORT, LISTEN_HOST, () => { | |
| process.stdout.write( | |
| `lmstudio reasoning shim listening on ${LISTEN_HOST}:${LISTEN_PORT}, upstream ${UPSTREAM.href}\n`, | |
| ); | |
| }); | |
| ---------------------------------- |
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment