Skip to content

Instantly share code, notes, and snippets.

@J-GainSec
Created April 26, 2026 08:35
Show Gist options
  • Select an option

  • Save J-GainSec/abc563d2bc0063530711e4342edf7537 to your computer and use it in GitHub Desktop.

Select an option

Save J-GainSec/abc563d2bc0063530711e4342edf7537 to your computer and use it in GitHub Desktop.
Fix I used to get Qwen 3.6 working with tool calls on my DGX Spark, LM Studio and Openclaw.
# openclaw-lmstudio-qwen36-fix
Small compatibility shim that makes **OpenClaw** work with **LM Studio** serving **Qwen 3.6**.
## What this fixes
In this setup, OpenClaw was failing against LM Studio even though:
- LM Studio itself was healthy
- the model loaded correctly
- direct requests to the server worked
The breakage came from two integration mismatches:
1. OpenClaw sent requests to:
- `POST /chat/completions`
but LM Studio expects:
- `POST /v1/chat/completions`
2. OpenClaw did not send:
- `reasoning_effort: "none"`
and Qwen 3.6 would emit reasoning-heavy responses that OpenClaw often treated as empty or incomplete turns.
This shim fixes both problems without modifying OpenClaw or LM Studio.
## What the shim does
The proxy sits in front of LM Studio and:
- rewrites `POST /chat/completions` to `POST /v1/chat/completions`
- injects `reasoning_effort: "none"` if it is missing
- forwards everything else unchanged
## Working topology
```text
OpenClaw -> shim :4001 -> LM Studio :4000
```
## Files
- `lmstudio_reasoning_shim.js` - the Node shim
- `lmstudio-reasoning-shim.service` - optional `systemd --user` service
## Requirements
- Node.js 18+ recommended
- LM Studio already running and reachable locally
- OpenClaw already configured to use an OpenAI-compatible provider
## LM Studio settings
These are the settings used in the working setup:
- Model: `qwen/qwen3.6-35b-a3b`
- Context length: `65536`
- Eval batch size: `8192`
- Flash Attention: `On`
- Offload KV cache to GPU: `On`
- Serve on local network: `On` if OpenClaw is on another machine
- LM Studio port: `4000`
Important:
- Simply setting `reasoning: false` inside OpenClaw was **not** sufficient
- The effective fix was the request-level injection of `reasoning_effort: "none"`
## OpenClaw settings
Point OpenClaw at the shim instead of LM Studio directly:
```json
{
"models": {
"providers": {
"litellm": {
"baseUrl": "http://192.168.13.37:4001",
"api": "openai-completions",
"models": [
{
"id": "qwen/qwen3.6-35b-a3b",
"name": "Qwen 3.6 35B A3B (Spark)",
"reasoning": false,
"input": ["text"],
"contextWindow": 65536,
"maxTokens": 8192
}
]
}
}
}
}
```
The default model used in the working setup was:
- `litellm/qwen/qwen3.6-35b-a3b`
## Install
### 1. Put the shim somewhere stable
Example:
```bash
mkdir -p ~/bin
cp lmstudio_reasoning_shim.js ~/bin/
chmod +x ~/bin/lmstudio_reasoning_shim.js
```
### 2. Run it manually
```bash
LMSTUDIO_SHIM_HOST=0.0.0.0 \
LMSTUDIO_SHIM_PORT=4001 \
LMSTUDIO_UPSTREAM=http://127.0.0.1:4000 \
node ~/bin/lmstudio_reasoning_shim.js
```
### 3. Or install the included systemd user service
```bash
mkdir -p ~/.config/systemd/user
cp lmstudio-reasoning-shim.service ~/.config/systemd/user/
systemctl --user daemon-reload
systemctl --user enable --now lmstudio-reasoning-shim.service
```
## Verify the shim directly
```bash
curl http://127.0.0.1:4001/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"qwen/qwen3.6-35b-a3b","messages":[{"role":"user","content":"Reply with exactly: OPENCLAW_OK"}],"max_tokens":64}'
```
Expected behavior:
- assistant `content` is present
- `reasoning_content` is empty
- reasoning token count is `0`
## Verify OpenClaw end to end
Example:
```bash
openclaw infer model run \
--model litellm/qwen/qwen3.6-35b-a3b \
--prompt "Reply with exactly: OPENCLAW_OK" \
--gateway --json
```
Expected output:
```json
{
"outputs": [
{
"text": "OPENCLAW_OK"
}
]
}
```
## Why not just change OpenClaw config?
Because in this setup OpenClaw:
- did not send `reasoning_effort: "none"`
- and sent the wrong path for LM Studio compatibility
So the problem was not only model config. It was request shaping.
--------------------------
lmstudio-reasoning-shim.service
---------------------------
[Unit]
Description=LM Studio reasoning-effort shim
After=network.target
[Service]
ExecStart=/usr/bin/node /home/nigel/bin/lmstudio_reasoning_shim.js
Restart=always
RestartSec=2
Environment=LMSTUDIO_SHIM_HOST=0.0.0.0
Environment=LMSTUDIO_SHIM_PORT=4001
Environment=LMSTUDIO_UPSTREAM=http://127.0.0.1:4000
[Install]
WantedBy=default.target
------------------------------
#!/usr/bin/env node
"use strict";
/**
* LM Studio compatibility shim for OpenClaw + Qwen 3.6.
*
* Root cause:
* - OpenClaw sent POST /chat/completions
* - LM Studio expects POST /v1/chat/completions for its OpenAI-compatible API
* - OpenClaw also did not send reasoning_effort="none"
* - Qwen 3.6 then emitted reasoning-heavy responses that OpenClaw often
* interpreted as empty or incomplete turns
*
* This shim fixes both issues:
* 1. rewrites /chat/completions -> /v1/chat/completions
* 2. injects reasoning_effort="none" when it is missing
*
* Everything else is passed through unchanged.
*/
const http = require("node:http");
const { URL } = require("node:url");
const LISTEN_HOST = process.env.LMSTUDIO_SHIM_HOST || "0.0.0.0";
const LISTEN_PORT = Number(process.env.LMSTUDIO_SHIM_PORT || "4001");
const UPSTREAM = new URL(process.env.LMSTUDIO_UPSTREAM || "http://127.0.0.1:4000");
function upstreamPathFor(reqUrl) {
if (reqUrl === "/chat/completions") {
return "/v1/chat/completions";
}
return reqUrl;
}
function rewriteJsonBody(req, bodyBuffer) {
const contentType = String(req.headers["content-type"] || "");
if (!contentType.includes("application/json")) {
return bodyBuffer;
}
let parsed;
try {
parsed = JSON.parse(bodyBuffer.toString("utf8"));
} catch {
return bodyBuffer;
}
if (req.method !== "POST") {
return bodyBuffer;
}
if (req.url !== "/v1/chat/completions" && req.url !== "/chat/completions") {
return bodyBuffer;
}
// OpenClaw did not send this, but LM Studio + Qwen 3.6 behaved correctly
// once this was present on the OpenAI-compatible request.
if (parsed.reasoning_effort == null) {
parsed.reasoning_effort = "none";
return Buffer.from(JSON.stringify(parsed));
}
return bodyBuffer;
}
const server = http.createServer((req, res) => {
const chunks = [];
req.on("data", (chunk) => chunks.push(chunk));
req.on("end", () => {
const originalBody = Buffer.concat(chunks);
const rewrittenBody = rewriteJsonBody(req, originalBody);
const upstreamPath = upstreamPathFor(req.url);
const headers = { ...req.headers };
headers.host = UPSTREAM.host;
headers["content-length"] = String(rewrittenBody.length);
const upstreamReq = http.request(
{
protocol: UPSTREAM.protocol,
hostname: UPSTREAM.hostname,
port: UPSTREAM.port,
method: req.method,
path: upstreamPath,
headers,
},
(upstreamRes) => {
res.writeHead(upstreamRes.statusCode || 502, upstreamRes.headers);
upstreamRes.pipe(res);
},
);
upstreamReq.on("error", (error) => {
res.writeHead(502, { "content-type": "application/json" });
res.end(
JSON.stringify({
error: "upstream_request_failed",
message: error.message,
}),
);
});
upstreamReq.end(rewrittenBody);
});
});
server.listen(LISTEN_PORT, LISTEN_HOST, () => {
process.stdout.write(
`lmstudio reasoning shim listening on ${LISTEN_HOST}:${LISTEN_PORT}, upstream ${UPSTREAM.href}\n`,
);
});
----------------------------------
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment