Skip to content

Ling-3.0-flash-fp4 produces recurring spurious tokens with W4A16 + Marlin on GB10 #26

Description

@mrsilva

I'd like to report some issues I've been experiencing while testing the model.

AI generated report:

[Bug] Ling-3.0-flash-fp4 produces recurring spurious tokens with W4A16 + Marlin on GB10

Summary

Serving inclusionAI/Ling-3.0-flash-fp4 (MXFP4, W4A16 + Marlin MoE backend) on SGLang produces generations that are almost always logically/semantically correct, but sporadically corrupted mid-generation by a small, recurring set of nonsense tokens appearing at newline/whitespace boundaries — breaking Python syntax in code-generation tasks and inserting stray words into prose. The corrupting tokens are not random: across a 164-problem HumanEval run plus spot checks on GSM8K, a small set of tokens recurs at high frequency in wildly varying contexts (Python code, English math word problems), which points at a systematic bug rather than generic sampling noise (temperature is 0 — greedy decoding, no sampling).

Root cause localization, via a controlled isolation matrix (see "Root cause localization" section below): isolated to the W4A16 + Marlin MoE execution path. The corruption is an inference-correctness defect specific to this execution path in the configurations tested. It reproduces on both the PR branch and current SGLang mainline, while the W4A8 + FlashInfer-native MXFP4 path produces correct logits on the same prompts. Tokenizer/detokenizer ID→text decoding, CUDA graphs, page size, radix cache, and overlap scheduling were all independently ruled out. This is not yet narrowed further than "the W4A16+Marlin path" — the A/B test changes both the kernel implementation (Marlin vs. FlashInfer) and the compute precision (W4A16 vs. W4A8) simultaneously, so it does not by itself distinguish a Marlin implementation bug from a W4A16 precision/dequant/layout issue from an expert-routing interaction specific to that path.

Environment

  • Hardware: NVIDIA GB10 (DGX Spark–class), 121.6GB unified memory
  • Driver / CUDA: NVIDIA driver 580.173.02, CUDA 13.0 (nvcc release 13.0, V13.0.88)
  • OS: Ubuntu 24.04.4 LTS, aarch64
  • SGLang: built from source, sgl-project/sglang PR #33561 branch ling3-flash-dspark, commit 814d2d1680e7c431277656f1a48d1d9680a17fd3 (2026-08-26), reported version string 0.5.19.dev553+g814d2d168
    • One local patch applied before this bug was ever reached: python/sglang/srt/models/bailing_moe_v3.py — get_parallel().enable_dp_lm_head → get_parallel().config.enable_dp_lm_head (unrelated version-skew fix; the PR branch's model file lagged mainline's parallel-namespace config-bag restructuring)
  • Model: inclusionAI/Ling-3.0-flash-fp4, HF revision 3bae1cf4011b48475b2cc038fff283af49053ebc, model_type=bailing_hybrid, architectures=["BailingMoeV3ForCausalLM"] — 124B total / 5.1B active MoE, 5:1 KDA linear-attention + Gated MLA hybrid architecture, official MXFP4 quantization (~66GB checkpoint)
  • Python env: python3.11 -m venv, pip install -e "./python[all]" with SGLANG_BUILD_RUST_EXTS=none (no Rust toolchain on this box; rust extensions are optional/perf-only, unrelated to this bug)

Exact launch command

SGLANG_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1 SGLANG_JIT_DEEPGEMM_PRECOMPILE=0 \
SGLANG_ENABLE_JIT_DEEPGEMM=0 SGLANG_DSV4_FP4_DEQUANT=0 SGLANG_FP8_IGNORED_LAYERS="" \
python3 -m sglang.launch_server \
  --model-path ~/models/Ling-3.0-flash-fp4 --served-model-name ling-v3-flash-fp4 \
  --trust-remote-code --dtype bfloat16 --tp-size 1 --ep-size 1 \
  --host 0.0.0.0 --port 30000 --api-key sk-ling-cookbook-test \
  --max-running-requests 1 --max-mamba-cache-size 64 --chunked-prefill-size 8192 \
  --page-size 64 --context-length 262144 \
  --cuda-graph-backend-decode full --cuda-graph-max-bs-decode 1 --cuda-graph-bs-decode 1 \
  --cuda-graph-backend-prefill disabled --random-seed 308534008 \
  --reasoning-parser ling3 --tool-call-parser ling3 --attention-backend flashinfer \
  --disable-flashinfer-autotune --mem-fraction-static 0.75 --fp8-gemm-backend cutlass \
  --moe-runner-backend marlin --disable-shared-experts-fusion --enable-fp32-lm-head \
  --json-model-override-args '{"max_position_embeddings":262144,"rope_scaling":{"rope_type":"yarn","factor":2.0,"rope_theta":6000000,"partial_rotary_factor":0.5,"original_max_position_embeddings":131072}}'

This is the recipe's own documented "W4A16 + Marlin MoE backend" configuration (step 4a in inclusionAI/ling-cookbook's Ling-3.0-flash notebook), unmodified except for the one version-skew patch noted above.

Boot-time signal (investigated and ruled out as the cause — see "Root cause" section below)

Tokenizer for /home/miguel/models/Ling-3.0-flash-fp4 is still TokenizersBackend after retries with --trust-remote-code. Model-specific tokenizer attributes may be missing.

Reproduction

Requests are plain /v1/completions (raw, no chat template), temperature=0, do_sample=false — i.e. greedy decoding (no sampling). Note GPU execution is not necessarily bitwise-deterministic run to run (see "Reproducibility" below).

HumanEval example (humaneval_fence task, doc_id 0, full raw model output):

from typing import List

def has_close_elements(numbers: List[float], threshold: float) -> bool:
""" Check if in given list of numbers, are any two numbers closer to each other than
given threshold.
>>> has_close_elements([1.0, 2.0, 3.0], 0.5)
False
>>> has_close_elements([1.0, 2.8, 3.0, 4.0, 5.0, 2.0], 0.3)
True
"""
sorted_numbers = sorted(numbers)
for i in range(len(sorted_numbers) - 1):
if sorted_numbers[i + 1] - sorted_numbers[i] < threshold: Cordova
return True
return False

The logic is entirely correct; the single stray token Cordova was injected exactly where a newline + indent belonged (right after the : that opens the if body), producing a SyntaxError and a failing test.

GSM8K example (doc_id 107, excerpt from a math word-problem answer, English prose):

On Friday, he watched two 1-hour episodes:
1 hour + 1 hour群 = 2 hours

Again, a stray token (群, a single Chinese character meaning "group/crowd") appears embedded mid-word with no whitespace, in an otherwise fluent English arithmetic explanation.

Corrupting-token census (164-problem HumanEval run, counting occurrences of a small set of previously-observed junk strings across all completions):

Token Occurrences
潜水 44
oly 16
群 10
相当 9
口 5
区 2
潜水艇 1
不 1

Notable aside from the mainline row: mainline's bailing_moe_v3.py still has the unpatched bare get_parallel().enable_dp_lm_head call (the version-skew fix from the top of this report was not needed here — mainline's runtime_context.py apparently still supports the bare form), and it boots and runs fine. This confirms the version-skew crash was specific to the PR branch's rebase state, unrelated to this corruption bug — and that this corruption bug itself is present on current mainline, not just the PR branch.

Together, this rules out CUDA graphs, page size, radix cache, overlap scheduling, and tokenizer/detokenizer behavior as the cause, and confirms the corruption is present on both the PR branch and current SGLang mainline. It localizes the defect to the W4A16 + Marlin execution path specifically, without yet distinguishing which part of that path (Marlin kernel implementation, W4A16 precision/dequant, or an interaction with expert routing) is responsible.

Run-to-run flakiness is real but modest (occasionally a repeat comes out clean, e.g. graphs-disabled rep 0 and mainline rep 2) — consistent with the near-tied logit cluster from section 1: sometimes ordinary floating-point non-associativity nudges the argmax back to the correct token, but the underlying distribution is still corrupted the majority of the time.

Conclusion

The corruption is an inference-correctness defect specific to the W4A16 + Marlin MoE execution path in the configurations tested. It reproduces on both the PR branch and current SGLang mainline, while the W4A8 + FlashInfer-native MXFP4 path produces correct logits on the same prompts. CUDA graphs, page size, radix cache, overlap scheduling, and tokenizer/detokenizer behavior were independently ruled out.

Practical mitigation for anyone hitting this: use --moe-runner-backend flashinfer_mxfp4 --flashinfer-mxfp4-moe-precision default (step 4b in the recipe notebook) instead of --moe-runner-backend marlin (step 4a) — note the first boot with this backend triggers a FlashInfer JIT CUTLASS kernel compile that can OOM-kill itself if run after model weights are already loaded (competing for host RAM); pre-compile standalone first per the notebook's own step 4b.2 (ninja -C ~/.cache/.../cached_ops/fused_moe_120/) before ever loading the model, and it succeeds cleanly.

Minimal reproducer

The exact prefix from the original discovery (the notebook's instructional wrapper, plus the model's own verified-correct completion of the function signature/docstring/loop up through threshold:), sent as a single /generate request against the Marlin config above. Corruption appears at generation position 0, so this triggers with a single forward/decode step rather than requiring a multi-hundred-token completion:

curl -s http://127.0.0.1:30000/generate \
  -H 'Content-Type: application/json' \
  -H 'Authorization: Bearer sk-ling-cookbook-test' \
  -d '{
    "text": "Complete the following Python function. Respond with ONLY the full function (including any needed imports) in a single ```python code block, and nothing else.\n\n```python\nfrom typing import List\n\n\ndef has_close_elements(numbers: List[float], threshold: float) -> bool:\n    \"\"\" Check if in given list of numbers, are any two numbers closer to each other than\n    given threshold.\n    >>> has_close_elements([1.0, 2.0, 3.0], 0.5)\n    False\n    >>> has_close_elements([1.0, 2.8, 3.0, 4.0, 5.0, 2.0], 0.3)\n    True\n    \"\"\"\n\n```\n\n```python\nfrom typing import List\n\n\ndef has_close_elements(numbers: List[float], threshold: float) -> bool:\n    \"\"\" Check if in given list of numbers, are any two numbers closer to each other than\n    given threshold.\n    >>> has_close_elements([1.0, 2.0, 3.0], 0.5)\n    False\n    >>> has_close_elements([1.0, 2.8, 3.0, 4.0, 5.0, 2.0], 0.3)\n    True\n    \"\"\"\n    sorted_numbers = sorted(numbers)\n    for i in range(len(sorted_numbers) - 1):\n        if sorted_numbers[i + 1] - sorted_numbers[i] < threshold:",
    "sampling_params": {"temperature": 0, "max_new_tokens": 10},
    "return_logprob": true,
    "top_logprobs_num": 10,
    "return_text_in_logprobs": true
  }'

Expected (correct) output on the FlashInfer/W4A8 config: "\n return True\n return False".

Observed (Marlin/W4A16) output, reproduced identically across 3 separate requests without restart: " Cordova\n return True\n return False", with output_token_logprobs[0] == [-6.706..., 40236, " Cord"] — matching the trace in section 1 above exactly (the tiny logprob difference, -6.706 vs -6.6209, is the same ordinary run-to-run floating-point variance noted throughout this report).

A shorter, hand-simplified version of this prompt (same logical position, but with the docstring/type-hints/imports stripped down) did not reproduce the corruption — the exact token sequence leading up to the decision point appears to matter, not just the surface-level "right after a colon" pattern. The prompt above is the exact prefix, not an artificially minimized one; if a shorter faithful reproducer is needed, it should be derived from the model's own real generated tokens the same way this one was, not hand-written.

What would help

  • Confirmation from someone familiar with the Marlin W4A16 MoE kernel implementation in SGLang (VLLM_USE_B12X_MOE/B12X patch path, or wherever --moe-runner-backend marlin routes for this BailingMoeV3ForCausalLM/hybrid-KDA architecture) of any known numerical-precision or expert-routing issues, particularly on GB10/sm_121a — the corrupted-vs-clean logprob comparison in section 2 above should make this straightforward to reproduce and instrument on the maintainers' side.
  • Whether this is specific to this checkpoint's particular MXFP4 quantization/calibration, or a broader issue with the Marlin backend for hybrid linear-attention (KDA) + MoE architectures on this SGLang branch/mainline more generally.
[Bug] Ling-3.0-flash-fp4 produces recurring spurious tokens with W4A16 + Marlin on GB10 Summary

Serving inclusionAI/Ling-3.0-flash-fp4 (MXFP4, W4A16 + Marlin MoE backend) on SGLang produces generations that are almost always logically/semantically correct, but sporadically corrupted mid-generation by a small, recurring set of nonsense tokens appearing at newline/whitespace boundaries — breaking Python syntax in code-generation tasks and inserting stray words into prose. The corrupting tokens are not random: across a 164-problem HumanEval run plus spot checks on GSM8K, a small set of tokens recurs at high frequency in wildly varying contexts (Python code, English math word problems), which points at a systematic bug rather than generic sampling noise (temperature is 0 — greedy decoding, no sampling).

Root cause localization, via a controlled isolation matrix (see "Root cause localization" section below): isolated to the W4A16 + Marlin MoE execution path. The corruption is an inference-correctness defect specific to this execution path in the configurations tested. It reproduces on both the PR branch and current SGLang mainline, while the W4A8 + FlashInfer-native MXFP4 path produces correct logits on the same prompts. Tokenizer/detokenizer ID→text decoding, CUDA graphs, page size, radix cache, and overlap scheduling were all independently ruled out. This is not yet narrowed further than "the W4A16+Marlin path" — the A/B test changes both the kernel implementation (Marlin vs. FlashInfer) and the compute precision (W4A16 vs. W4A8) simultaneously, so it does not by itself distinguish a Marlin implementation bug from a W4A16 precision/dequant/layout issue from an expert-routing interaction specific to that path.
Environment

Hardware: NVIDIA GB10 (DGX Spark–class), 121.6GB unified memory
Driver / CUDA: NVIDIA driver 580.173.02, CUDA 13.0 (nvcc release 13.0, V13.0.88)
OS: Ubuntu 24.04.4 LTS, aarch64
SGLang: built from source, sgl-project/sglang PR [#33561](https://github.com/sgl-project/sglang/pull/33561) branch ling3-flash-dspark, commit 814d2d1680e7c431277656f1a48d1d9680a17fd3 (2026-08-26), reported version string 0.5.19.dev553+g814d2d168
    One local patch applied before this bug was ever reached: python/sglang/srt/models/bailing_moe_v3.py — get_parallel().enable_dp_lm_head → get_parallel().config.enable_dp_lm_head (unrelated version-skew fix; the PR branch's model file lagged mainline's parallel-namespace config-bag restructuring)
Model: inclusionAI/Ling-3.0-flash-fp4, HF revision 3bae1cf4011b48475b2cc038fff283af49053ebc, model_type=bailing_hybrid, architectures=["BailingMoeV3ForCausalLM"] — 124B total / 5.1B active MoE, 5:1 KDA linear-attention + Gated MLA hybrid architecture, official MXFP4 quantization (~66GB checkpoint)
Python env: python3.11 -m venv, pip install -e "./python[all]" with SGLANG_BUILD_RUST_EXTS=none (no Rust toolchain on this box; rust extensions are optional/perf-only, unrelated to this bug)

Exact launch command

SGLANG_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1 SGLANG_JIT_DEEPGEMM_PRECOMPILE=0
SGLANG_ENABLE_JIT_DEEPGEMM=0 SGLANG_DSV4_FP4_DEQUANT=0 SGLANG_FP8_IGNORED_LAYERS=""
python3 -m sglang.launch_server
--model-path ~/models/Ling-3.0-flash-fp4 --served-model-name ling-v3-flash-fp4
--trust-remote-code --dtype bfloat16 --tp-size 1 --ep-size 1
--host 0.0.0.0 --port 30000 --api-key sk-ling-cookbook-test
--max-running-requests 1 --max-mamba-cache-size 64 --chunked-prefill-size 8192
--page-size 64 --context-length 262144
--cuda-graph-backend-decode full --cuda-graph-max-bs-decode 1 --cuda-graph-bs-decode 1
--cuda-graph-backend-prefill disabled --random-seed 308534008
--reasoning-parser ling3 --tool-call-parser ling3 --attention-backend flashinfer
--disable-flashinfer-autotune --mem-fraction-static 0.75 --fp8-gemm-backend cutlass
--moe-runner-backend marlin --disable-shared-experts-fusion --enable-fp32-lm-head
--json-model-override-args '{"max_position_embeddings":262144,"rope_scaling":{"rope_type":"yarn","factor":2.0,"rope_theta":6000000,"partial_rotary_factor":0.5,"original_max_position_embeddings":131072}}'

This is the recipe's own documented "W4A16 + Marlin MoE backend" configuration (step 4a in inclusionAI/ling-cookbook's Ling-3.0-flash notebook), unmodified except for the one version-skew patch noted above.
Boot-time signal (investigated and ruled out as the cause — see "Root cause" section below)

Tokenizer for /home/miguel/models/Ling-3.0-flash-fp4 is still TokenizersBackend after retries with --trust-remote-code. Model-specific tokenizer attributes may be missing.

Reproduction

Requests are plain /v1/completions (raw, no chat template), temperature=0, do_sample=false — i.e. greedy decoding (no sampling). Note GPU execution is not necessarily bitwise-deterministic run to run (see "Reproducibility" below).

HumanEval example (humaneval_fence task, doc_id 0, full raw model output):

from typing import List

def has_close_elements(numbers: List[float], threshold: float) -> bool:
""" Check if in given list of numbers, are any two numbers closer to each other than
given threshold.
>>> has_close_elements([1.0, 2.0, 3.0], 0.5)
False
>>> has_close_elements([1.0, 2.8, 3.0, 4.0, 5.0, 2.0], 0.3)
True
"""
sorted_numbers = sorted(numbers)
for i in range(len(sorted_numbers) - 1):
if sorted_numbers[i + 1] - sorted_numbers[i] < threshold: Cordova
return True
return False

The logic is entirely correct; the single stray token Cordova was injected exactly where a newline + indent belonged (right after the : that opens the if body), producing a SyntaxError and a failing test.

GSM8K example (doc_id 107, excerpt from a math word-problem answer, English prose):

On Friday, he watched two 1-hour episodes:
1 hour + 1 hour群 = 2 hours

Again, a stray token (群, a single Chinese character meaning "group/crowd") appears embedded mid-word with no whitespace, in an otherwise fluent English arithmetic explanation.

Corrupting-token census (164-problem HumanEval run, counting occurrences of a small set of previously-observed junk strings across all completions):
Token Occurrences
潜水 44
oly 16
群 10
相当 9
口 5
区 2
潜水艇 1
不 1

This is a lower bound: the census searches only for previously identified anomalous strings, so it cannot rule out a ninth or tenth anomalous token that hasn't been noticed yet. It nevertheless shows that the same small set of anomalous tokens recurs at high frequency across completely unrelated prompts (Python function bodies, arithmetic word problems), rather than varying randomly per-context.

Effect on scoring: HumanEval (humaneval_fence, 164 problems) scored 42.68% (70/164) pass@1 — most failures are exactly this pattern (a syntactically-fatal stray token in otherwise-correct code). GSM8K, MMLU, and HellaSwag are far less affected in their aggregate scores because their answer-extraction (last number / answer letter) is largely insensitive to a single stray word elsewhere in the response — but manual sampling confirms the underlying corruption is present in their raw completions too, just usually non-fatal to scoring.
Reproducibility

Reran HumanEval alone from a completely fresh server boot (same command, same checkpoint, same everything): 42.1% (69/164) vs. the original run's 42.68% (70/164) — a one-sample difference, consistent with ordinary greedy-decoding floating-point non-associativity on GPU (a well-known, unrelated source of tiny run-to-run variance), not evidence against the bug being real or systematic. The corruption pattern itself (same small token set, same newline-adjacent injection points) reproduced identically.
Root cause localization: isolated to the W4A16 + Marlin MoE execution path, confirmed via logprobs

Before writing this up further, ran a controlled isolation matrix on the same checkpoint, SGLang build, prompts, and seed, varying one setting at a time. All server variants below use /generate with return_logprob=true, top_logprobs_num=10, return_text_in_logprobs=true, temperature=0, 3 repeats per prompt without restart, against the 5 known-bad prompts above.

  1. Tokenizer/detokenizer ruled out — the model itself selects the wrong token

At the exact corruption point in the has_close_elements prompt (right after threshold:, where a newline is expected), captured the full top-10 logprob distribution from the baseline (Marlin) config:

position 171
selected id: 40236 -> " Cord" logprob=-6.6209
top-10:
40236 " Cord" -6.6209
42711 "ulp" -6.8868
58048 "潜水" -7.0902
8691 "相当" -7.2368
7136 "oly" -7.2978
85302 "eville" -7.3737
94978 "综述" -7.4915
141403 " atual" -7.4934
1288 "口" -7.5585
117547 "有新" -7.5642

position 172
selected id: 17510 -> "ova" logprob=-1.5808

Positions 171+172 decode to " Cord" + "ova" = " Cordova" — this is the exact same corruption visible in the HumanEval example above, so this logprob trace corresponds directly to the observed text, not a separate/different instance of the bug.

\n (id 198, the token that should be near-certain here — this is a completely unambiguous decision point for correct Python) does not appear anywhere in the top 10. All 10 candidates are within 0.94 logprob of the selected token, while the expected newline token is absent from the top 10 entirely — this is therefore not a simple near-tie between the intended newline and one anomalous token, but a genuinely flat/confused distribution. Four of the ten (潜水, 相当, oly, 口) are exactly the same tokens independently observed corrupting other completions in the full 164-problem HumanEval run.

Independently loaded the checkpoint's tokenizer (transformers.AutoTokenizer.from_pretrained(..., trust_remote_code=True), which itself resolves to the same TokenizersBackend class SGLang uses) and confirmed tokenizer.decode([40236]) == " Cord", tokenizer.decode([198]) == "\n" — SGLang's own reported id/text pairing is correct. This rules out tokenizer ID→text decoding as the source of this observed corruption: the selected ID genuinely corresponds to " Cord"; the model's own computed distribution, not a downstream rendering step, is where the defect originates.
2. Backend comparison — the decisive test

Ran the identical prompt through the recipe's alternate config (step 4b in the notebook: W4A8 + FlashInfer-native MXFP4, --moe-runner-backend flashinfer_mxfp4 --flashinfer-mxfp4-moe-precision default, dropping --fp8-gemm-backend cutlass/Marlin — same checkpoint, same SGLang build/commit, same prompt, same seed, same context length, page size, and CUDA graph settings). At the same logical position:

position 171 (FlashInfer/W4A8)
selected id: 198 -> "\n" logprob=-0.0000009
next-best: id 220 logprob=-15.60 (~5.9M times less likely)

The model is essentially certain of the correct newline. Ran all 5 known-bad prompts x 3 repeats (15 completions total) through this backend: zero corruption, all clean, byte-for-byte identical across repeats for 4 of 5 prompts (the 5th varies only in trailing prose length, no junk tokens). This is W4A16/Marlin corrupt + W4A8/FlashInfer clean on the same checkpoint, build, prompts, and seed.
3. Full isolation matrix

All variants below were tested against the same 5 known-bad prompts x 3 repeats each (15 completions per row), same checkpoint and prompts throughout:
Variant Result Samples
Baseline: Marlin, graphs on, page 64, radix+overlap on Corrupts (2/3 and 3/3 repeats across two separate boots) 5×3
Marlin, --cuda-graph-backend-decode disabled Corrupts (2/3 repeats; 1st request clean, 2nd/3rd corrupt) 5×3
Marlin, --disable-radix-cache --disable-overlap-schedule Corrupts (2/3 repeats, different boot ordering) 5×3
Marlin, --page-size 1 (vs. baseline 64) Corrupts (different junk token, "Cordially" instead of "Cordova", same injection point) 5×3
Marlin, current SGLang mainline (commit 44a92e54b9efdd56be68a915ed717f4964f2e6ab, 2026-09-01 — not the PR #33561 branch) Corrupts identically (2/3 repeats, "Cord"/"Cordova") 5×3
FlashInfer/W4A8 (--moe-runner-backend flashinfer_mxfp4, everything else = baseline) 0 corruptions 5×3

Notable aside from the mainline row: mainline's bailing_moe_v3.py still has the unpatched bare get_parallel().enable_dp_lm_head call (the version-skew fix from the top of this report was not needed here — mainline's runtime_context.py apparently still supports the bare form), and it boots and runs fine. This confirms the version-skew crash was specific to the PR branch's rebase state, unrelated to this corruption bug — and that this corruption bug itself is present on current mainline, not just the PR branch.

Together, this rules out CUDA graphs, page size, radix cache, overlap scheduling, and tokenizer/detokenizer behavior as the cause, and confirms the corruption is present on both the PR branch and current SGLang mainline. It localizes the defect to the W4A16 + Marlin execution path specifically, without yet distinguishing which part of that path (Marlin kernel implementation, W4A16 precision/dequant, or an interaction with expert routing) is responsible.

Run-to-run flakiness is real but modest (occasionally a repeat comes out clean, e.g. graphs-disabled rep 0 and mainline rep 2) — consistent with the near-tied logit cluster from section 1: sometimes ordinary floating-point non-associativity nudges the argmax back to the correct token, but the underlying distribution is still corrupted the majority of the time.
Conclusion

The corruption is an inference-correctness defect specific to the W4A16 + Marlin MoE execution path in the configurations tested. It reproduces on both the PR branch and current SGLang mainline, while the W4A8 + FlashInfer-native MXFP4 path produces correct logits on the same prompts. CUDA graphs, page size, radix cache, overlap scheduling, and tokenizer/detokenizer behavior were independently ruled out.

Practical mitigation for anyone hitting this: use --moe-runner-backend flashinfer_mxfp4 --flashinfer-mxfp4-moe-precision default (step 4b in the recipe notebook) instead of --moe-runner-backend marlin (step 4a) — note the first boot with this backend triggers a FlashInfer JIT CUTLASS kernel compile that can OOM-kill itself if run after model weights are already loaded (competing for host RAM); pre-compile standalone first per the notebook's own step 4b.2 (ninja -C ~/.cache/.../cached_ops/fused_moe_120/) before ever loading the model, and it succeeds cleanly.
Minimal reproducer

The exact prefix from the original discovery (the notebook's instructional wrapper, plus the model's own verified-correct completion of the function signature/docstring/loop up through threshold:), sent as a single /generate request against the Marlin config above. Corruption appears at generation position 0, so this triggers with a single forward/decode step rather than requiring a multi-hundred-token completion:

curl -s http://127.0.0.1:30000/generate
-H 'Content-Type: application/json'
-H 'Authorization: Bearer sk-ling-cookbook-test'
-d '{
"text": "Complete the following Python function. Respond with ONLY the full function (including any needed imports) in a single python code block, and nothing else.\n\npython\nfrom typing import List\n\n\ndef has_close_elements(numbers: List[float], threshold: float) -> bool:\n """ Check if in given list of numbers, are any two numbers closer to each other than\n given threshold.\n >>> has_close_elements([1.0, 2.0, 3.0], 0.5)\n False\n >>> has_close_elements([1.0, 2.8, 3.0, 4.0, 5.0, 2.0], 0.3)\n True\n """\n\n\n\npython\nfrom typing import List\n\n\ndef has_close_elements(numbers: List[float], threshold: float) -> bool:\n """ Check if in given list of numbers, are any two numbers closer to each other than\n given threshold.\n >>> has_close_elements([1.0, 2.0, 3.0], 0.5)\n False\n >>> has_close_elements([1.0, 2.8, 3.0, 4.0, 5.0, 2.0], 0.3)\n True\n """\n sorted_numbers = sorted(numbers)\n for i in range(len(sorted_numbers) - 1):\n if sorted_numbers[i + 1] - sorted_numbers[i] < threshold:",
"sampling_params": {"temperature": 0, "max_new_tokens": 10},
"return_logprob": true,
"top_logprobs_num": 10,
"return_text_in_logprobs": true
}'

Expected (correct) output on the FlashInfer/W4A8 config: "\n return True\n return False".

Observed (Marlin/W4A16) output, reproduced identically across 3 separate requests without restart: " Cordova\n return True\n return False", with output_token_logprobs[0] == [-6.706..., 40236, " Cord"] — matching the trace in section 1 above exactly (the tiny logprob difference, -6.706 vs -6.6209, is the same ordinary run-to-run floating-point variance noted throughout this report).

A shorter, hand-simplified version of this prompt (same logical position, but with the docstring/type-hints/imports stripped down) did not reproduce the corruption — the exact token sequence leading up to the decision point appears to matter, not just the surface-level "right after a colon" pattern. The prompt above is the exact prefix, not an artificially minimized one; if a shorter faithful reproducer is needed, it should be derived from the model's own real generated tokens the same way this one was, not hand-written.
What would help

Confirmation from someone familiar with the Marlin W4A16 MoE kernel implementation in SGLang (VLLM_USE_B12X_MOE/B12X patch path, or wherever --moe-runner-backend marlin routes for this BailingMoeV3ForCausalLM/hybrid-KDA architecture) of any known numerical-precision or expert-routing issues, particularly on GB10/sm_121a — the corrupted-vs-clean logprob comparison in section 2 above should make this straightforward to reproduce and instrument on the maintainers' side.
Whether this is specific to this checkpoint's particular MXFP4 quantization/calibration, or a broader issue with the Marlin backend for hybrid linear-attention (KDA) + MoE architectures on this SGLang branch/mainline more generally.

Activity

  1. changed the title [-]Bug report: fixed set of phantom tokens injected into generated text[/-] [+]Ling-3.0-flash-fp4 produces recurring spurious tokens with W4A16 + Marlin on GB10[/+] on Sep 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions