Skip to content

bitnet-b1.58-2B-4T: weight_scale precision (bf16 vs I2_S f32) and packed trits vs the bf16 master weights #632

Description

Summary

We read BitNet b1.58 2B4T from its three Hugging Face repositories at their current revisions with independent decoders. We compared the scales and I2_S trailers of all 210 ternary tensors, and the trits of two tensors (layer 0 q_proj and down_proj, 24.2M weights) across all three repositories. Three observations; data and a standalone reproduction below.

  1. The packed checkpoint stores the I2_S scale rounded to bf16. In all 210 ternary tensors, weight_scale in microsoft/bitnet-b1.58-2B-4T (bf16) equals the f32 scale stored after the matching I2_S tensor in microsoft/bitnet-b1.58-2B-4T-gguf, rounded to the nearest bf16. The two therefore differ by 0.14% at the median and 0.36% at most (blk.4.attn_output.weight: 2.140625 vs 2.1330159). In the two tensors we compared, the packed and I2_S trits are identical, so for the same int8 activations the integer products agree and only the per-tensor factor differs. We did not run inference.

    Tensor Packed weight_scale (bf16) I2_S scale (f32) mean|W| of the bf16 master weights
    layer 0 q_proj 1.21875 1.2188548 1.2188505
    layer 0 down_proj 2.15625 2.1631613 2.1631434
  2. The packed trits cannot be recomputed from the published bf16 weights. Applying WeightQuant from transformers (s = 1 / mean|w| in float32, round(w * s), clamp) to microsoft/bitnet-b1.58-2B-4T-bf16 gives trits that differ from the packed checkpoint for 79,719 of 6,553,600 weights (1.22%) of layer 0 q_proj and 101,673 of 17,694,720 (0.57%) of layer 0 down_proj. All of them have the single bf16 value nearest the rounding threshold, |w| = 0.5 × weight_scale (0.609375 and 1.078125, where |w·s| is 0.49996 and 0.49841). At that value the packed checkpoint is mixed: 79,719 of the 163,223 q_proj weights with it are ±1 and the rest are 0; in down_proj, 101,673 of 1,009,468. Weights with the same bf16 value get different trits, so no rule applied to the bf16 values reproduces the packed trits. That fits trits computed from higher-precision master weights, of which the bf16 repository is a rounded copy; we cannot check this. The model card describes the bf16 repository as master weights for training or fine-tuning; its config sets quantization_mode: online, so loading it as configured starts from ternary weights that differ from the packed checkpoint in these tensors.

  3. I2_S trailers carry leftover bytes from other tensors. Each I2_S tensor in ggml-model-i2_s.gguf ends with 32 bytes: the f32 scale and 28 more. In all 210 I2_S tensors those 28 bytes are nonzero and equal the bytes an earlier, at least as large tensor in the file holds at the same offset (token_embd or a tensor of the same layer), which is what a reused output buffer leaves behind: the C quantize_i2_s that llama-quantize used when this GGUF was uploaded (April 2025) writes only the packed bytes and the scale. quantize_to_i2_s in utils/convert-hf-to-gguf-bitnet.py writes zeros there, so the two conversion paths give different bytes. The dequantizer does not read them, so inference is unaffected; mentioning it in case zeroing the tail in the C path is preferred.

Questions

  • Is the f32 scale in the GGUF the reference scale for this model, with weight_scale its bf16 rounding?
  • Were the packed trits and the GGUF scales computed from higher-precision master weights than the bf16 repository holds? If so, is the bf16 repository expected to reproduce the packed trits in online mode?

Revisions

Repository Revision File sha256 published by the Hub
microsoft/bitnet-b1.58-2B-4T 04c3b9ad9361b824064a1f25ea60a8be9599b127 model.safetensors 8143ae115ed6babe5e5ada8fb8c5b769d8f417802b2db042ad98b4f7ed73975b
microsoft/bitnet-b1.58-2B-4T-bf16 276681394656abdadb8e80e5b2c3db5e5d7fcaff model.safetensors 529637ff6dab1f5890767356928693f69ffe61d3b6040a43de9306b37bfd5ae1
microsoft/bitnet-b1.58-2B-4T-gguf a1f2f1c765812aa8af3f6eda4a313707064bba15 ggml-model-i2_s.gguf 4221b252fdd5fd25e15847adfeb5ee88886506ba50b8a34548374492884c2162

We read file headers and the listed tensors with HTTP range requests; the sha256 of every range read by our t27 run is recorded in fixtures/manifest.lock.json, and the audit script below reads the remaining trailer ranges from the same revisions. Write-up: docs/ternary-check.md.

Reproduce

tools/bitnet_audit.py needs only the Python standard library. It reads about 63 MB of byte ranges from the three revisions above and prints all 210 scales and trailers (with the tensor each trailer was copied from) and the layer-0 counts as JSON:

curl -sO https://raw.githubusercontent.com/dmitrii-f-t27/trinity-memory/master/tools/bitnet_audit.py
python3 bitnet_audit.py > audit.json

Output of that run: reports/ternary-check/bitnet-audit-2026-09-23.json. Our t27 decoders (t27/formats.t27, python3 -m trinity_memory.ternary_check, which needs the pinned t27 compiler and a Rust toolchain) give the same layer-0 trits, scales and mismatch counts.

Related: #608

At the revisions above the tensors named in #608 have nonzero scales where we read them: packed weight_scale of layer 0 gate_proj 1.5546875, up_proj 1.828125, down_proj 2.15625; I2_S scales of blk.0.ffn_up 1.8309356 and blk.2.ffn_up 1.0256335. Layer 0 down_proj decodes to about 31% −1, 38% 0 and 31% +1 in the packed checkpoint, and its bf16 master weights contain no zero words. We read only these byte ranges and did not run inference.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions