Summary
We read BitNet b1.58 2B4T from its three Hugging Face repositories at their current revisions with independent decoders. We compared the scales and I2_S trailers of all 210 ternary tensors, and the trits of two tensors (layer 0 q_proj and down_proj, 24.2M weights) across all three repositories. Three observations; data and a standalone reproduction below.
-
The packed checkpoint stores the I2_S scale rounded to bf16. In all 210 ternary tensors, weight_scale in microsoft/bitnet-b1.58-2B-4T (bf16) equals the f32 scale stored after the matching I2_S tensor in microsoft/bitnet-b1.58-2B-4T-gguf, rounded to the nearest bf16. The two therefore differ by 0.14% at the median and 0.36% at most (blk.4.attn_output.weight: 2.140625 vs 2.1330159). In the two tensors we compared, the packed and I2_S trits are identical, so for the same int8 activations the integer products agree and only the per-tensor factor differs. We did not run inference.
| Tensor |
Packed weight_scale (bf16) |
I2_S scale (f32) |
mean|W| of the bf16 master weights |
layer 0 q_proj |
1.21875 |
1.2188548 |
1.2188505 |
layer 0 down_proj |
2.15625 |
2.1631613 |
2.1631434 |
-
The packed trits cannot be recomputed from the published bf16 weights. Applying WeightQuant from transformers (s = 1 / mean|w| in float32, round(w * s), clamp) to microsoft/bitnet-b1.58-2B-4T-bf16 gives trits that differ from the packed checkpoint for 79,719 of 6,553,600 weights (1.22%) of layer 0 q_proj and 101,673 of 17,694,720 (0.57%) of layer 0 down_proj. All of them have the single bf16 value nearest the rounding threshold, |w| = 0.5 × weight_scale (0.609375 and 1.078125, where |w·s| is 0.49996 and 0.49841). At that value the packed checkpoint is mixed: 79,719 of the 163,223 q_proj weights with it are ±1 and the rest are 0; in down_proj, 101,673 of 1,009,468. Weights with the same bf16 value get different trits, so no rule applied to the bf16 values reproduces the packed trits. That fits trits computed from higher-precision master weights, of which the bf16 repository is a rounded copy; we cannot check this. The model card describes the bf16 repository as master weights for training or fine-tuning; its config sets quantization_mode: online, so loading it as configured starts from ternary weights that differ from the packed checkpoint in these tensors.
-
I2_S trailers carry leftover bytes from other tensors. Each I2_S tensor in ggml-model-i2_s.gguf ends with 32 bytes: the f32 scale and 28 more. In all 210 I2_S tensors those 28 bytes are nonzero and equal the bytes an earlier, at least as large tensor in the file holds at the same offset (token_embd or a tensor of the same layer), which is what a reused output buffer leaves behind: the C quantize_i2_s that llama-quantize used when this GGUF was uploaded (April 2025) writes only the packed bytes and the scale. quantize_to_i2_s in utils/convert-hf-to-gguf-bitnet.py writes zeros there, so the two conversion paths give different bytes. The dequantizer does not read them, so inference is unaffected; mentioning it in case zeroing the tail in the C path is preferred.
Questions
- Is the f32 scale in the GGUF the reference scale for this model, with
weight_scale its bf16 rounding?
- Were the packed trits and the GGUF scales computed from higher-precision master weights than the bf16 repository holds? If so, is the bf16 repository expected to reproduce the packed trits in online mode?
Revisions
| Repository |
Revision |
File |
sha256 published by the Hub |
microsoft/bitnet-b1.58-2B-4T |
04c3b9ad9361b824064a1f25ea60a8be9599b127 |
model.safetensors |
8143ae115ed6babe5e5ada8fb8c5b769d8f417802b2db042ad98b4f7ed73975b |
microsoft/bitnet-b1.58-2B-4T-bf16 |
276681394656abdadb8e80e5b2c3db5e5d7fcaff |
model.safetensors |
529637ff6dab1f5890767356928693f69ffe61d3b6040a43de9306b37bfd5ae1 |
microsoft/bitnet-b1.58-2B-4T-gguf |
a1f2f1c765812aa8af3f6eda4a313707064bba15 |
ggml-model-i2_s.gguf |
4221b252fdd5fd25e15847adfeb5ee88886506ba50b8a34548374492884c2162 |
We read file headers and the listed tensors with HTTP range requests; the sha256 of every range read by our t27 run is recorded in fixtures/manifest.lock.json, and the audit script below reads the remaining trailer ranges from the same revisions. Write-up: docs/ternary-check.md.
Reproduce
tools/bitnet_audit.py needs only the Python standard library. It reads about 63 MB of byte ranges from the three revisions above and prints all 210 scales and trailers (with the tensor each trailer was copied from) and the layer-0 counts as JSON:
curl -sO https://raw.githubusercontent.com/dmitrii-f-t27/trinity-memory/master/tools/bitnet_audit.py
python3 bitnet_audit.py > audit.json
Output of that run: reports/ternary-check/bitnet-audit-2026-09-23.json. Our t27 decoders (t27/formats.t27, python3 -m trinity_memory.ternary_check, which needs the pinned t27 compiler and a Rust toolchain) give the same layer-0 trits, scales and mismatch counts.
Related: #608
At the revisions above the tensors named in #608 have nonzero scales where we read them: packed weight_scale of layer 0 gate_proj 1.5546875, up_proj 1.828125, down_proj 2.15625; I2_S scales of blk.0.ffn_up 1.8309356 and blk.2.ffn_up 1.0256335. Layer 0 down_proj decodes to about 31% −1, 38% 0 and 31% +1 in the packed checkpoint, and its bf16 master weights contain no zero words. We read only these byte ranges and did not run inference.
Summary
We read BitNet b1.58 2B4T from its three Hugging Face repositories at their current revisions with independent decoders. We compared the scales and I2_S trailers of all 210 ternary tensors, and the trits of two tensors (layer 0
q_projanddown_proj, 24.2M weights) across all three repositories. Three observations; data and a standalone reproduction below.The packed checkpoint stores the I2_S scale rounded to bf16. In all 210 ternary tensors,
weight_scaleinmicrosoft/bitnet-b1.58-2B-4T(bf16) equals the f32 scale stored after the matching I2_S tensor inmicrosoft/bitnet-b1.58-2B-4T-gguf, rounded to the nearest bf16. The two therefore differ by 0.14% at the median and 0.36% at most (blk.4.attn_output.weight: 2.140625 vs 2.1330159). In the two tensors we compared, the packed and I2_S trits are identical, so for the same int8 activations the integer products agree and only the per-tensor factor differs. We did not run inference.weight_scale(bf16)q_projdown_projThe packed trits cannot be recomputed from the published bf16 weights. Applying
WeightQuantfromtransformers(s = 1 / mean|w|in float32,round(w * s), clamp) tomicrosoft/bitnet-b1.58-2B-4T-bf16gives trits that differ from the packed checkpoint for 79,719 of 6,553,600 weights (1.22%) of layer 0q_projand 101,673 of 17,694,720 (0.57%) of layer 0down_proj. All of them have the single bf16 value nearest the rounding threshold, |w| = 0.5 ×weight_scale(0.609375 and 1.078125, where |w·s| is 0.49996 and 0.49841). At that value the packed checkpoint is mixed: 79,719 of the 163,223q_projweights with it are ±1 and the rest are 0; indown_proj, 101,673 of 1,009,468. Weights with the same bf16 value get different trits, so no rule applied to the bf16 values reproduces the packed trits. That fits trits computed from higher-precision master weights, of which the bf16 repository is a rounded copy; we cannot check this. The model card describes the bf16 repository as master weights for training or fine-tuning; its config setsquantization_mode: online, so loading it as configured starts from ternary weights that differ from the packed checkpoint in these tensors.I2_S trailers carry leftover bytes from other tensors. Each I2_S tensor in
ggml-model-i2_s.ggufends with 32 bytes: the f32 scale and 28 more. In all 210 I2_S tensors those 28 bytes are nonzero and equal the bytes an earlier, at least as large tensor in the file holds at the same offset (token_embdor a tensor of the same layer), which is what a reused output buffer leaves behind: the Cquantize_i2_sthatllama-quantizeused when this GGUF was uploaded (April 2025) writes only the packed bytes and the scale.quantize_to_i2_sinutils/convert-hf-to-gguf-bitnet.pywrites zeros there, so the two conversion paths give different bytes. The dequantizer does not read them, so inference is unaffected; mentioning it in case zeroing the tail in the C path is preferred.Questions
weight_scaleits bf16 rounding?Revisions
microsoft/bitnet-b1.58-2B-4T04c3b9ad9361b824064a1f25ea60a8be9599b127model.safetensors8143ae115ed6babe5e5ada8fb8c5b769d8f417802b2db042ad98b4f7ed73975bmicrosoft/bitnet-b1.58-2B-4T-bf16276681394656abdadb8e80e5b2c3db5e5d7fcaffmodel.safetensors529637ff6dab1f5890767356928693f69ffe61d3b6040a43de9306b37bfd5ae1microsoft/bitnet-b1.58-2B-4T-ggufa1f2f1c765812aa8af3f6eda4a313707064bba15ggml-model-i2_s.gguf4221b252fdd5fd25e15847adfeb5ee88886506ba50b8a34548374492884c2162We read file headers and the listed tensors with HTTP range requests; the sha256 of every range read by our t27 run is recorded in
fixtures/manifest.lock.json, and the audit script below reads the remaining trailer ranges from the same revisions. Write-up:docs/ternary-check.md.Reproduce
tools/bitnet_audit.pyneeds only the Python standard library. It reads about 63 MB of byte ranges from the three revisions above and prints all 210 scales and trailers (with the tensor each trailer was copied from) and the layer-0 counts as JSON:curl -sO https://raw.githubusercontent.com/dmitrii-f-t27/trinity-memory/master/tools/bitnet_audit.py python3 bitnet_audit.py > audit.jsonOutput of that run:
reports/ternary-check/bitnet-audit-2026-09-23.json. Our t27 decoders (t27/formats.t27,python3 -m trinity_memory.ternary_check, which needs the pinned t27 compiler and a Rust toolchain) give the same layer-0 trits, scales and mismatch counts.Related: #608
At the revisions above the tensors named in #608 have nonzero scales where we read them: packed
weight_scaleof layer 0gate_proj1.5546875,up_proj1.828125,down_proj2.15625; I2_S scales ofblk.0.ffn_up1.8309356 andblk.2.ffn_up1.0256335. Layer 0down_projdecodes to about 31% −1, 38% 0 and 31% +1 in the packed checkpoint, and its bf16 master weights contain no zero words. We read only these byte ranges and did not run inference.