Lost in the bf16 Cast: Exporting Ternary Language Models Can Revert Most Low-Learning-Rate Code Changes
Ternary language models such as BitNet b1.58, Falcon-E and BitCPM are fine-tuned with higher-precision latent weights and deployed as ternary codes produced by an export step that, in the labs' documented pipelines, first casts the latents to bf16.
Key points
- We audit those pipelines across three labs.
- At fine-tuned endpoints, with learning rates selected to match a nominal learning-rate-to-bf16-ULP ratio, the documented export lowers greedy GSM8K strict accuracy from 58.79% to 0.78% for Falcon-E-1B-Base and from 36.13% to 0.39% for BitCPM-CANN-0.5B, and a bf16 save and reload lowers BitNet 2B-4T's strict accuracy by 27.54 points while its last-number accuracy rises.
- Two compatibility remedies, writing the training quantizer's codes directly or adjusting the bf16 inputs until the unchanged tools emit them, each met a 4-point strict-accuracy non-inferiority criterion against online evaluation in all three models.
- In two model families, randomized interventions on the initial distance from the threshold support distance-dependent selection of the codes that fine-tuning changes.
Sources (1)
- [1]Lost in the bf16 Cast: Exporting Ternary Language Models Can Revert Most Low-Learning-Rate Code ChangesarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 06:58 AM
Ternary language models such as BitNet b1.58, Falcon-E and BitCPM are fine-tuned with higher-precision latent weights and deployed as ternary codes produced by an export step that, in the labs' documented pipelines, first casts the latents to bf16.
We audit those pipelines across three labs.
Extractive summary: sentences quoted from the sources.
Before this
- Sep 30, 2026Expanding AI Storage Access with NVIDIA cuObject and the NVIDIA SCADA Server SDK
- Sep 28, 2026openai/openai-python v3.20.0
- Sep 22, 2026vllm-project/vllm v0.30.0
- Aug 22, 2026sgl-project/sglang v0.5.18
- Aug 10, 2026vllm-project/vllm v0.27.0
- Jun 10, 2026DiffusionGemma: 4x faster text generation