AION
Research paperLarge Language Models · Efficiency & Inference1 source · Oct 6, 2026

Lost in the bf16 Cast: Exporting Ternary Language Models Can Revert Most Low-Learning-Rate Code Changes

Ternary language models such as BitNet b1.58, Falcon-E and BitCPM are fine-tuned with higher-precision latent weights and deployed as ternary codes produced by an export step that, in the labs' documented pipelines, first casts the latents to bf16.

Key points

  • We audit those pipelines across three labs.
  • At fine-tuned endpoints, with learning rates selected to match a nominal learning-rate-to-bf16-ULP ratio, the documented export lowers greedy GSM8K strict accuracy from 58.79% to 0.78% for Falcon-E-1B-Base and from 36.13% to 0.39% for BitCPM-CANN-0.5B, and a bf16 save and reload lowers BitNet 2B-4T's strict accuracy by 27.54 points while its last-number accuracy rises.
  • Two compatibility remedies, writing the training quantizer's codes directly or adjusting the bf16 inputs until the unchanged tools emit them, each met a 4-point strict-accuracy non-inferiority criterion against online evaluation in all three models.
  • In two model families, randomized interventions on the initial distance from the threshold support distance-dependent selection of the codes that fine-tuning changes.

Sources (1)

Extractive summary: sentences quoted from the sources.

Before this

  1. Sep 30, 2026Expanding AI Storage Access with NVIDIA cuObject and the NVIDIA SCADA Server SDK
  2. Sep 28, 2026openai/openai-python v3.20.0
  3. Sep 22, 2026vllm-project/vllm v0.30.0
  4. Aug 22, 2026sgl-project/sglang v0.5.18
  5. Aug 10, 2026vllm-project/vllm v0.27.0
  6. Jun 10, 2026DiffusionGemma: 4x faster text generation

Related