AnalysisOpinion / analysisLarge Language Models · Efficiency & Inference1 source · Oct 11, 2026

I trained a 102M recursive BitNet-v2 model from scratch: 64K context, trained on less than 5B tokens

Hiya, I’m releasing Recursive BitNet N-Gram 102M, a small experiment combining ternary weights, shared transformer layers, and hashed n-gram embeddings, trained with a whooping budget of 100€

Proof1 community thread

Key points

  • DISCLAIMER: the post and the model card was made with the assist of AI.
  • All model weights started from random initialization.
  • • 65,536-token training context during the continuation phase.
  • ──────────────────────── N-gram continuation 1× NVIDIA B300 64K 1.216B

Sources (1)

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 11, 2026Converting dense models into Mixture-of-Experts
  2. Oct 9, 2026Qwen/Qwen-Image-2.1-Turbo
  3. Oct 9, 2026Introducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash
  4. Oct 8, 2026huggingface/trl v1.15.0
  5. Oct 8, 2026Microsoft event debuts new AI-friendly hardware and Windows changes
  6. Oct 7, 2026Multimodal open d1 decision models for the edge

Related