I trained a 102M recursive BitNet-v2 model from scratch: 64K context, trained on less than 5B tokens
Hiya, I’m releasing Recursive BitNet N-Gram 102M, a small experiment combining ternary weights, shared transformer layers, and hashed n-gram embeddings, trained with a whooping budget of 100€
Proof1 community thread
Key points
- DISCLAIMER: the post and the model card was made with the assist of AI.
- All model weights started from random initialization.
- • 65,536-token training context during the continuation phase.
- ──────────────────────── N-gram continuation 1× NVIDIA B300 64K 1.216B
Sources (1)
- [1]I trained a 102M recursive BitNet-v2 model from scratch: 64K context, trained on less than 5B tokensr/LocalLLaMA (top, daily) · Oct 11, 10:19 AM
Hiya, I’m releasing Recursive BitNet N-Gram 102M, a small experiment combining ternary weights, shared transformer layers, and hashed n-gram embeddings, trained with a whooping budget of 100€
DISCLAIMER: the post and the model card was made with the assist of AI.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 11, 2026Converting dense models into Mixture-of-Experts
- Oct 9, 2026Qwen/Qwen-Image-2.1-Turbo
- Oct 9, 2026Introducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash
- Oct 8, 2026huggingface/trl v1.15.0
- Oct 8, 2026Microsoft event debuts new AI-friendly hardware and Windows changes
- Oct 7, 2026Multimodal open d1 decision models for the edge