Now you can grow Bonsai on your potato
Full 27B-class reasoning in ternary transformer weights, for llama.cpp (CUDA, Metal, CPU)

Proof1 community thread
Key points
- No more excuses for GPU-poor folks not to start LLMing!
- ~9.3x smaller than FP16 (ideal) | 98.2% of FP16 intelligence retained | ~47 tok/s on an Apple M5 Max laptop
- ~5.9 GB language model (down from ~54 GB FP16) — full 27B-class reasoning on a standard laptop or a single GPU
- Two GGUF packings with custom ternary hybrid-attention kernels for llama.cpp (CUDA, Metal) — PTQ10 packs trits densely (1.75 bits/weight, 5.95 GB), PQ20 stores each trit in a 2-bit slot (2.13 bits/weight, 7.21 GB); packed weights are consumed directly, never expanded back to FP16
Sources (1)
- [1]Now you can grow Bonsai on your potator/LocalLLaMA (top, daily) · Oct 11, 12:58 PM
Full 27B-class reasoning in ternary transformer weights, for llama.cpp (CUDA, Metal, CPU)
No more excuses for GPU-poor folks not to start LLMing!
Extractive summary: sentences quoted from the sources.
Before this
- Oct 10, 2026Open-source Mac app that runs EmbeddingGemma 2 locally to search your files by what’s in them
- Oct 10, 2026Nace AI Open-Sources Drex 1.5: A 9B Decision Model That Scores Options, Not Text
- Oct 9, 2026Qwen/Qwen-Image-2.1-Turbo
- Oct 9, 2026Introducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash
- Oct 8, 2026ConwayResearch/Underdog-Saluki-27B-1.0
- Oct 7, 2026[AINews] Claude Haiku 5.5 — better than GPT-6 Luna at the same pricing