AnalysisOpinion / analysisEfficiency & Inference1 source · Oct 11, 2026

Now you can grow Bonsai on your potato

Full 27B-class reasoning in ternary transformer weights, for llama.cpp (CUDA, Metal, CPU)

Proof1 community thread

Key points

  • No more excuses for GPU-poor folks not to start LLMing!
  • ~9.3x smaller than FP16 (ideal) | 98.2% of FP16 intelligence retained | ~47 tok/s on an Apple M5 Max laptop
  • ~5.9 GB language model (down from ~54 GB FP16) — full 27B-class reasoning on a standard laptop or a single GPU
  • Two GGUF packings with custom ternary hybrid-attention kernels for llama.cpp (CUDA, Metal) — PTQ10 packs trits densely (1.75 bits/weight, 5.95 GB), PQ20 stores each trit in a 2-bit slot (2.13 bits/weight, 7.21 GB); packed weights are consumed directly, never expanded back to FP16

Sources (1)

  • [1]Now you can grow Bonsai on your potato
    r/LocalLLaMA (top, daily) · Oct 11, 12:58 PM
    Full 27B-class reasoning in ternary transformer weights, for llama.cpp (CUDA, Metal, CPU)
    No more excuses for GPU-poor folks not to start LLMing!

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 10, 2026Open-source Mac app that runs EmbeddingGemma 2 locally to search your files by what’s in them
  2. Oct 10, 2026Nace AI Open-Sources Drex 1.5: A 9B Decision Model That Scores Options, Not Text
  3. Oct 9, 2026Qwen/Qwen-Image-2.1-Turbo
  4. Oct 9, 2026Introducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash
  5. Oct 8, 2026ConwayResearch/Underdog-Saluki-27B-1.0
  6. Oct 7, 2026[AINews] Claude Haiku 5.5 — better than GPT-6 Luna at the same pricing

Related