AION
Repository / library

llama.cpp

Also known as: ggml, llamacpp

13stories this week
17last 30 days
21all time

Timeline

  1. Oct 11, 2026 · Opinion / analysis · 1 source
    OpenMed 3.0 is out: Apache-2.0 clinical AI that runs fully local and never falls back to the cloud. 422 open issues if you want in on 3.1
    Quick recap: it's an open-source (Apache-2.0) medical AI toolkit with one rule we never break: patient data stays on your machine.
  2. Oct 11, 2026 · Opinion / analysis · 1 source
    UPDATE: Qwen 3.8 27B 140 tok/s on single RTX 3090 Megakernel: KL divergence 0.0009 vs llama.cpp
    RECAP: The megakernel is a CUDA engine for Qwen3.8-27B that runs 1.4-1.9x faster than llama.cpp on a single 3090.
  3. Oct 11, 2026 · Opinion / analysis · 1 source
    Reminder: try probabilistic MTP if you missed it. Decode +14% on prose
    Optimal draft-n-max / draft-p-min seem to be in line with greedy sampling.
  4. Oct 10, 2026 · Opinion / analysis · 1 source
    [Model] Support MiniCPM-V 4.7 by tc-mb · Pull Request #29416 · ggml-org/llama.cpp
    Let me remind you that MiniCPM-V-4.7-35B-A3B was spotted on r/LocalLLaMA a few days ago (but the model was later hidden on HF).
  5. Oct 10, 2026 · Opinion / analysis · 1 source
    Is there a better option than llama.cpp for 4GB VRAM for Higher tokens/sec?
    I want something that utilizes my system resources more efficiently—for instance, by managing my RTX with 4GB VRAM more intelligently—and delivers higher tokens-per-second, all without the burden of heavy dependencies like PyTorch.
  6. Oct 10, 2026 · Opinion / analysis · 1 source
    Qwen3.8-27B on a single 3090: 140 tok/s on code with a custom megakernel
    I've been using Claude Opus 5.5 to speed up Qwen3.8-27B on my PC (rtx 3090), it wrote a CUDA megakernel that is 1.4-1.9x faster than llama.cpp depending on the task/context length.
  7. Oct 10, 2026 · Opinion / analysis · 1 source
    Open-source Mac app that runs EmbeddingGemma 2 locally to search your files by what’s in them
    DigUp is a free Mac app that runs Google DeepMind’s new EmbeddingGemma 2 locally over your own files.
  8. Oct 10, 2026 · Opinion / analysis · 1 source
    Improve token per second without touching quant
    Spent the past month tweaking and experimenting with many different numbers to achieve 30tps.
  9. Oct 8, 2026 · Open-source release · 1 source
    ollama/ollama v0.40.2
    Models downloaded with earlier versions of Ollama are upgraded in the background the first time you run them, for better performance and compatibility when running on llama.cpp.
  10. Oct 8, 2026 · Open-source release · 1 source
    unslothai/unsloth v0.1.905-beta: Sandboxing is here!
    We're introducing Windows, Mac and Linux sandboxing in Unsloth!
  11. Oct 7, 2026 · Open-source release · 1 source
    unslothai/unsloth v0.1.904-beta: Train your own Decision model
    Turn any text or vision LLM into a Jev-style decision model in Unsloth, with decision accuracy going from 30% to 80%.
  12. Oct 7, 2026 · Research paper · 2 sources
    System Switch: When Should a Fast Decision Model Stop and Think?
    Dual-process agents pair a fast policy with a slow deliberative model.
  13. Oct 6, 2026 · Open-source release · 1 source
    unslothai/unsloth v0.1.903-beta: New Browser + Voice Cloning
    This release adds a browser inside Unsloth (browser use coming very soon), so files, web pages and pages the model writes open right beside your chat.
  14. Sep 29, 2026 · Open-source release · 1 source
    ollama/ollama v0.35.1
    Clef (27B) and Clef Flash (9B) are multimodal: requests can now include images alongside the text state, shared by all questions and scored jointly with it.
  15. Sep 23, 2026 · Open-source release · 1 source
    ollama/ollama v0.34.4
    Qwen 3.8 prompt processing is faster on Apple Silicon.
  16. Sep 15, 2026 · Open-source release · 1 source
    ollama/ollama v0.34.2
    Added first-run setup when running ollama, with options to sign in or continue locally.
  17. Sep 14, 2026 · Open-source release · 1 source
    ollama/ollama v0.34.1
    GGUF model creation now requires using llama.cpp tooling for safetensor conversion and quantization.
  18. Sep 2, 2026 · Open-source release · 1 source
    ollama/ollama v0.33.3
    gemma4 now supports images and audio on MLX engine
  19. Aug 26, 2026 · Open-source release · 1 source
    ollama/ollama v0.33.1
    mlxrunner: avoid Metal GPU timeouts when loading models from slow storage
  20. Aug 19, 2026 · Open-source release · 1 source
    ollama/ollama v0.32.15
    New desktop onboarding flow on first launch
  21. Jun 9, 2026 · Model release · 1 source
    Introducing Gemma 4 12B: a unified, encoder-free multimodal model
    Introducing Gemma 4 12B: a unified, encoder-free multimodal model

Often appears with