llama.cpp
Also known as: ggml, llamacpp
13stories this week
17last 30 days
21all time
Timeline
- Oct 11, 2026 · Opinion / analysis · 1 sourceOpenMed 3.0 is out: Apache-2.0 clinical AI that runs fully local and never falls back to the cloud. 422 open issues if you want in on 3.1Quick recap: it's an open-source (Apache-2.0) medical AI toolkit with one rule we never break: patient data stays on your machine.
- Oct 11, 2026 · Opinion / analysis · 1 sourceUPDATE: Qwen 3.8 27B 140 tok/s on single RTX 3090 Megakernel: KL divergence 0.0009 vs llama.cppRECAP: The megakernel is a CUDA engine for Qwen3.8-27B that runs 1.4-1.9x faster than llama.cpp on a single 3090.
- Oct 11, 2026 · Opinion / analysis · 1 sourceReminder: try probabilistic MTP if you missed it. Decode +14% on proseOptimal draft-n-max / draft-p-min seem to be in line with greedy sampling.
- Oct 10, 2026 · Opinion / analysis · 1 source[Model] Support MiniCPM-V 4.7 by tc-mb · Pull Request #29416 · ggml-org/llama.cppLet me remind you that MiniCPM-V-4.7-35B-A3B was spotted on r/LocalLLaMA a few days ago (but the model was later hidden on HF).
- Oct 10, 2026 · Opinion / analysis · 1 sourceIs there a better option than llama.cpp for 4GB VRAM for Higher tokens/sec?I want something that utilizes my system resources more efficiently—for instance, by managing my RTX with 4GB VRAM more intelligently—and delivers higher tokens-per-second, all without the burden of heavy dependencies like PyTorch.
- Oct 10, 2026 · Opinion / analysis · 1 sourceQwen3.8-27B on a single 3090: 140 tok/s on code with a custom megakernelI've been using Claude Opus 5.5 to speed up Qwen3.8-27B on my PC (rtx 3090), it wrote a CUDA megakernel that is 1.4-1.9x faster than llama.cpp depending on the task/context length.
- Oct 10, 2026 · Opinion / analysis · 1 sourceOpen-source Mac app that runs EmbeddingGemma 2 locally to search your files by what’s in themDigUp is a free Mac app that runs Google DeepMind’s new EmbeddingGemma 2 locally over your own files.
- Oct 10, 2026 · Opinion / analysis · 1 sourceImprove token per second without touching quantSpent the past month tweaking and experimenting with many different numbers to achieve 30tps.
- Oct 8, 2026 · Open-source release · 1 sourceollama/ollama v0.40.2Models downloaded with earlier versions of Ollama are upgraded in the background the first time you run them, for better performance and compatibility when running on llama.cpp.
- Oct 8, 2026 · Open-source release · 1 sourceunslothai/unsloth v0.1.905-beta: Sandboxing is here!We're introducing Windows, Mac and Linux sandboxing in Unsloth!
- Oct 7, 2026 · Open-source release · 1 sourceunslothai/unsloth v0.1.904-beta: Train your own Decision modelTurn any text or vision LLM into a Jev-style decision model in Unsloth, with decision accuracy going from 30% to 80%.
- Oct 7, 2026 · Research paper · 2 sourcesSystem Switch: When Should a Fast Decision Model Stop and Think?Dual-process agents pair a fast policy with a slow deliberative model.
- Oct 6, 2026 · Open-source release · 1 sourceunslothai/unsloth v0.1.903-beta: New Browser + Voice CloningThis release adds a browser inside Unsloth (browser use coming very soon), so files, web pages and pages the model writes open right beside your chat.
- Sep 29, 2026 · Open-source release · 1 sourceollama/ollama v0.35.1Clef (27B) and Clef Flash (9B) are multimodal: requests can now include images alongside the text state, shared by all questions and scored jointly with it.
- Sep 23, 2026 · Open-source release · 1 sourceollama/ollama v0.34.4Qwen 3.8 prompt processing is faster on Apple Silicon.
- Sep 15, 2026 · Open-source release · 1 sourceollama/ollama v0.34.2Added first-run setup when running ollama, with options to sign in or continue locally.
- Sep 14, 2026 · Open-source release · 1 sourceollama/ollama v0.34.1GGUF model creation now requires using llama.cpp tooling for safetensor conversion and quantization.
- Sep 2, 2026 · Open-source release · 1 sourceollama/ollama v0.33.3gemma4 now supports images and audio on MLX engine
- Aug 26, 2026 · Open-source release · 1 sourceollama/ollama v0.33.1mlxrunner: avoid Metal GPU timeouts when loading models from slow storage
- Aug 19, 2026 · Open-source release · 1 sourceollama/ollama v0.32.15New desktop onboarding flow on first launch
- Jun 9, 2026 · Model release · 1 sourceIntroducing Gemma 4 12B: a unified, encoder-free multimodal modelIntroducing Gemma 4 12B: a unified, encoder-free multimodal model