Gemma
14stories this week
18last 30 days
42all time
Timeline
- Oct 10, 2026 · Opinion / analysis · 1 sourceEngineer / developer observations of Gemma4-31B, Qwen3.8-27B, and 6.1-Sol for software engineering workModels: Gemma4-31B vs Qwen3.8-27B at the same quantization (an Unsloth flavor of Q4).
- Oct 8, 2026 · Research paper · 1 sourceBeliefScope: Diagnosing Evidence-Driven Revision and Pressure-Induced Shifts in Large Language ModelsWe introduce BeliefScope, a controlled black-box framework for separating these two sources of influence around a fixed target proposition.
- Oct 7, 2026 · Research paper · 1 sourceAmortized Off-Policy Evaluation for LLMsTo address this, we propose PFN-OPE, a prior-data fitted network that amortizes OPE across a distribution of contextual-bandit tasks.
- Oct 7, 2026 · Research paper · 2 sourcesReal Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than RecomputeWe test a memory layer, the public package galahad-kv, that saves the KV state of each block of about 16,000 tokens to encrypted local NVMe disk and loads it back later, byte-exact, without recomputing it.
- Oct 7, 2026 · Research paper · 1 sourceInsights from Autoresearch for Solar Panel SegmentationThis paper investigates AutoResearch, a protocol in which a coding language model edits a training program under a one-hour GPU budget and retains a change only if validation IoU improves.
- Oct 7, 2026 · Research paper · 1 sourceDecoupling Logic from Persona: Structural Immunity of Edge LLM Agents to Context PollutionWe study what happens to the logical part of such an agent when that history is long, misleading and persona-heavy (persona-logic interference), and present a Decoupling Architecture (AO-DA) that separates logical inference ("What") from persona expression ("How") into two inference paths on one INT4 base model with hot-swappable LoRA adapters.
- Oct 6, 2026 · Open-source release · 1 sourcehuggingface/trl v1.14.2Patch release fixing two cases of silently wrong training and three crashes.
- Oct 6, 2026 · Product / feature launch · 1 sourceEmbeddingGemma 2: an open, lightweight multimodal embedding modelEmbeddingGemma 2: an open, lightweight multimodal embedding model
- Oct 6, 2026 · Research paper · 1 sourceNot Every Call Needs a Frontier Model: Per-Call-Site Evaluation of Small Language Models in a Deployed Agentic Home-Automation SystemAn agentic system issues several structurally different kinds of LLM calls.
- Oct 6, 2026 · Research paper · 1 sourceCoverage, Not Difficulty, Sets How Much Synthetic Data an Activation Probe NeedsActivation probes that monitor deployed language models are trained on synthetic conversations, and how many a probe needs is open.
- Oct 6, 2026 · Open-source release · 1 sourcehuggingface/transformers v5.19.0: Release v5.19.0EmbeddingGemma 2 is a multimodal embedding model from Google built on the Gemma 4 architecture.
- Oct 6, 2026 · Research paper · 1 sourceMINDSET: Energy-based Schema Evolution for Long Conversational Agent MemoryWe introduce MINDSET, a memory controller that stores a conversation as immutable episodes and organizes them into versioned schemas through minimum-energy state transitions.
- Oct 6, 2026 · Research paper · 1 sourceMASKerade: Token-Routed Mask Experts for Dense-to-MoE UpcyclingWe introduce MASKerade, a dense-to-MoE training method that instead learns experts as sparse subnetworks of a frozen pretrained FFN.
- Oct 5, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.31.0Fast restart: the new vllm preload CLI launches the weight-cache daemon that keeps post-quantized weights resident in GPU memory across engine restarts (#56680), now with data parallelism (#57386), MTP draft models (#57312), a /health endpoint (#58552) and a readiness wait (#58370).
- Oct 1, 2026 · Product / feature launch · 1 sourceunslothai/unsloth v0.1.902-beta: Command Palette + Desktop UI/UXThis release brings faster navigation, shareable run settings, and clearer errors to Unsloth Desktop.
- Sep 29, 2026 · Open-source release · 1 sourceNVIDIA/TensorRT-LLM v1.3.0rc29Expose Nemotron-H vision-language LoRA configuration for supported inference paths #19151
- Sep 28, 2026 · Open-source release · 1 sourceunslothai/unsloth v0.1.900-beta: Laya Decision Models + LibraryRun and serve Decision Models like Laya (open-source Jev) locally
- Sep 23, 2026 · Open-source release · 1 sourceollama/ollama v0.34.4Qwen 3.8 prompt processing is faster on Apple Silicon.
- Sep 8, 2026 · Opinion / analysis · 1 sourceAlphaGenome Atlas: A predictive map of every possible DNA letter change in the human genomeWe are grateful to our research collaborators at University of Exeter, Broad Institute, Boston Children’s Hospital, Stowers Institute for Medical Research, Harvard University, Memorial Sloan Kettering Cancer Center, Center for Genomic Medicine at Massachusetts General Hospital, and the University of Kansas Medical Center.
- Aug 10, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.27.0Kimi K3 support with a full stack landing in one release: core model files and kernels (#50089, #50000), Python (#50093) and Rust (#50104) frontends, AttnRes kernels (#50090), DeepGEMM support (#50458), compressed-tensors quantized checkpoints (#50500), DSpark AR fusion (#50242), and an option to shard the shared expert instead of replicating it (#50656).
- Aug 10, 2026 · Open-source release · 1 sourcehuggingface/transformers v5.15.0: Release: v5.15.0Muse Glimmer, released today, is Meta’s new multimodal model, especially designed for agentic use cases.
- Jul 27, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.26.0New Inkling model family with a full support stack: base modeling (#48799), piecewise CUDA graph support (#48822), Hopper FA4 relative attention (#48858), MTP=1 speculative decoding (#48869), LoRA (#48884), and standard ModelOpt NVFP4 quantization (#48990).
- Jul 14, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.25.1Avoid blocking model launching when no system FFmpeg is available for TorchCodec (#47888).
- Jul 11, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.25.0Building on quantized-model support from the previous release, MRv2 is now the standard execution path, with new support for EVS (#46535), realtime embeddings (#46762), prefix caching for Mamba hybrid models (#42406), multimodal-prefix bidirectional attention (#46942), and dynamic speculative decoding compatible with full CUDA graphs (#45953).
- Jun 29, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.24.0MiniMax-M3: Added support for the new MiniMax-M3 model (#45381), with a fast follow-on of BF16/FP8 indexer via MSA (#45892), MXFP4 support (#45896), FP8 sparse GQA (#45744), and extensive AMD/ROCm tuning — mxfp8 MoE/linear on gfx950 (#45725), fp8perchannel for bf16 weights on MI300X (#45854), FP8 KV-cache fix (#45720), and packed-modules mapping (#45794).
- Jun 15, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.23.0DeepSeek-V4 matures across backends: Following its introduction in v0.22.0, DeepSeek-V4 received another large hardening and optimization pass.
- Jun 10, 2026 · Open-source release · 1 sourcehuggingface/transformers v5.11.0: Release v5.11.0DiffusionGemma is engineered to reduce the sequential bottlenecks of standard causal language models by employing an encoder-decoder architecture specifically optimized for inference speed.
- Jun 10, 2026 · Model release · 1 sourceDiffusionGemma: 4x faster text generationToday, we’re introducing DiffusionGemma, an experimental open model that explores text diffusion, an exceptionally fast approach to text generation.
- Jun 9, 2026 · Model release · 1 sourceIntroducing Gemma 4 12B: a unified, encoder-free multimodal modelIntroducing Gemma 4 12B: a unified, encoder-free multimodal model
- Jun 3, 2026 · Open-source release · 1 sourcehuggingface/transformers v5.10.1: Release v5.10.1Sorry everyone, this happens when we rush a release!!!