DeepSeek
9stories this week
10last 30 days
31all time
Timeline
- Oct 11, 2026 · Opinion / analysis · 1 sourcePSA: DeepSeek V4.1 Flash habitually exfiltrates API keys. It is dangerously misaligned and may be hazardous to useEDIT: since people keep calling it out, this is API key abuse but not exfiltration.
- Oct 11, 2026 · Opinion / analysis · 1 sourceQwen3.8 Flash Next fixed my GNOME extensionI love Dash2Dock Lite, but Icedman is always a week or two before updates.
- Oct 8, 2026 · Research paper · 1 sourceQUILT: Rethinking Sparse-Attention Prefill through Shared Query ExecutionWe present QUILT, a workload-aware sparse-attention execution mechanism that jointly processes neighboring queries and reuses shared KV entries to reduce redundant memory traffic and computation.
- Oct 7, 2026 · Research paper · 1 sourceEngramEdit: Decoupled Knowledge Updates in LLMs through Conditional MemoryWe propose EngramEdit for decoupled knowledge updates through conditional memory.
- Oct 7, 2026 · Research paper · 1 sourceQuSema: Detecting Silent Bugs in Quantum Libraries via Quantum-knowledge-enhanced AgentsHere we present QuSema, an autonomous testing agent for finding silent bugs in quantum libraries.
- Oct 7, 2026 · Research paper · 1 sourceLearning Situation-Conditioned Thinking Policies for Long-Term LLM AgentsLong-running autonomous agents must reuse accumulated reasoning experience without allowing explicit historical memory and LLM context to grow indefinitely.
- Oct 6, 2026 · Product / feature launch · 1 sourceIntroducing Mistral Large 4: Le chonkIntroducing Mistral Large 4: Le chonk
- Oct 6, 2026 · Research paper · 1 sourceHow Much Evidence Should a Coding Agent's Self-Correction Carry? Adaptive Dirichlet Evidence for Self-DistillationExecution feedback lets coding agents revise programs and learn from their own corrections.
- Oct 6, 2026 · Research paper · 1 sourceSCOPE: Certified Theorem Proving with a Language Model as the Policy PlannerDirect generation fails on multi-step numeric propositions: a proof is valid only if every content integer is correct, so the pass rate is bounded by the k-th power of the per-integer accuracy.
- Oct 2, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.21| Model | Type | Cookbook |
- Sep 9, 2026 · Open-source release · 1 sourcehuggingface/transformers v5.17.0: Release 5.17.0Hy4-Preview is a 780B-parameter mixture-of-experts language model that activates 49B parameters per
- Sep 9, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.29.0MRV2 also gained CUDA graph memory profiling for KV cache auto-sizing (#53306), batch-sharded sampling that cuts per-step logits memory by 1/TP (#50465), prompt embeds (#42963), extracthiddenstates speculation (#49811), padded FULL cudagraph dispatch for uniform decode under spec decode (#53407), and DP-sync skipping before EAGLE/MTP draft prefill (#53694).
- Aug 21, 2026 · Open-source release · 1 sourceollama/ollama v0.33.0Developers can now easily configure Claude Desktop to seamlessly work with Ollama as a third-party gateway provider.
- Aug 14, 2026 · Open-source release · 1 sourceollama/ollama v0.32.11The OpenAI-compatible Responses API now supports web search
- Aug 10, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.27.0Kimi K3 support with a full stack landing in one release: core model files and kernels (#50089, #50000), Python (#50093) and Rust (#50104) frontends, AttnRes kernels (#50090), DeepGEMM support (#50458), compressed-tensors quantized checkpoints (#50500), DSpark AR fusion (#50242), and an option to shard the shared expert instead of replicating it (#50656).
- Aug 8, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.17Kimi K3 day-0 support: A 2.8T-parameter multimodal LatentMoE (896 experts, top-16, routed in a 3584-dim latent space) with a 1M-token context, 69 KDA linear-attention layers interleaved with 24 MLA layers, and a MoonViT3d vision tower, shipping as a native MXFP4 checkpoint.
- Jul 11, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.25.0Building on quantized-model support from the previous release, MRv2 is now the standard execution path, with new support for EVS (#46535), realtime embeddings (#46762), prefix caching for Mamba hybrid models (#42406), multimodal-prefix bidirectional attention (#46942), and dynamic speculative decoding compatible with full CUDA graphs (#45953).
- Jul 10, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.15GLM-5.2 NVFP4, tuned for production: We took time this cycle to tune GLM-5.2 NVFP4 on Blackwell for optimized production serving.
- Jun 29, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.24.0MiniMax-M3: Added support for the new MiniMax-M3 model (#45381), with a fast follow-on of BF16/FP8 indexer via MSA (#45892), MXFP4 support (#45896), FP8 sparse GQA (#45744), and extensive AMD/ROCm tuning — mxfp8 MoE/linear on gfx950 (#45725), fp8perchannel for bf16 weights on MI300X (#45854), FP8 KV-cache fix (#45720), and packed-modules mapping (#45794).
- Jun 26, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.14Full release notes by category below.
- Jun 13, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.13DeepSeek V4 — context parallelism & sparse-attention kernels: Building on the v0.5.12 Day-0 path, v0.5.13 extends DeepSeek-V4 to context-parallel serving and adds its sparse-attention kernels:
- Jun 10, 2026 · Open-source release · 1 sourcehuggingface/transformers v5.11.0: Release v5.11.0DiffusionGemma is engineered to reduce the sequential bottlenecks of standard causal language models by employing an encoder-decoder architecture specifically optimized for inference speed.
- Jun 3, 2026 · Open-source release · 1 sourcehuggingface/transformers v5.10.1: Release v5.10.1Sorry everyone, this happens when we rush a release!!!
- May 16, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.12DeepSeek V4 support: Full inference path for DeepSeek-V4 (#23882), including:
- May 15, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.21.0Transformers v4 deprecated: This release formally deprecates transformers v4 support (#40389).
- May 5, 2026 · Open-source release · 1 sourcehuggingface/transformers v5.8.0: Release 5.8.0DeepSeek-V4 is the next-generation MoE (Mixture of Experts) language model from DeepSeek that introduces several architectural innovations over DeepSeek-V3.
- Apr 6, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.10Piecewise CUDA Graph Enabled by Default: Piecewise CUDA graph capture is now the default execution mode, reducing memory overhead and improving throughput for models with complex control flow patterns: #16331
- Mar 28, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.10rc0Piecewise CUDA Graph Enabled by Default: Piecewise CUDA graph capture is now the default execution mode, reducing memory overhead and improving throughput for models with complex control flow patterns: #16331
- Feb 24, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.9TRT-LLM NSA Kernel Integration for DeepSeek V3.2: Integrate TRT-LLM DSA kernels for Native Sparse Attention, boosting DeepSeek V3.2 performance by 3x-5x on Blackwell platforms with trtllm for both --nsa-prefill-backend and --nsa-decode-backend
- Jan 23, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.8Qwen3-VL-Embedding & Qwen3-VL-Reranker model support: #16635, #16403