AION
Organization

DeepSeek

9stories this week
10last 30 days
31all time

Timeline

  1. Oct 11, 2026 · Opinion / analysis · 1 source
    PSA: DeepSeek V4.1 Flash habitually exfiltrates API keys. It is dangerously misaligned and may be hazardous to use
    EDIT: since people keep calling it out, this is API key abuse but not exfiltration.
  2. Oct 11, 2026 · Opinion / analysis · 1 source
    Qwen3.8 Flash Next fixed my GNOME extension
    I love Dash2Dock Lite, but Icedman is always a week or two before updates.
  3. Oct 8, 2026 · Research paper · 1 source
    QUILT: Rethinking Sparse-Attention Prefill through Shared Query Execution
    We present QUILT, a workload-aware sparse-attention execution mechanism that jointly processes neighboring queries and reuses shared KV entries to reduce redundant memory traffic and computation.
  4. Oct 7, 2026 · Research paper · 1 source
    EngramEdit: Decoupled Knowledge Updates in LLMs through Conditional Memory
    We propose EngramEdit for decoupled knowledge updates through conditional memory.
  5. Oct 7, 2026 · Research paper · 1 source
    QuSema: Detecting Silent Bugs in Quantum Libraries via Quantum-knowledge-enhanced Agents
    Here we present QuSema, an autonomous testing agent for finding silent bugs in quantum libraries.
  6. Oct 7, 2026 · Research paper · 1 source
    Learning Situation-Conditioned Thinking Policies for Long-Term LLM Agents
    Long-running autonomous agents must reuse accumulated reasoning experience without allowing explicit historical memory and LLM context to grow indefinitely.
  7. Oct 6, 2026 · Product / feature launch · 1 source
    Introducing Mistral Large 4: Le chonk
    Introducing Mistral Large 4: Le chonk
  8. Oct 6, 2026 · Research paper · 1 source
    How Much Evidence Should a Coding Agent's Self-Correction Carry? Adaptive Dirichlet Evidence for Self-Distillation
    Execution feedback lets coding agents revise programs and learn from their own corrections.
  9. Oct 6, 2026 · Research paper · 1 source
    SCOPE: Certified Theorem Proving with a Language Model as the Policy Planner
    Direct generation fails on multi-step numeric propositions: a proof is valid only if every content integer is correct, so the pass rate is bounded by the k-th power of the per-integer accuracy.
  10. Oct 2, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.21
    | Model | Type | Cookbook |
  11. Sep 9, 2026 · Open-source release · 1 source
    huggingface/transformers v5.17.0: Release 5.17.0
    Hy4-Preview is a 780B-parameter mixture-of-experts language model that activates 49B parameters per
  12. Sep 9, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.29.0
    MRV2 also gained CUDA graph memory profiling for KV cache auto-sizing (#53306), batch-sharded sampling that cuts per-step logits memory by 1/TP (#50465), prompt embeds (#42963), extracthiddenstates speculation (#49811), padded FULL cudagraph dispatch for uniform decode under spec decode (#53407), and DP-sync skipping before EAGLE/MTP draft prefill (#53694).
  13. Aug 21, 2026 · Open-source release · 1 source
    ollama/ollama v0.33.0
    Developers can now easily configure Claude Desktop to seamlessly work with Ollama as a third-party gateway provider.
  14. Aug 14, 2026 · Open-source release · 1 source
    ollama/ollama v0.32.11
    The OpenAI-compatible Responses API now supports web search
  15. Aug 10, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.27.0
    Kimi K3 support with a full stack landing in one release: core model files and kernels (#50089, #50000), Python (#50093) and Rust (#50104) frontends, AttnRes kernels (#50090), DeepGEMM support (#50458), compressed-tensors quantized checkpoints (#50500), DSpark AR fusion (#50242), and an option to shard the shared expert instead of replicating it (#50656).
  16. Aug 8, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.17
    Kimi K3 day-0 support: A 2.8T-parameter multimodal LatentMoE (896 experts, top-16, routed in a 3584-dim latent space) with a 1M-token context, 69 KDA linear-attention layers interleaved with 24 MLA layers, and a MoonViT3d vision tower, shipping as a native MXFP4 checkpoint.
  17. Jul 11, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.25.0
    Building on quantized-model support from the previous release, MRv2 is now the standard execution path, with new support for EVS (#46535), realtime embeddings (#46762), prefix caching for Mamba hybrid models (#42406), multimodal-prefix bidirectional attention (#46942), and dynamic speculative decoding compatible with full CUDA graphs (#45953).
  18. Jul 10, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.15
    GLM-5.2 NVFP4, tuned for production: We took time this cycle to tune GLM-5.2 NVFP4 on Blackwell for optimized production serving.
  19. Jun 29, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.24.0
    MiniMax-M3: Added support for the new MiniMax-M3 model (#45381), with a fast follow-on of BF16/FP8 indexer via MSA (#45892), MXFP4 support (#45896), FP8 sparse GQA (#45744), and extensive AMD/ROCm tuning — mxfp8 MoE/linear on gfx950 (#45725), fp8perchannel for bf16 weights on MI300X (#45854), FP8 KV-cache fix (#45720), and packed-modules mapping (#45794).
  20. Jun 26, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.14
    Full release notes by category below.
  21. Jun 13, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.13
    DeepSeek V4 — context parallelism & sparse-attention kernels: Building on the v0.5.12 Day-0 path, v0.5.13 extends DeepSeek-V4 to context-parallel serving and adds its sparse-attention kernels:
  22. Jun 10, 2026 · Open-source release · 1 source
    huggingface/transformers v5.11.0: Release v5.11.0
    DiffusionGemma is engineered to reduce the sequential bottlenecks of standard causal language models by employing an encoder-decoder architecture specifically optimized for inference speed.
  23. Jun 3, 2026 · Open-source release · 1 source
    huggingface/transformers v5.10.1: Release v5.10.1
    Sorry everyone, this happens when we rush a release!!!
  24. May 16, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.12
    DeepSeek V4 support: Full inference path for DeepSeek-V4 (#23882), including:
  25. May 15, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.21.0
    Transformers v4 deprecated: This release formally deprecates transformers v4 support (#40389).
  26. May 5, 2026 · Open-source release · 1 source
    huggingface/transformers v5.8.0: Release 5.8.0
    DeepSeek-V4 is the next-generation MoE (Mixture of Experts) language model from DeepSeek that introduces several architectural innovations over DeepSeek-V3.
  27. Apr 6, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.10
    Piecewise CUDA Graph Enabled by Default: Piecewise CUDA graph capture is now the default execution mode, reducing memory overhead and improving throughput for models with complex control flow patterns: #16331
  28. Mar 28, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.10rc0
    Piecewise CUDA Graph Enabled by Default: Piecewise CUDA graph capture is now the default execution mode, reducing memory overhead and improving throughput for models with complex control flow patterns: #16331
  29. Feb 24, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.9
    TRT-LLM NSA Kernel Integration for DeepSeek V3.2: Integrate TRT-LLM DSA kernels for Native Sparse Attention, boosting DeepSeek V3.2 performance by 3x-5x on Blackwell platforms with trtllm for both --nsa-prefill-backend and --nsa-decode-backend
  30. Jan 23, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.8
    Qwen3-VL-Embedding & Qwen3-VL-Reranker model support: #16635, #16403

Often appears with