AION
Model

OLMo

Also known as: Olmo

1stories this week
1last 30 days
3all time

Timeline

  1. Oct 8, 2026 · Research paper · 1 source
    Which Skill to Distill? SGUID: Selecting a Compact Skill Bank for Model-Skill Co-Evolution
    Skills, reusable procedural guidance added at inference, can substantially improve LLM downstream performance (Li et al., 2026).
  2. Sep 9, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.29.0
    MRV2 also gained CUDA graph memory profiling for KV cache auto-sizing (#53306), batch-sharded sampling that cuts per-step logits memory by 1/TP (#50465), prompt embeds (#42963), extracthiddenstates speculation (#49811), padded FULL cudagraph dispatch for uniform decode under spec decode (#53407), and DP-sync skipping before EAGLE/MTP draft prefill (#53694).
  3. Jul 27, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.26.0
    New Inkling model family with a full support stack: base modeling (#48799), piecewise CUDA graph support (#48822), Hopper FA4 relative attention (#48858), MTP=1 speculative decoding (#48869), LoRA (#48884), and standard ModelOpt NVFP4 quantization (#48990).

Often appears with