OLMo
Also known as: Olmo
1stories this week
1last 30 days
3all time
Timeline
- Oct 8, 2026 · Research paper · 1 sourceWhich Skill to Distill? SGUID: Selecting a Compact Skill Bank for Model-Skill Co-EvolutionSkills, reusable procedural guidance added at inference, can substantially improve LLM downstream performance (Li et al., 2026).
- Sep 9, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.29.0MRV2 also gained CUDA graph memory profiling for KV cache auto-sizing (#53306), batch-sharded sampling that cuts per-step logits memory by 1/TP (#50465), prompt embeds (#42963), extracthiddenstates speculation (#49811), padded FULL cudagraph dispatch for uniform decode under spec decode (#53407), and DP-sync skipping before EAGLE/MTP draft prefill (#53694).
- Jul 27, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.26.0New Inkling model family with a full support stack: base modeling (#48799), piecewise CUDA graph support (#48822), Hopper FA4 relative attention (#48858), MTP=1 speculative decoding (#48869), LoRA (#48884), and standard ModelOpt NVFP4 quantization (#48990).