ModelsBenchmark resultEfficiency & Inference · Evaluation & Benchmarks · Reinforcement Learning1 source · Oct 11, 2026

Same Model, Different Speed: Why Your Inference Provider Matters — FriendliAI

Open-weight models are good enough now.

Proof1 independent outlet

Key points

  • Yunmo Koo, founding engineer at FriendliAI, argues that open-weight models like GLM 5.2 and MiniMax M3 now compete with the best proprietary models.
  • He explains where the gap comes from: prefix caching with cache-aware routing and a GPU, CPU and NVMe cache hierarchy for long, repetitive agent inputs, plus hybrid speculative decoding that adapts to live traffic.
  • • Why the same model feels different on different providers
  • 2:24 GLM 5.2 vs Opus: a tower defense game

Sources (1)

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 11, 2026Can Your Agent Hear You Now? Building Live Voice Agents with Gemini — Thor Schaeff
  2. Oct 11, 2026Agents That Own Their Inference — Du'an Lightfoot & Khaja Omer, Akamai Technologies
  3. Oct 11, 2026UPDATE: Qwen 3.8 27B 140 tok/s on single RTX 3090 Megakernel: KL divergence 0.0009 vs llama.cpp
  4. Oct 8, 2026Microsoft Joins the Local AI Push
  5. Oct 8, 2026Text is so 2023
  6. Oct 7, 2026[AINews] Claude Haiku 5.5 — better than GPT-6 Luna at the same pricing

Related