Same Model, Different Speed: Why Your Inference Provider Matters — FriendliAI
Open-weight models are good enough now.

Proof1 independent outlet
Key points
- Yunmo Koo, founding engineer at FriendliAI, argues that open-weight models like GLM 5.2 and MiniMax M3 now compete with the best proprietary models.
- He explains where the gap comes from: prefix caching with cache-aware routing and a GPU, CPU and NVMe cache hierarchy for long, repetitive agent inputs, plus hybrid speculative decoding that adapts to live traffic.
- • Why the same model feels different on different providers
- 2:24 GLM 5.2 vs Opus: a tower defense game
Sources (1)
- [1]Same Model, Different Speed: Why Your Inference Provider Matters — FriendliAIAI Engineer (YouTube) · Oct 11, 07:30 PM
Open-weight models are good enough now.
Yunmo Koo, founding engineer at FriendliAI, argues that open-weight models like GLM 5.2 and MiniMax M3 now compete with the best proprietary models.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 11, 2026Can Your Agent Hear You Now? Building Live Voice Agents with Gemini — Thor Schaeff
- Oct 11, 2026Agents That Own Their Inference — Du'an Lightfoot & Khaja Omer, Akamai Technologies
- Oct 11, 2026UPDATE: Qwen 3.8 27B 140 tok/s on single RTX 3090 Megakernel: KL divergence 0.0009 vs llama.cpp
- Oct 8, 2026Microsoft Joins the Local AI Push
- Oct 8, 2026Text is so 2023
- Oct 7, 2026[AINews] Claude Haiku 5.5 — better than GPT-6 Luna at the same pricing
