Models
New and updated models, and how they score.
VideoAI Engineer (YouTube)3h ago
Same Model, Different Speed: Why Your Inference Provider Matters — FriendliAI
Open-weight models are good enough now.
VideoAI Engineer (YouTube)9h ago
Agents That Own Their Inference — Du'an Lightfoot & Khaja Omer, Akamai Technologies
Speculative decoding drops a demo model from roughly 58 tokens per second to 16 instead of speeding it up.