Models
New and updated models, and how they score.
VideoNewAI Engineer (YouTube)2h ago
Same Model, Different Speed: Why Your Inference Provider Matters — FriendliAI
Open-weight models are good enough now.

The Kaitchup (Benjamin Marie)2d ago
Qwen3.8 27B Xhigh or Flash Next Medium
In my Qwen3.8 27B reasoning benchmarks, Xhigh generally produced the strongest results, but it needed much longer responses than Medium.
NVIDIA Technical Blog5d ago
Scale Bitwise-Deterministic Pretraining with NVIDIA Megatron Core
Bitwise determinism makes large-scale pretraining easier to debug, validate, and resume reproducibly.
Vendor claim only
SiliconANGLE: AI2d ago
CoreWeave targets AI inference bottlenecks with full-stack optimization
The post CoreWeave targets AI inference bottlenecks with full-stack optimization appeared first on SiliconANGLE.
VideoAI Engineer (YouTube)7h ago
Agents That Own Their Inference — Du'an Lightfoot & Khaja Omer, Akamai Technologies
Speculative decoding drops a demo model from roughly 58 tokens per second to 16 instead of speeding it up.

The Kaitchup (Benjamin Marie)5d ago
Qwen3.8 Flash Next Reasoning Modes: Off vs Low vs Medium vs Xhigh
Like Qwen3.8 27B, Qwen3.8 Flash Next has three reasoning efforts: low, medium, and xhigh.