Models
New and updated models, and how they score.

HF: Qwen4 sources9h ago
Qwen/Qwen-Image-2.1-Turbo
Qwen published the model Qwen-Image-2.1-Turbo on Hugging Face.
Weights
VideoNewAI Engineer (YouTube)2h ago
Same Model, Different Speed: Why Your Inference Provider Matters — FriendliAI
Open-weight models are good enough now.

The Decoder6h ago
AI agent teams waste massive tokens for barely measurable quality gains, research finds
Teams of AI agents barely outperform solo agents but cost up to 5.1x more, according to Vals AI.

MarkTechPost15h ago
OrcaRouter Releases OrcaCyber Zero 1.5 Cybersecurity Model With 1M Context
OrcaRouter has released OrcaCyber Zero 1.5, a model for authorized vulnerability research.
VideoAI Engineer (YouTube)4h ago
From Zero to Leaderboard: Agent Evaluation — Wolfram Ravenwolf, Weights & Biases
An agent repairs a broken git repository and earns a perfect score, but the run stays out of the default leaderboard because it tested only one task.
VideoAI Engineer (YouTube)7h ago
Agents That Own Their Inference — Du'an Lightfoot & Khaja Omer, Akamai Technologies
Speculative decoding drops a demo model from roughly 58 tokens per second to 16 instead of speeding it up.
VideoAI Engineer (YouTube)8h ago
Vector Isn't Enough: Hybrid Search & Retrieval — Jeff Vestal & James Williams, Elastic
Jeff Vestal and James Williams work backward from that agent behavior to the retrieval layer in Elasticsearch.