Models
New and updated models, and how they score.
VideoNewAI Engineer (YouTube)1h ago
Hill-Climbing Skills: Improve Agents Without Changing the Model — Shubhankar Srivastava, Browserbase
Shubhankar Srivastava uses that uneven progress to show how browser agents can learn a task without changing model weights.
VideoAI Engineer (YouTube)3h ago
Same Model, Different Speed: Why Your Inference Provider Matters — FriendliAI
Open-weight models are good enough now.
VideoAI Engineer (YouTube)6h ago
From Zero to Leaderboard: Agent Evaluation — Wolfram Ravenwolf, Weights & Biases
An agent repairs a broken git repository and earns a perfect score, but the run stays out of the default leaderboard because it tested only one task.