Models

New and updated models, and how they score.

Video
AI Engineer (YouTube)3h ago

Same Model, Different Speed: Why Your Inference Provider Matters — FriendliAI

Open-weight models are good enough now.

Latent Space8d ago

[AINews] not much happened today

AINews’ website lets you search all past issues.

Simon Willison's Weblog5d ago

Mistral Large 4

llm -m gpt-6.1-sol 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars'

The Kaitchup (Benjamin Marie)2d ago

Qwen3.8 27B Xhigh or Flash Next Medium

In my Qwen3.8 27B reasoning benchmarks, Xhigh generally produced the strongest results, but it needed much longer responses than Medium.

Video
Sam Witteveen (YouTube)7d ago

Which is The Best Qwen3.8-27B?

Which is the best fine-tune of Qwen3.8 27B that reduces the amount of thinking and reasoning tokens but still keeps the best accuracy for your particular use case?

MarkTechPost1d ago

Nace AI Open-Sources Drex 1.5: A 9B Decision Model That Scores Options, Not Text

Nace.AI has open-sourced Drex 1.5, a 9B decision model for agents and backend workflows.

GitHub Blog: AI and ML6d ago

ReviewBench: An open benchmark for AI code review

Agentic code review is becoming an essential piece of how development happens.

Hugging Face Blog11d ago

Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning

TLDR 👉 new TTS leaderboard focused on open-source and multilingual

Vendor claim only
Video
Machine Learning Street Talk (YouTube)12d ago

Why Deep Learning Failed on Tables for a Decade - Frank Hutter

Frank Hutter, co-founder of Prior Labs, on why deep learning struggled with tabular data for a decade: tables are messy and heterogeneous, hyped models like TabNet did not generalise to new datasets, and there was no ImageNet of tables.