Analysis
Opinion, explainers and guides from people worth reading.
I Made Terrible Games With Google’s AI Playground
A long day’s haul in the video game slop mines.
Qwen 3.8 27B Q5 vs Qwen 3.8 Next Q3_S for document analysis
So far I've been using Qwen 3.8 27B Q5 with a 150K context window, but I'm wondering whether I should switch to Qwen 3.8 Next Q3S, since it has much more knowledge and could extract data much better than the 27B.
[P] Pecision models that score every allowed label from the logits: Jebadiah v2.1 (27B, 9B), open weights and self-run benchmark results [P]
I've been building open models that treat a decision as a closed-set scoring problem rather than text generation.
VideoLet Your Agent Cook: Using Skills to Evaluate and Improve Your App — Ankur Duggal, Arize AI
A financial agent invents answers, repeats tool calls and produces summaries that its own evaluations flag as wrong.
4090 VS 5090 MINIMAX H3 SPEED
To those who used to have a 4090, what percentage of speed improvement have you noticed after upgrading to a 5090?
Refs
Give Claude better references for your videos
Qwen3.8 27B addition in words
Research: Qwen3.8 27B addition in words
Has machine learning research gotten more "wordy"? [D]
I feel like I am unable to digest most machine learning research these days.
NeurIPS 2026 Paris complimentary registration already full. Did any Top Reviewers get a spot? [D]
Did any reviewers actually manage to secure a spot in Paris, or were all the available places already taken during the earlier registration rounds for SACs and ACs?
Do I need to do anything now to use an ARR August review for NAACL 2027? [D]
If we want to commit to NAACL 2027 instead of EACL, do we need to do anything with the August ARR submission now, or can we simply wait and commit it when the NAACL commitment site opens?
The Computer Game
Build the best computer, from stone tools to your own chips
GitGlow
Review your code and your agents' changes before you ship