Analysis

Opinion, explainers and guides from people worth reading.

Video
AI Engineer (YouTube)5h ago

Teaching Agents to Search with NVIDIA Data Designer — Dhruv Nathawani, NVIDIA

Dhruv Nathawani uses that funnel to explain how NVIDIA builds synthetic data that teaches a model to search, rather than answer from memory.

r/LocalLLaMA (top, daily)14h ago

Qwen 3.8 27B Q5 vs Qwen 3.8 Next Q3_S for document analysis

So far I've been using Qwen 3.8 27B Q5 with a 150K context window, but I'm wondering whether I should switch to Qwen 3.8 Next Q3S, since it has much more knowledge and could extract data much better than the 27B.

r/MachineLearning (top, daily)12h ago

[P] Pecision models that score every allowed label from the logits: Jebadiah v2.1 (27B, 9B), open weights and self-run benchmark results [P]

I've been building open models that treat a decision as a closed-set scoring problem rather than text generation.

Lobsters: ai23h ago

Byte Language Models: Scaling, Emergent Abstractions, and Information Allocation

The paper challenges the assumption that language models need explicit tokenizers to be efficient demonstrating that standard flat Transformers can process raw byte sequences and actually outperform traditional subword models as parameter sizes scale.

Paper
Latent Space1d ago

Why AlphaFold Didn't Solve Protein Folding — Pushmeet Kohli, Google DeepMind & Sal Candido, Biohub

From the Bitter Lesson of AI scaling to the unsolved mysteries of protein folding, Google DeepMind’s Pushmeet Kohli and Biohub’s Sal Candido are rethinking what it takes to build AI that truly understands biology.

r/StableDiffusion (top, day)10h ago

Building prompts with LLMs

I’ve been into AI image and video generation for a while, and I’ve been a bit obsessed with H3 for the past two months.

Simon Willison's Weblog2d ago

ttok 1.0

I released ttok 0.4, ran uv tool upgrade ttok, piped a file into the new version... and realized that it was defaulting to the GPT-4 tokenizer when it should very clearly default to GPT-5/GPT-6 instead!

OpenAI News3d ago

How Oracle turns days of work into minutes with ChatGPT and Codex

Across recruiting, engineering, and operations, Oracle turns specialist knowledge into fast, repeatable workflows with ChatGPT Work and Codex.

The Kaitchup (Benjamin Marie)1d ago

Agentic Coding with Qwen3.8 Flash Next GGUFs and Notes from COLM

Updates on Qwen3.8 Flash Next GGUF evaluations for agentic coding

Ben's Bites3d ago

Text is so 2023

Hark (from Brett Adcock - Figure robots) launched which suggests tasks as one-tap action buttons.

The Decoder12h ago

ArXiv caps submissions at two per month as AI paper flood overwhelms the preprint server

Starting October 2026, arXiv will cap submissions at two per person per month.

Video
bycloud (YouTube)1d ago

DeepSeek-V4.1 Flash: The Most Insane Optimization So Far

DeepSeek-V4.1-Flash is probably the craziest architecture revamp to date.

Ars Technica: AI5d ago

MCP for agent-to-agent comms may be the riskiest protocol you've never heard of

The adoption of AI agents in millions of organizations is creating new opportunities for attackers to make them take malicious actions, such as exfiltrating database contents and sensitive business and personal information.

Hacker News: LLM threads (40+ points)4d ago

Show HN: TerrainSR – fast, realistic heightmap upscaling model

I wanted to have a 1:1 scale model of Europe, but my problem was that 100m data was too low-res while 10m LIDAR data was patchy, took hundreds of GBs to store and was full of manmade objects like mines, buildings and so on.

Weights
r/MachineLearning (top, daily)7h ago

I trained a 414k-parameter transformer to fly a boids flock, then tested whether the rules a probe can read are the ones it uses [P]

I wrote a small boid simulator (12 birds), recorded it flying, and trained a transformer to predict each bird's next move without it knowing about any boid rules.

r/LocalLLaMA (top, daily)17h ago

Running Next Flash IQ3_XXS at ~70 tok/s with 100k context or 2 instances of Qwen 3.6 35B A3B IQ4 at ~145 tok/s with 256k all on $500 of ex mining BC-250 boards

This will be my third update on the bc-250 cluster and for my first forray into local ai I have been having a blast.

r/LocalLLaMA (top, daily)11h ago

I trained a 102M recursive BitNet-v2 model from scratch: 64K context, trained on less than 5B tokens

Hiya, I’m releasing Recursive BitNet N-Gram 102M, a small experiment combining ternary weights, shared transformer layers, and hashed n-gram embeddings, trained with a whooping budget of 100€

r/LocalLLaMA (top, daily)21h ago

OMG! If you have a Mac with 64GB, try Qwen3.8-Flash-Next-oQ4e-mtp with oMLX!

I was able to run Qwen3.8-Flash-Next-oQ4e-mtp on M3Max 64GB with oMLX!

r/LocalLLaMA (top, daily)1d ago

Benefits of using bigger models than Qwen 3.8 flash next?

Qwen 3.8 27b was the first model I tried on Ninfer at NVFP4 and then shifted to Flash next after seeing issues with 27b such as not willing to yield to instructions set in AGENTS.md or agent skills.

r/LocalLLaMA (top, daily)1d ago

Engineer / developer observations of Gemma4-31B, Qwen3.8-27B, and 6.1-Sol for software engineering work

Models: Gemma4-31B vs Qwen3.8-27B at the same quantization (an Unsloth flavor of Q4).

r/LocalLLaMA (top, daily)1d ago

Is anyone running Qwen3.8 Flash Next with a 1M context?

So, while doing a search to read the model card again, and because I didn't memorize the huggingface URL, I saw an AI "answer" at the top of the search, stating that while it natively supports a ~244k context, it could go to 1M using YaRN.

r/LocalLLaMA (top, daily)6h ago

Reverse Engineering w/ Local?

Can a local model like Qwen 3.8 27B reverse engineer games and programs?

r/StableDiffusion (top, day)22h ago

Krea2 Turbo Distill 2 step LoRA - FINAL checkpoint released (chk51195)

Krea 2 Turbo — 2-Step Distillation LoRA (FINAL Version)

r/LocalLLaMA (top, daily)12h ago

Reminder: try probabilistic MTP if you missed it. Decode +14% on prose

Optimal draft-n-max / draft-p-min seem to be in line with greedy sampling.

r/StableDiffusion (top, day)15h ago

Heretic or Abliterated?

Looking for recommendations for local LLM to convert an image into a usable prompt for Krea2 with max accuracy.

Latent Space1d ago

[AINews] TypeSafe/Jev at >$100M ARR, $7.5B valuation 3 weeks after launch

As you can see in the AINews X recap section below, everyone on earth has cloned the Jev API, but only one company can ever create the category.

r/MachineLearning (top, daily)1d ago

I Built a O(NlogN) attention system that retains 97% accuracy over long context (MQAR)[P]

ALHR- Adaptive learnable Hierarchical Routing is a static binary tree based system that uses learnable functions to REDUCE the amount of keys used.

r/LocalLLaMA (top, daily)1d ago

Open-source Mac app that runs EmbeddingGemma 2 locally to search your files by what’s in them

DigUp is a free Mac app that runs Google DeepMind’s new EmbeddingGemma 2 locally over your own files.

r/MachineLearning (top, daily)1d ago

Help With Choosing Hardware [P]

I am going to fine-tune a satire model, with the base model being Qwen3.5-14B-Base.

r/LocalLLaMA (top, daily)15h ago

M5 Ultra 256GB or 2x DGX Spark? Which one?

Currently contemplating adding either M5 Ultra 256GB or dual DGX Spark in addition to existing 5090.

Simon Willison's Weblog4d ago

Anti-Patterns in Software Blogging

Some excellent writing advice from Michael Lynch.

Simon Willison's Weblog6d ago

Qwen3.8 27B addition in words

Research: Qwen3.8 27B addition in words

Lobsters: ai4d ago

Best Books/Courses/Channels to Leapfrog on AI/ML Material

Assuming I have been on hibernation since 2018-19 time frame, which materials - books, MOOC courses and YT channels would be recommended for me to get started with understanding all the current progress on AI/ML.