Explore

Everything AION read, in seven sections. Pick one, a topic or a time window.

Video
NewAI Engineer (YouTube)1h ago

Stop Renting Intelligence: The Train-to-Deploy Loop for Specialized AI — Fireworks AI

Jetashree Ravi, who leads part of the applied machine learning team at Fireworks AI, explains how teams move from closed models to open ones without losing quality.

Newr/StableDiffusion (top, day)1h ago

A/B Blind Taste Test - Krea2-Turbo vs. Qwen-Image 2.1 Turbo

Since Qwen 2.1 Turbo dropped suddenly right after my last taste test, figured might as well do one for it as well.

r/LocalLLaMA (top, daily)3h ago

I tested different Qwen 3.8 27B quants

There is a lot of discussion which quant to use.

NewMarkTechPost2h ago

What is Decision 3.0? vLLM Semantic Router’s New Open Decision Models

The vLLM Semantic Router team has released Decision 3.0, a family of multimodal decision models.

Newr/MachineLearning (top, daily)1h ago

I built a browser tool to visualize how a small LLM picks its next word and see what it's attending to [P]

You give it a prompt, it generates one token at a time, and for each step you get the next-token candidates with probabilities, plus the attention from the current token back to the earlier ones drawn as links with the weights labeled.

The Decoder14h ago

ArXiv caps submissions at two per month as AI paper flood overwhelms the preprint server

Starting October 2026, arXiv will cap submissions at two per person per month.

Video
NewAI Engineer (YouTube)2h ago

Parameter Golf with AutoResearch — Vayum Arora, Zhengyao Jiang, Dixing Xu & Dhruv Srikanth, Weco AI

Zhengyao Jiang introduces autoresearch as repeated proposals and evaluations, and Dixing Xu explains the team's Aiden system and its contributions to OpenAI's Parameter Golf challenge.

Video
AI Engineer (YouTube)7h ago

Teaching Agents to Search with NVIDIA Data Designer — Dhruv Nathawani, NVIDIA

Dhruv Nathawani uses that funnel to explain how NVIDIA builds synthetic data that teaches a model to search, rather than answer from memory.

r/MachineLearning (top, daily)9h ago

I trained a 414k-parameter transformer to fly a boids flock, then tested whether the rules a probe can read are the ones it uses [P]

I wrote a small boid simulator (12 birds), recorded it flying, and trained a transformer to predict each bird's next move without it knowing about any boid rules.

r/LocalLLaMA (top, daily)8h ago

Reverse Engineering w/ Local?

Can a local model like Qwen 3.8 27B reverse engineer games and programs?

MarkTechPost17h ago

OrcaRouter Releases OrcaCyber Zero 1.5 Cybersecurity Model With 1M Context

OrcaRouter has released OrcaCyber Zero 1.5, a model for authorized vulnerability research.

r/MachineLearning (top, daily)14h ago

[P] Pecision models that score every allowed label from the logits: Jebadiah v2.1 (27B, 9B), open weights and self-run benchmark results [P]

I've been building open models that treat a decision as a closed-set scoring problem rather than text generation.

r/LocalLLaMA (top, daily)13h ago

I trained a 102M recursive BitNet-v2 model from scratch: 64K context, trained on less than 5B tokens

Hiya, I’m releasing Recursive BitNet N-Gram 102M, a small experiment combining ternary weights, shared transformer layers, and hashed n-gram embeddings, trained with a whooping budget of 100€

r/LocalLLaMA (top, daily)16h ago

Qwen 3.8 27B Q5 vs Qwen 3.8 Next Q3_S for document analysis

So far I've been using Qwen 3.8 27B Q5 with a 150K context window, but I'm wondering whether I should switch to Qwen 3.8 Next Q3S, since it has much more knowledge and could extract data much better than the 27B.

r/StableDiffusion (top, day)12h ago

Building prompts with LLMs

I’ve been into AI image and video generation for a while, and I’ve been a bit obsessed with H3 for the past two months.

r/LocalLLaMA (top, daily)18h ago

Converting dense models into Mixture-of-Experts

For the past few weeks I've been trying out converting existing dense models to sparse Mixture-of-Experts models, with no pretraining from scratch.

r/LocalLLaMA (top, daily)14h ago

Reminder: try probabilistic MTP if you missed it. Decode +14% on prose

Optimal draft-n-max / draft-p-min seem to be in line with greedy sampling.

r/LocalLLaMA (top, daily)19h ago

Running Next Flash IQ3_XXS at ~70 tok/s with 100k context or 2 instances of Qwen 3.6 35B A3B IQ4 at ~145 tok/s with 256k all on $500 of ex mining BC-250 boards

This will be my third update on the bc-250 cluster and for my first forray into local ai I have been having a blast.

r/StableDiffusion (top, day)17h ago

Heretic or Abliterated?

Looking for recommendations for local LLM to convert an image into a usable prompt for Krea2 with max accuracy.

r/LocalLLaMA (top, daily)23h ago

OMG! If you have a Mac with 64GB, try Qwen3.8-Flash-Next-oQ4e-mtp with oMLX!

I was able to run Qwen3.8-Flash-Next-oQ4e-mtp on M3Max 64GB with oMLX!

r/LocalLLaMA (top, daily)17h ago

M5 Ultra 256GB or 2x DGX Spark? Which one?

Currently contemplating adding either M5 Ultra 256GB or dual DGX Spark in addition to existing 5090.