Analysis

Opinion, explainers and guides from people worth reading.

Video
Two Minute Papers (YouTube)6h ago

Why DeepSeek Wants AI To Forget

📝 The DeepSeek OCR paper is available here:

r/LocalLLaMA (top, daily)20h ago

Qwen3.8 Flash Next fixed my GNOME extension

I love Dash2Dock Lite, but Icedman is always a week or two before updates.

Epoch AI: Gradient Updates3d ago

The Epoch Brief - October 8, 2026

Welcome back to the Epoch Brief.

Video
NewAI Engineer (YouTube)2h ago

Taming the AI Hardware Cambrian Explosion — Abdul Dakkak, Modular

Abdul Dakkak, chief scientist at Modular, explains why today's AI software stack is a mess, with vLLM, SGLang, TensorRT-LLM and llama.cpp each patching around different hardware, and how Modular rebuilt the stack from the ground up.

NVIDIA Technical Blog4d ago

Validate AI Factory Changes with Digital Twins and AI Agents

AI factories are some of the most complex operations in the world, combining GPUs, CPUs, switches, DPUs, and SuperNICs alongside schedulers, orchestration...AI factories are some of the most complex operations in the world, combining GPUs, CPUs, switches, DPUs, and SuperNICs alongside schedulers, orchestration services, security controls, and a rapidly changing software stack.

r/StableDiffusion (top, day)10h ago

Updated ComfyUI-SeedVR2-VideoUpscaler-with-TensorRT v1.6.6 - Native ConvRot W4A8 Support & DisTorch2 Scope

The other day, when I published information about v1.5.7, I received a request in the comments regarding support for w4a8.

r/MachineLearning (top, daily)1d ago

Help With Choosing Hardware [P]

I am going to fine-tune a satire model, with the base model being Qwen3.5-14B-Base.

The Decoder14h ago

Cheaper AI tokens are driving more demand, and that's Jensen Huang's best-case scenario

Data from a16z shows a Jevons paradox in the AI market: token prices keep falling, but H100 GPU rental prices hold steady or climb.

Ars Technica: AI9d ago

US arrests tech CEO accused of smuggling $300M in Nvidia chips into China

The US has arrested another suspect accused of smuggling high-end computer servers containing export-controlled Nvidia chips into China.

TechCrunch: AI4h ago

48 hours to TechCrunch Disrupt 2026 — hear directly from the people building what’s next

Less than 48 hours, TechCrunch Disrupt 2026 starts in San Francisco’s Moscone West.

SiliconANGLE: AI1d ago

What to expect during the AI Data Pipeline Forum: Join theCUBE Oct. 13

The post What to expect during the AI Data Pipeline Forum: Join theCUBE Oct.

Simon Willison's Weblog9d ago

Rex's Dino Store

Located just before the turnstiles in the Grand Army Plaza subway station at the north end of Brooklyn's Prospect Park is this former newsstand which is now operated by a dinosaur.

Product Hunt: AI launches2d ago

The Computer Game

Build the best computer, from stone tools to your own chips

IEEE Spectrum: AI2d ago

Master AI Chip Principles With New IEEE Design Program

Today’s engineers face an unprecedented acceleration in AI hardware complexity, as explained in the recent research article “Revisiting Edge AI: Opportunities and Challenges.” The article examines the rapid growth of edge AI and the challenges it creates, including resource constraints, model architecture limitations, and network demands across edge-AI deployments.

Google DeepMind Blog18d ago

Advancing Private AI Compute with secure, server-side memory

A technical update on our Private AI Compute architecture, which will enable persistent, cross-device AI memory with on-device privacy standards.

Google Blog (AI)27d ago

Watch astronaut Christina Koch and Google’s James Manyika discuss space, technology, and discovery.

Christina Koch sits down with James Manyika, Google’s Senior Vice President of Research, Labs, Technology & Society.

r/LocalLLaMA (top, daily)19h ago

Running Next Flash IQ3_XXS at ~70 tok/s with 100k context or 2 instances of Qwen 3.6 35B A3B IQ4 at ~145 tok/s with 256k all on $500 of ex mining BC-250 boards

This will be my third update on the bc-250 cluster and for my first forray into local ai I have been having a blast.

r/LocalLLaMA (top, daily)23h ago

Building a 4x R9700 setup for a 10 person startup

Just wanted to share a build I am doing for a client.

r/LocalLLaMA (top, daily)1d ago

Strata with Qwen3.8 Flash Next UD-Q4_K_XL

Most of the benchmarks I've seen are using IQ2 or IQ3 quants, so I wanted to see how Unsloth's UD-Q4KXL performs instead.

r/LocalLLaMA (top, daily)20h ago

AMD Reportedly Raises GDDR6 Prices for Board Partners

Is an AMD price incoming as well?

r/LocalLLaMA (top, daily)6h ago

vllm-ascend updates (GLM5.3-flash) on dual 310p Ascend cards

Alright, I learend a few valuable lessons since I posted about a week ago, and I've made more strides with the DDR4x Ascend 96GB 310P cards I purchased, so here we go, meet the new friendly and less verbose me.

r/LocalLLaMA (top, daily)1d ago

Is there a better option than llama.cpp for 4GB VRAM for Higher tokens/sec?

I want something that utilizes my system resources more efficiently—for instance, by managing my RTX with 4GB VRAM more intelligently—and delivers higher tokens-per-second, all without the burden of heavy dependencies like PyTorch.

NVIDIA Technical Blog4d ago

The Machines that Make the Machines

However, the process of assembling GB300 trays requires skilled physical labor in factories across the world.

NVIDIA Technical Blog3d ago

Building Reliable Data Analytics Agents: Lessons from the KDD Cup

The NVIDIA KGMON team placed second in the KDD Cup 2026 Data Agents competition with a system built around a simple idea of making an agent's harness smaller,...The NVIDIA KGMON team placed second in the KDD Cup 2026 Data Agents competition with a system built around a simple idea of making an agent’s harness smaller, clearer, and easier to verify.

r/LocalLLaMA (top, daily)1d ago

No more RTX 5090

Nvidia reportedly halts GeForce RTX 5090 production in favor of AI data center and professional GPUs — impending supply drought expected to drive up prices, RTX 5080 24GB rumored as new gaming flagship

Epoch AI: Gradient Updates9d ago

Hundreds of millions of AI agents are coming. Is there work for them?

AI companies are shelling out hundreds of billions of dollars each year on chips and data centers, betting that the massive compute buildouts will run AI agents capable of doing work that people do today.

Ars Technica: AI9d ago

Amazon’s $1B plan to combat data center backlash draws more backlash

On Friday, Amazon committed to donating more than $1 billion over the next five years to communities neighboring data centers.

Epoch AI: Gradient Updates13d ago

AI is getting cheaper faster than any other transformative technology

Because that’s how fast AI is getting cheaper.

TechCrunch: AI1d ago

3 days to TechCrunch Disrupt 2026: Meet the startups before they hit mainstream

TechCrunch Disrupt 2026 takes place October 13-15 in San Francisco.

r/StableDiffusion (top, day)16h ago

4090 VS 5090 MINIMAX H3 SPEED

To those who used to have a 4090, what percentage of speed improvement have you noticed after upgrading to a 5090?

r/LocalLLaMA (top, daily)17h ago

M5 Ultra 256GB or 2x DGX Spark? Which one?

Currently contemplating adding either M5 Ultra 256GB or dual DGX Spark in addition to existing 5090.

r/LocalLLaMA (top, daily)20h ago

Be careful fam - New RTX 4090 fakes hit the market, with GPUs and memory almost indistinguishable from real chips

Be careful out there of deals now.

r/LocalLLaMA (top, daily)21h ago

M5 Ultra Mac 256GB Studio - 64 core vs. 80 core version, is the inference speed difference worth the extra 10 weeks wait time?

I’m seeing 6-7 week delivery wait time for M5 Ultra Mac Studio 64 core version versus 15-17 weeks for 80 core version.

r/LocalLLaMA (top, daily)1d ago

big or small?

what size do you want? tell them on X:

Product Hunt: AI launches6d ago

Ghost Core

A personal AI computer that runs entirely on-device