Explore

Everything AION read, in seven sections. Pick one, a topic or a time window.

Video
Two Minute Papers (YouTube)4h ago

Why DeepSeek Wants AI To Forget

📝 The DeepSeek OCR paper is available here:

r/LocalLLaMA (top, daily)4h ago

vllm-ascend updates (GLM5.3-flash) on dual 310p Ascend cards

Alright, I learend a few valuable lessons since I posted about a week ago, and I've made more strides with the DDR4x Ascend 96GB 310P cards I purchased, so here we go, meet the new friendly and less verbose me.

Video
AI Engineer (YouTube)3h ago

AI Apps in a Flash: Ship to GPUs Without Docker — Dean Quiñanola, Runpod

Dean Quiñanola, staff engineer at Runpod, introduces Runpod Flash, which lets you run Python on cloud GPUs as if the GPU were local, with no Dockerfiles, image builds or registry pushes.

r/StableDiffusion (top, day)8h ago

Updated ComfyUI-SeedVR2-VideoUpscaler-with-TensorRT v1.6.6 - Native ConvRot W4A8 Support & DisTorch2 Scope

The other day, when I published information about v1.5.7, I received a request in the comments regarding support for w4a8.

Ars Technica: AI2 sources3d ago

Microsoft event debuts new AI-friendly hardware and Windows changes

In its first live event in two years, Microsoft today announced its newest Surface laptop, outlined a host of changes coming to Windows 11, and shared a vision of how local-AI and agentic workflows could reshape personal computing for those in the dev community as well as home enthusiasts.

2 outlets
The Decoder11h ago

Cheaper AI tokens are driving more demand, and that's Jensen Huang's best-case scenario

Data from a16z shows a Jevons paradox in the AI market: token prices keep falling, but H100 GPU rental prices hold steady or climb.

NVIDIA Technical Blog4d ago

Scaling Decision Optimization to 100 Million Variables and Beyond with mPDLP in NVIDIA cuOpt

NVIDIA cuOpt GPU-accelerated decision optimization can already deliver speedups of more than 10x over CPU...

Vendor claim only
NewTechCrunch: AI2h ago

48 hours to TechCrunch Disrupt 2026 — hear directly from the people building what’s next

Less than 48 hours, TechCrunch Disrupt 2026 starts in San Francisco’s Moscone West.

r/MachineLearning (top, daily)1d ago

Help With Choosing Hardware [P]

I am going to fine-tune a satire model, with the base model being Qwen3.5-14B-Base.

Epoch AI: Gradient Updates3d ago

The Epoch Brief - October 8, 2026

Welcome back to the Epoch Brief.

Hugging Face Blog2d ago

Impactful scheduling for GPU clusters

On the AI Infrastructure team at Ai2, we’re responsible for providing the institute’s GPU compute capacity, specifically targeting large, distributed training workloads.

Vendor claim only
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

GLIO2: A GPU-Parallelized Tightly-Coupled LiDAR-Inertial-GNSS System for Robust and Real-Time Global Localization and Mapping

We propose GLIO2, a tightly-coupled LiDAR-Inertial-GNSS system whose GPU-parallel front-end jointly optimizes scan-to-multiscan LiDAR, IMU pre-integration, and raw GNSS measurements in a single sliding-window factor graph, sustaining real-time operation on edge hardware.

Paper
SiliconANGLE: AI23h ago

What to expect during the AI Data Pipeline Forum: Join theCUBE Oct. 13

The post What to expect during the AI Data Pipeline Forum: Join theCUBE Oct.

Together AI Blog5d ago

Expanding our enterprise inference capacity with IBM Cloud and NVIDIA

Enterprises can now run open models at production scale on a dedicated B300 inference cluster, built by Together AI, IBM Cloud, and NVIDIA

Product Hunt: AI launches2d ago

The Computer Game

Build the best computer, from stone tools to your own chips

IEEE Spectrum: AI2d ago

Master AI Chip Principles With New IEEE Design Program

Today’s engineers face an unprecedented acceleration in AI hardware complexity, as explained in the recent research article “Revisiting Edge AI: Opportunities and Challenges.” The article examines the rapid growth of edge AI and the challenges it creates, including resource constraints, model architecture limitations, and network demands across edge-AI deployments.

r/LocalLLaMA (top, daily)17h ago

Qwen3.8 Flash Next fixed my GNOME extension

I love Dash2Dock Lite, but Icedman is always a week or two before updates.

r/LocalLLaMA (top, daily)17h ago

Running Next Flash IQ3_XXS at ~70 tok/s with 100k context or 2 instances of Qwen 3.6 35B A3B IQ4 at ~145 tok/s with 256k all on $500 of ex mining BC-250 boards

This will be my third update on the bc-250 cluster and for my first forray into local ai I have been having a blast.

r/LocalLLaMA (top, daily)21h ago

Building a 4x R9700 setup for a 10 person startup

Just wanted to share a build I am doing for a client.

r/LocalLLaMA (top, daily)18h ago

AMD Reportedly Raises GDDR6 Prices for Board Partners

Is an AMD price incoming as well?

r/LocalLLaMA (top, daily)1d ago

Strata with Qwen3.8 Flash Next UD-Q4_K_XL

Most of the benchmarks I've seen are using IQ2 or IQ3 quants, so I wanted to see how Unsloth's UD-Q4KXL performs instead.

r/LocalLLaMA (top, daily)1d ago

Is there a better option than llama.cpp for 4GB VRAM for Higher tokens/sec?

I want something that utilizes my system resources more efficiently—for instance, by managing my RTX with 4GB VRAM more intelligently—and delivers higher tokens-per-second, all without the burden of heavy dependencies like PyTorch.

r/LocalLLaMA (top, daily)1d ago

No more RTX 5090

Nvidia reportedly halts GeForce RTX 5090 production in favor of AI data center and professional GPUs — impending supply drought expected to drive up prices, RTX 5080 24GB rumored as new gaming flagship

r/StableDiffusion (top, day)14h ago

4090 VS 5090 MINIMAX H3 SPEED

To those who used to have a 4090, what percentage of speed improvement have you noticed after upgrading to a 5090?

r/LocalLLaMA (top, daily)15h ago

M5 Ultra 256GB or 2x DGX Spark? Which one?

Currently contemplating adding either M5 Ultra 256GB or dual DGX Spark in addition to existing 5090.

r/LocalLLaMA (top, daily)18h ago

Be careful fam - New RTX 4090 fakes hit the market, with GPUs and memory almost indistinguishable from real chips

Be careful out there of deals now.

r/LocalLLaMA (top, daily)18h ago

M5 Ultra Mac 256GB Studio - 64 core vs. 80 core version, is the inference speed difference worth the extra 10 weeks wait time?

I’m seeing 6-7 week delivery wait time for M5 Ultra Mac Studio 64 core version versus 15-17 weeks for 80 core version.

NVIDIA Technical Blog3d ago

Building Reliable Data Analytics Agents: Lessons from the KDD Cup

The NVIDIA KGMON team placed second in the KDD Cup 2026 Data Agents competition with a system built around a simple idea of making an agent's harness smaller,...The NVIDIA KGMON team placed second in the KDD Cup 2026 Data Agents competition with a system built around a simple idea of making an agent’s harness smaller, clearer, and easier to verify.

NVIDIA Technical Blog4d ago

Faster Scientific Image Analysis with NVIDIA cuPhoton

Observatories and telescopes, lasers and X-ray light sources, and other high-throughput instruments generate image data faster than CPU-bound pipelines can...Observatories and telescopes, lasers and X-ray light sources, and other high-throughput instruments generate image data faster than CPU-bound pipelines can process it to support timely decisions.

TechCrunch: AI1d ago

3 days to TechCrunch Disrupt 2026: Meet the startups before they hit mainstream

TechCrunch Disrupt 2026 takes place October 13-15 in San Francisco.

NVIDIA Technical Blog4d ago

Validate AI Factory Changes with Digital Twins and AI Agents

AI factories are some of the most complex operations in the world, combining GPUs, CPUs, switches, DPUs, and SuperNICs alongside schedulers, orchestration...AI factories are some of the most complex operations in the world, combining GPUs, CPUs, switches, DPUs, and SuperNICs alongside schedulers, orchestration services, security controls, and a rapidly changing software stack.

NVIDIA Technical Blog4d ago

The Machines that Make the Machines

However, the process of assembling GB300 trays requires skilled physical labor in factories across the world.

r/LocalLLaMA (top, daily)1d ago

big or small?

what size do you want? tell them on X:

NVIDIA Technical Blog5d ago

How DOCA GPUNetIO Unifies GPU-Initiated Networking Across the NVIDIA Software Stack

GPU applications increasingly need networking and data movement to behave like first-class GPU-controlled operations rather than host-driven services.

Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Democratizing MoE inference on commodity GPUs with CoMoE

We present CoMoE, a communication-efficient MoE inference system that resolves this mismatch through novel host-centric routing.

Paper
TechCrunch: AI2d ago

TechCrunch Disrupt 2026 starts in 4 days — lock in your pass savings of up to $100 before prices rise

Four days until TechCrunch Disrupt 2026 starts, when 10,000 founders, investors, and tech leaders gather in San Francisco's Moscone West on October 13-15.

Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Expert Coupling in MoE Pretraining: Reducing All-to-All Overhead with Correlated Placement and Token Shuffling

Mixture-of-Experts (MoE) layers replace the feed-forward block of a Transformer with E expert networks, and each token is routed to k of these experts.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)3d ago

Cova-PINN: Cross-Domain Conservation Physics-Informed Neural Network for Fluid-Solid Conjugate Heat Transfer in Complex Geometries

Multi-domain physics-informed neural networks (PINNs) flexibly model medium-specific representations to solve fluid--solid conjugate heat transfer (CHT).

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

CPU-Auth: Device Fingerprinting for Authentication via DVFS Side-Channel

This work explores CPU-Auth, a novel authentication mechanism based on unique variations in the physical characteristics of the CPU of a computing device.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

PoreML: A Data-Driven Framework for Learning Multiphase Flow in Porous Media

Multiphase flow in porous microstructures is central to CO$2$ storage, fuel-cell operation, and flip-chip packaging.

Paper
Product Hunt: AI launches2d ago

iwant

Find available GPUs and host open models in one command

Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)4d ago

Embedded Evaluation of Task Admission Coalescing in Decentralized Multi-Robot Systems

Multi-robot task allocators in dynamic missions commonly admit newly released tasks immediately, potentially invoking allocation for each new arrival.

Paper
NVIDIA Technical Blog5d ago

AICR v1.0: Open, stable, and verifiable GPU cluster configuration

GPU-accelerated Kubernetes clusters depend on compatible versions across dozens of components, each on its own release cycle: host kernels, GPU drivers,...GPU-accelerated Kubernetes clusters depend on compatible versions across dozens of components, each on its own release cycle: host kernels, GPU drivers, container runtimes, networking, storage, operators, and workload frameworks.

NVIDIA Technical Blog5d ago

Control How Your GPU Shares Work with Green Contexts

Controlling how GPU resources are shared between them remains difficult.

Product Hunt: AI launches3d ago

Pine Computer

A cloud computer built for AI to get jobs done

Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

X-OPM: Explainable Automatic Digital On-Chip Power Modeling for Enhanced Robustness

Grounded in the design principles of synchronous digital VLSI circuits, X-OPM introduces a robust feature engineering framework that uses tree-based models to capture feature interactions and linear models for prediction.

Paper
Paper
arXiv (AI, ML, NLP, CV, robotics, multi-agent)5d ago

Accelerating the Development of PLGA In Situ Forming Depots Through AI-Driven Multi-Objective Optimization

Developing long-acting injectable formulations requires the simultaneous optimization of drug loading, release kinetics, viscosity, injectability, stability and other objectives.

Paper
Product Hunt: AI launches5d ago

HeyPi

Build and ship anything live, in-meetings