AION
Tutorial / explainerEfficiency & Inference · Hardware & Compute · Large Language Models1 source · Oct 11, 2026

Running Next Flash IQ3_XXS at ~70 tok/s with 100k context or 2 instances of Qwen 3.6 35B A3B IQ4 at ~145 tok/s with 256k all on $500 of ex mining BC-250 boards

This will be my third update on the bc-250 cluster and for my first forray into local ai I have been having a blast.

Key points

  • I am running 4 BC-250 boards connected by 2.5gb ethernet adapters to a 2.5gb switch using llama with vulkan and RPC.
  • After starting out at 30 tok/s on flash next Iq2 and 80 tok/s with 3.6 35B I pointed Claude Code at the cluster and we are now running a better quant next flash as well as maintaining 60-70 tok/s (~150 ppt) at context.
  • The 4 boards can run 2 instances of 3.6 35B at 145tok/s (~450 ppt) starting and going below 100 tok/s at 150k context.
  • The next step for me is to use AIO coolers for each board as now I have been seeing some thermal throttling as the boards are better utilized .

Sources (1)

Extractive summary: sentences quoted from the sources.