Running Next Flash IQ3_XXS at ~70 tok/s with 100k context or 2 instances of Qwen 3.6 35B A3B IQ4 at ~145 tok/s with 256k all on $500 of ex mining BC-250 boards
This will be my third update on the bc-250 cluster and for my first forray into local ai I have been having a blast.
Key points
- I am running 4 BC-250 boards connected by 2.5gb ethernet adapters to a 2.5gb switch using llama with vulkan and RPC.
- After starting out at 30 tok/s on flash next Iq2 and 80 tok/s with 3.6 35B I pointed Claude Code at the cluster and we are now running a better quant next flash as well as maintaining 60-70 tok/s (~150 ppt) at context.
- The 4 boards can run 2 instances of 3.6 35B at 145tok/s (~450 ppt) starting and going below 100 tok/s at 150k context.
- The next step for me is to use AIO coolers for each board as now I have been seeing some thermal throttling as the boards are better utilized .
Sources (1)
- [1]Running Next Flash IQ3_XXS at ~70 tok/s with 100k context or 2 instances of Qwen 3.6 35B A3B IQ4 at ~145 tok/s with 256k all on $500 of ex mining BC-250 boardsr/LocalLLaMA (top, daily) · Oct 11, 04:27 AM
This will be my third update on the bc-250 cluster and for my first forray into local ai I have been having a blast.
I am running 4 BC-250 boards connected by 2.5gb ethernet adapters to a 2.5gb switch using llama with vulkan and RPC.
Extractive summary: sentences quoted from the sources.