This person runs Qwen Flash Next on a home server at about 15 tok/s, and says that despite the modest speed it outperforms all their cloud models on agent tasks. They acknowledge having an unusually lucky setup, with 600GB of RAM and some custom servers gifted to them, but they bought an old P100 GPU and argue that going fully local is doable for under $3,000 these days.
They point to Qwen 3.8 27B in an IQ3 quant fitting in 16GB of VRAM as the budget option, or running something as large as Qwen Flash Next if there is enough memory headroom.
Reported anonymously by an r/LocalLLM contributor · score 1