This developer runs Qwen 3.8 Flash Next on a Strix Halo at 40-50 tok/s decode and about 1300 tok/s prefill, and uses it to maintain several large projects: a full digital audio workstation that supports large sample libraries without VSTs, an ad-blocking proxy with more features than uBlock Origin, and a card-focused note app with Obsidian-level capability. They say there is much more, and that the model running on a Pi harness performs at roughly the level of Opus 4.8 for their work (and they consider 4.6 better than 4.8). Their message is that what matters is how you adjust the setup to your minimum hardware, not the raw benchmark.
Reported anonymously by an r/LocalLLM contributor · score 1