Ran Qwen 3.8 Flash Next on a 2021 M1 Max laptop (32 GPU cores, 64 GB unified memory) through a Splash/ds4 fork for real OpenCode agent sessions. With the GSQ-RCO Q2_0 quant (35 GiB) and MTP, decode held 44 tok/s at 4K, 38.8 at 256K and 35.4 at 398K, with prefill falling from 328 to 292 tok/s, while IQ3_XXS (44 GiB) managed 35.2 tok/s at 4K and 31.7 at 259K. Sustained load stayed flat at 43.2/43.1 tok/s over five minutes at 73-75 C with no throttling and 0.70 J per token. Limits: the fork is unofficial, Q2_0 loses about 6 points on code versus BF16 in ISTA's evals (81.1 vs 87.4 LiveCodeBench v6, 89.1 vs 93.1 overall), and Flash Next still needs a 64 GB Mac.
Reported anonymously by an r/LocalLLM contributor · 1 upvotes at capture
Useful references
Community-provided links related to this setup, workflow or measurements.