Back to directory
Coding · 2026-10-03

Qwen 3.8 Flash Next on a 2021 M1 Max: 44 tok/s and 35 tok/s at 398K context with Splash

Long-context decode benchmark of Qwen 3.8 Flash Next on an M1 Max 64 GB through an unofficial Splash/ds4 fork in real OpenCode agent sessions.

Ran Qwen 3.8 Flash Next on a 2021 M1 Max laptop (32 GPU cores, 64 GB unified memory) through a Splash/ds4 fork for real OpenCode agent sessions. With the GSQ-RCO Q2_0 quant (35 GiB) and MTP, decode held 44 tok/s at 4K, 38.8 at 256K and 35.4 at 398K, with prefill falling from 328 to 292 tok/s, while IQ3_XXS (44 GiB) managed 35.2 tok/s at 4K and 31.7 at 259K. Sustained load stayed flat at 43.2/43.1 tok/s over five minutes at 73-75 C with no throttling and 0.70 J per token. Limits: the fork is unofficial, Q2_0 loses about 6 points on code versus BF16 in ISTA's evals (81.1 vs 87.4 LiveCodeBench v6, 89.1 vs 93.1 overall), and Flash Next still needs a 64 GB Mac.

VERIFIABLE SOURCE

Reported anonymously by an r/LocalLLM contributor · 1 upvotes at capture

View source

Useful references

Community-provided links related to this setup, workflow or measurements.

SETUP HISTORY

Snapshots over time

1 version
CURRENT2026-10-03Qwen 3.8 Flash Next on a 2021 M1 Max: 44 tok/s and 35 tok/s at 398K context with Splash

Long-context decode benchmark of Qwen 3.8 Flash Next on an M1 Max 64 GB through an unofficial Splash/ds4 fork in real OpenCode agent sessions.

Coding1 machineQwen 3.8 Flash Next, Qwen 3.8 27BSplash

This is the currently published snapshot.