Back to directory
Coding · 2026-10-03

Qwen3.8 Flash Next 177B at 11-15 tok/s on a single RTX 5070 12GB with 32GB DDR4-2400

A llama.cpp expert-streaming build runs the 177B Qwen 3.8 Flash Next MoE at roughly 11.5 tok/s on a 12GB RTX 5070 and 32GB DDR4-2400 by streaming experts from SSD.

A single RTX 5070 12GB with 32GB DDR4-2400, a Ryzen 5 5600GT and a PCIe Gen3 link runs Qwen 3.8 Flash Next 177B (UD-IQ3_XXS) through a llama.cpp expert-streaming setup on Windows. The author benchmarks about 11.5 tok/s, sees 14-15 tok/s in normal conversation, and generated 4,892 tokens at 10.15 tok/s on a coding prompt that produced a working single-file Snake game. Gains came from fixing Windows I/O queue depth, one file handle per worker, and a page-locked hot-expert tier. The author notes reloading a long chat from scratch is slow and the project targets machines with little RAM or VRAM.

VERIFIABLE SOURCE

Reported anonymously by an r/LocalLLM contributor · 1 upvotes at capture

View source

Useful references

Community-provided links related to this setup, workflow or measurements.

SETUP HISTORY

Snapshots over time

1 version
CURRENT2026-10-03Qwen3.8 Flash Next 177B at 11-15 tok/s on a single RTX 5070 12GB with 32GB DDR4-2400

A llama.cpp expert-streaming build runs the 177B Qwen 3.8 Flash Next MoE at roughly 11.5 tok/s on a 12GB RTX 5070 and 32GB DDR4-2400 by streaming experts from SSD.

Coding1 machineQwen 3.8 Flash Nextllama.cpp

This is the currently published snapshot.