A single RTX 5070 12GB with 32GB DDR4-2400, a Ryzen 5 5600GT and a PCIe Gen3 link runs Qwen 3.8 Flash Next 177B (UD-IQ3_XXS) through a llama.cpp expert-streaming setup on Windows. The author benchmarks about 11.5 tok/s, sees 14-15 tok/s in normal conversation, and generated 4,892 tokens at 10.15 tok/s on a coding prompt that produced a working single-file Snake game. Gains came from fixing Windows I/O queue depth, one file handle per worker, and a page-locked hot-expert tier. The author notes reloading a long chat from scratch is slow and the project targets machines with little RAM or VRAM.
Reported anonymously by an r/LocalLLM contributor · 1 upvotes at capture
Useful references
Community-provided links related to this setup, workflow or measurements.