A custom inference engine (Strata) tuned for one model family on one class of hardware reports measured numbers for Qwen 3.8 Flash Next on a 12GB RTX 5070 SFF with a Ryzen 5 7600 and 64GB DDR5, Windows. At 128K context, output runs Q2_0 65.1 tok/s, IQ2_XS 52.0, IQ3_XXS 44.8 and IQ3_S 42; prompt processing runs Q2_0 543, IQ2_XS 472, IQ3_XXS 414 and IQ3_S 374. The model is not resident in VRAM alone: minimum RAM+VRAM is 37.6 GB (Q2_0) rising to 54.8 GB (IQ3_S), plus 0.91 GB for the vision encoder, so the 12GB card is one part of a 64GB system rather than the whole story.
Reported anonymously by an r/LocalLLM contributor · 0 upvotes at capture
Useful references
Community-provided links related to this setup, workflow or measurements.