Back to directory
Research & analysis · 2026-09-25

Qwen 3.8 Flash Next on 12GB of VRAM via RAM+VRAM offload (Strata engine)

Four quants of a vision model measured on a 12GB RTX 5070 SFF with 64GB DDR5, through a purpose-built inference engine, at 37 to 55 GB combined RAM+VRAM.

A custom inference engine (Strata) tuned for one model family on one class of hardware reports measured numbers for Qwen 3.8 Flash Next on a 12GB RTX 5070 SFF with a Ryzen 5 7600 and 64GB DDR5, Windows. At 128K context, output runs Q2_0 65.1 tok/s, IQ2_XS 52.0, IQ3_XXS 44.8 and IQ3_S 42; prompt processing runs Q2_0 543, IQ2_XS 472, IQ3_XXS 414 and IQ3_S 374. The model is not resident in VRAM alone: minimum RAM+VRAM is 37.6 GB (Q2_0) rising to 54.8 GB (IQ3_S), plus 0.91 GB for the vision encoder, so the 12GB card is one part of a 64GB system rather than the whole story.

VERIFIABLE SOURCE

Reported anonymously by an r/LocalLLM contributor · 0 upvotes at capture

View source

Useful references

Community-provided links related to this setup, workflow or measurements.

SETUP HISTORY

Snapshots over time

1 version
CURRENT2026-09-25Qwen 3.8 Flash Next on 12GB of VRAM via RAM+VRAM offload (Strata engine)

Four quants of a vision model measured on a 12GB RTX 5070 SFF with 64GB DDR5, through a purpose-built inference engine, at 37 to 55 GB combined RAM+VRAM.

Research & analysis1 machineQwen 3.8 Flash NextStrata

This is the currently published snapshot.