A Radeon 7900XTX with 24GB VRAM and 64GB of DDR5 runs Qwen 3.8 Flash Next through Strata at about 60 tok/s, even at 250k context. The author uses it for frontend and backend coding, keeps a 6GB VRAM reserve so the GPU sits near 20GB and Windows 11 has about 4GB free, and games while the model runs in the background. They report q8 precision at longer context and Windows 11 performance matching Linux, and say it replaced their 27B dense model for accuracy and to avoid out-of-memory errors.
Reported anonymously by an r/LocalLLM contributor · 1 upvotes at capture