This contributor uses Qwen 3.8 Flash Next Q3 XXS for both hobby and professional projects on a laptop with an RTX 3500 Ada, 12GB of VRAM and 64GB of RAM. The runtime is llama.cpp and the surrounding workflow uses a DeepSeek harness. To achieve usable speed, the setup loads part of the MoE into CUDA-addressable system memory and keeps PLE data on an SSD, using lazy loading and a reduced load mode. The reported performance is around 10–15 tokens per second for generation and 80–120 tokens per second for prefill.
VERIFIABLE SOURCE
View submission form Submitted anonymously via Tally · reference zEOMJ20