After decades of software engineering experience, this user runs Qwen models on a 128GB Strix Halo for deep work. Generation stays below 30 tokens per second, but the author considers the experience excellent because they review every line of code. Qwen 3.8 is preferred for more demanding work, while Qwen 3.6 35B-A3B is used when faster, less demanding assistance is useful. The machine's unified-memory configuration was reported, but no separate VRAM allocation was given.
VERIFIABLE SOURCE
View comment Reported anonymously by an r/LocalLLM contributor · score 2