Back to directory
Coding · 2026-09-14

Deep coding work on a 128GB Strix Halo

An experienced software engineer finds sub-30 tok/s local inference on a 128GB Strix Halo more than adequate for deep coding work.

After decades of software engineering experience, this user runs Qwen models on a 128GB Strix Halo for deep work. Generation stays below 30 tokens per second, but the author considers the experience excellent because they review every line of code. Qwen 3.8 is preferred for more demanding work, while Qwen 3.6 35B-A3B is used when faster, less demanding assistance is useful. The machine's unified-memory configuration was reported, but no separate VRAM allocation was given.

VERIFIABLE SOURCE

Reported anonymously by an r/LocalLLM contributor · score 2

View comment