This person runs research workflows on a 128GB Strix Halo, mainly using Qwen 3.8 Flash Next and Gemma 4 31B, and can also run DeepSeek V4 Flash in a decent quant. Both reach at least 20 tok/s, probably more with MTP.
For research, meaning web search plus analysis, they use Qwen 3.8 Flash Next for the main strategy, search planning and summarising, while cheaper Qwen 3.6 35B handles single-page summarisation and verification. Their recommendation is blunt: they have the 128GB version and would not recommend anything lower for this kind of work.
Reported anonymously by an r/LocalLLM contributor · score 2