Back to directory
Research & analysis · 2026-09-13

A 128GB Strix Halo research stack, with a warning against smaller machines

A researcher runs Qwen 3.8 Flash Next and Gemma 4 31B on a 128GB Strix Halo, splitting strategy and summarisation across models and advising against 64GB machines.

This person runs research workflows on a 128GB Strix Halo, mainly using Qwen 3.8 Flash Next and Gemma 4 31B, and can also run DeepSeek V4 Flash in a decent quant. Both reach at least 20 tok/s, probably more with MTP.

For research, meaning web search plus analysis, they use Qwen 3.8 Flash Next for the main strategy, search planning and summarising, while cheaper Qwen 3.6 35B handles single-page summarisation and verification. Their recommendation is blunt: they have the 128GB version and would not recommend anything lower for this kind of work.

VERIFIABLE SOURCE

Reported anonymously by an r/LocalLLM contributor · score 2

View source
SETUP HISTORY

Snapshots over time

1 version
CURRENT2026-09-13A 128GB Strix Halo research stack, with a warning against smaller machines

A researcher runs Qwen 3.8 Flash Next and Gemma 4 31B on a 128GB Strix Halo, splitting strategy and summarisation across models and advising against 64GB machines.

Research & analysis1 machineQwen 3.8 Flash Next, Gemma 4 31B, DeepSeek V4 Flash, Qwen 3.6 35B-A3BRuntime unspecified

This is the currently published snapshot.