This person set out to find the minimum hardware for an agent that genuinely manipulates a computer. Their conclusion is at least a Strix Halo with an oculink-linked R9700, because such an agent needs two things at once: a large model answering quickly as the brain and orchestrator (DeepSeek V4 Flash 0731 or Vision, Qwen Flash Next, or at minimum Qwen 27B), and a smaller capable model on fast memory to handle concurrent work like memory recall, web crawling and subagent tasks (Qwen 27B dense, Gemma 4 31B, at minimum Gemma 12B, with Qwen 35B an acceptable compromise).
Their own agent runs Qwen Flash Next in the Strix Halo's main memory with the smaller model on a GPU connected over oculink (a V100 32GB), using ROCm FP4 or a Halogen server. A DGX Spark running DeepSeek V4 Flash in EXL3 quant is a viable alternative, but still needs a second machine with at least a 16GB GPU for Gemma 4 12B, plus tinkering to give the agent speech transcription and image recognition. They are clear that this is expensive and only worth it for cases where the agent really touches your machine and passwords, or where you hold sensitive or proprietary data, or simply because you refuse to hand everything to OpenAI or Anthropic.
Reported anonymously by an r/LocalLLM contributor · score 4