A product team developed their agentic tooling against local inference on a Strix Halo and midrange GPUs, then retargeted the same harness to cloud inference once it worked.
Start with your machine.
Capacity is the reported total for this setup. Multi-GPU and unified-memory layouts can behave differently.
After testing every model that fits a 128GB Strix Halo, this user keeps local strictly to embeddings, text-to-speech and speech-to-text, and relies on cloud for reasoning.
A developer runs Qwen 3.8 Flash Next on a Strix Halo through a Pi harness and reports 40-50 tok/s decode plus 1300 tok/s prefill while maintaining three substantial projects.
An experienced software engineer finds sub-30 tok/s local inference on a 128GB Strix Halo more than adequate for deep coding work.
A Hermes agent on a 128GB Strix Halo handles coding, open-source PRs, a docker-swarm-to-k8s migration and invoice processing on ERPNext.
A full personal-operations stack runs on one 128GB Strix Halo machine, with several named agents reachable through a single Telegram chat and each handling a different part of daily life.
A security professional runs uncensored Qwen models for client work on a 64GB MacBook and a 128GB Strix Halo, using a containerised OpenCode setup with ported code-review skills.
An experienced software engineer runs Qwen3-Coder 30B on a 64GB Strix Halo and finds local working better than newer, larger models for well-scoped tasks.
After experimenting, this user argues a real computer-manipulating agent needs at least a Strix Halo plus an oculink R9700, or a DGX Spark plus a second GPU for auxiliary models.
A researcher runs Qwen 3.8 Flash Next and Gemma 4 31B on a 128GB Strix Halo, splitting strategy and summarisation across models and advising against 64GB machines.
DevOps work, smart-home control, and paperwork all moved off the cloud once Strix Halo hardware arrived — with one exception the poster caught themselves forgetting.