A four-DGX-Spark cluster under vLLM serves several concurrent agentic coding and research sessions, with unquantized weights and honest notes on concurrency limits.
Start with your machine.
Capacity is the reported total for this setup. Multi-GPU and unified-memory layouts can behave differently.
With access to both Claude Fable and high-end local models, this developer chooses local because a less autonomous model keeps him involved in debugging and research work.
A freelance worker keeps restricted client work local on a large multi-machine setup that also doubles as a Blender and data-processing workstation.
A custom Linux workstation with an RTX Pro 6000 and 48GB of RAM runs programming, copy editing and summarisation through vLLM, with no cloud use all year.
After experimenting, this user argues a real computer-manipulating agent needs at least a Strix Halo plus an oculink R9700, or a DGX Spark plus a second GPU for auxiliary models.
A researcher runs Qwen 3.8 Flash Next and Gemma 4 31B on a 128GB Strix Halo, splitting strategy and summarisation across models and advising against 64GB machines.
A 512GB M3 Ultra Mac Studio running models through oMLX handles most agentic coding work, with measured speeds across several models listed by the owner.
Beyond the obvious coding use case, this setup handles website debugging, deep web research, and day-to-day Linux server maintenance.
Client confidentiality rules out cloud APIs for this freelancer, who spreads different model sizes across three separate machines.
Purpose-built tools around a personal research corpus, plus Hermes as an interactive research assistant, get this researcher roughly 80% of the way to frontier quality — locally.