With 49–96 GB,
here is what really runs.
Experiences reported by people using their machines every day. The figure is the reported setup capacity; multi-GPU and unified-memory layouts can behave differently.
in this range
Coaching elite athletes with a 2x DGX Spark cluster
A cycling and running coach runs client-facing coaching assistance on a small DGX Spark cluster, good enough for national-champion-level clients — at a cost the coach openly questions.
Two DGX Sparks as a daily driver, ROI aside
This person knows Qwen 3.8 Flash Next on two DGX Sparks can't match frontier intelligence, and says that's not really the point for them.
Three rigs for freelance work an NDA won't let leave the building
Client confidentiality rules out cloud APIs for this freelancer, who spreads different model sizes across three separate machines.
Coding a SaaS product on an M1 Max, one well-scoped feature at a time
Running Qwen 3.8 27B on the same Mac they work on, this developer credits per-feature domain glossaries and tight scoping — not the model — for making local coding reliable.
An RTX 6000 Pro and an M4 Max, benchmarked against an active Claude Max plan
This person kept paying for Claude Max specifically so they'd have a real comparison point for their own coding hardware — and admits the math still doesn't favor local.
Eight daily agent workflows, driven from Telegram on a Strix Halo
A full personal-operations stack runs on one 128GB Strix Halo machine, with several named agents reachable through a single Telegram chat and each handling a different part of daily life.
Building an agent stack locally before moving inference to the cloud
A product team developed their agentic tooling against local inference on a Strix Halo and midrange GPUs, then retargeted the same harness to cloud inference once it worked.
A senior engineer with no need for cloud on a 64GB Strix Halo
An experienced software engineer runs Qwen3-Coder 30B on a 64GB Strix Halo and finds local working better than newer, larger models for well-scoped tasks.
98% local, on a self-extended Pi Agent 'software factory'
DeepSeek V4 Flash Vision on two DGX Sparks, wrapped in a heavily extended Pi Agent setup, handles nearly all of this person's hands-off work.
Using a local model to stress-test the agent a coding agent is building
An unusual layered setup: the local model isn't the coding agent — it's the thing an outer coding agent uses to build and stress-test a separate 'inner' agent.
Academic and policy research on an M5 and a DGX Spark
Purpose-built tools around a personal research corpus, plus Hermes as an interactive research assistant, get this researcher roughly 80% of the way to frontier quality — locally.
A workstation GPU as a drop-in subagent for Codex and Claude
Qwen 3.8 Flash Next on an RTX Pro 6000 takes implementation and review tasks delegated from Codex or Claude, replacing Sonnet and Opus for that slice.
Local only for what is deterministic: embeddings, speech and TTS
After testing every model that fits a 128GB Strix Halo, this user keeps local strictly to embeddings, text-to-speech and speech-to-text, and relies on cloud for reasoning.
The minimum setup for a genuinely useful local agent
After experimenting, this user argues a real computer-manipulating agent needs at least a Strix Halo plus an oculink R9700, or a DGX Spark plus a second GPU for auxiliary models.
Maintaining a DAW, an adblock proxy and a notes app on a Strix Halo
A developer runs Qwen 3.8 Flash Next on a Strix Halo through a Pi harness and reports 40-50 tok/s decode plus 1300 tok/s prefill while maintaining three substantial projects.
Building a game iteratively with a local Pi harness
A non-professional developer uses a local model in small, iterative steps to build a long-term personal game project.
Coding and homelab maintenance on a dual-R9700 server
A garage server with two Radeon Pro AI R9700 cards supports coding, homelab maintenance and an evolving Cline/Open WebUI workflow.
Privacy-constrained freelance work on a four-3090 workstation
A freelance worker keeps restricted client work local on a large multi-machine setup that also doubles as a Blender and data-processing workstation.
A cloud-free workstation: RTX Pro 6000, vLLM and a DeepSeek harness
A custom Linux workstation with an RTX Pro 6000 and 48GB of RAM runs programming, copy editing and summarisation through vLLM, with no cloud use all year.
A 128GB Strix Halo research stack, with a warning against smaller machines
A researcher runs Qwen 3.8 Flash Next and Gemma 4 31B on a 128GB Strix Halo, splitting strategy and summarisation across models and advising against 64GB machines.
Agentic coding on a 512GB Mac Studio with oMLX
A 512GB M3 Ultra Mac Studio running models through oMLX handles most agentic coding work, with measured speeds across several models listed by the owner.
Migrating a business's own cloud onto a Strix Halo with Hermes
A Hermes agent on a 128GB Strix Halo handles coding, open-source PRs, a docker-swarm-to-k8s migration and invoice processing on ERPNext.