Full directory
GUIDE BY CAPACITY

With 97 GB+,
here is what really runs.

Experiences reported by people using their machines every day. The figure is the reported setup capacity; multi-GPU and unified-memory layouts can behave differently.

18documented uses
in this range
Business automation

Coaching elite athletes with a 2x DGX Spark cluster

A cycling and running coach runs client-facing coaching assistance on a small DGX Spark cluster, good enough for national-champion-level clients — at a cost the coach openly questions.

Personal assistant

Two DGX Sparks as a daily driver, ROI aside

This person knows Qwen 3.8 Flash Next on two DGX Sparks can't match frontier intelligence, and says that's not really the point for them.

Personal assistant

Eight daily agent workflows, driven from Telegram on a Strix Halo

A full personal-operations stack runs on one 128GB Strix Halo machine, with several named agents reachable through a single Telegram chat and each handling a different part of daily life.

Hybrid orchestration

Four DGX Sparks running agentic work all day

A four-DGX-Spark cluster under vLLM serves several concurrent agentic coding and research sessions, with unquantized weights and honest notes on concurrency limits.

Business automation

Building an agent stack locally before moving inference to the cloud

A product team developed their agentic tooling against local inference on a Strix Halo and midrange GPUs, then retargeted the same harness to cloud inference once it worked.

Coding

98% local, on a self-extended Pi Agent 'software factory'

DeepSeek V4 Flash Vision on two DGX Sparks, wrapped in a heavily extended Pi Agent setup, handles nearly all of this person's hands-off work.

Coding

Eight A100s in tensor-parallel, inside a company with a hundred more

A large enterprise deployment: eight A100 80GBs running Qwen 3.8 under SGLang with tensor parallelism, while the company brings newer B200s online for bigger models.

Regulated work

A HIPAA-style internal LLM platform, built on a cluster of DGX Sparks

Beyond personal daily use, this person stood up a Spark cluster for a client so their employees get an internal, compliance-friendly LLM platform.

Regulated work

A cardiologist's referral system, fully air-gapped

Referral triage that extracts details, standardizes them, and suggests categories and tests — with the final call always left to a doctor, and nothing ever leaving the network.

Research & analysis

Academic and policy research on an M5 and a DGX Spark

Purpose-built tools around a personal research corpus, plus Hermes as an interactive research assistant, get this researcher roughly 80% of the way to frontier quality — locally.

Research & analysis

Local only for what is deterministic: embeddings, speech and TTS

After testing every model that fits a 128GB Strix Halo, this user keeps local strictly to embeddings, text-to-speech and speech-to-text, and relies on cloud for reasoning.

Hybrid orchestration

The minimum setup for a genuinely useful local agent

After experimenting, this user argues a real computer-manipulating agent needs at least a Strix Halo plus an oculink R9700, or a DGX Spark plus a second GPU for auxiliary models.

Coding

Maintaining a DAW, an adblock proxy and a notes app on a Strix Halo

A developer runs Qwen 3.8 Flash Next on a Strix Halo through a Pi harness and reports 40-50 tok/s decode plus 1300 tok/s prefill while maintaining three substantial projects.

Hybrid orchestration

Building a two-server local AI ecosystem across 17 GPUs

A self-hosted AI ecosystem spans two Linux servers, assigning different GPUs to long-context coding, inference, speech, RAG and media workloads.

Business automation

Running sales work on GLM 5.3 Flash across two DGX Sparks, 200M tokens a week

A sales professional dropped every frontier subscription and runs work agents on two DGX Sparks, averaging about 200 million tokens per week.

Research & analysis

A 128GB Strix Halo research stack, with a warning against smaller machines

A researcher runs Qwen 3.8 Flash Next and Gemma 4 31B on a 128GB Strix Halo, splitting strategy and summarisation across models and advising against 64GB machines.

Coding

Agentic coding on a 512GB Mac Studio with oMLX

A 512GB M3 Ultra Mac Studio running models through oMLX handles most agentic coding work, with measured speeds across several models listed by the owner.

Business automation

Migrating a business's own cloud onto a Strix Halo with Hermes

A Hermes agent on a 128GB Strix Halo handles coding, open-source PRs, a docker-swarm-to-k8s migration and invoice processing on ERPNext.