A four-DGX-Spark cluster under vLLM serves several concurrent agentic coding and research sessions, with unquantized weights and honest notes on concurrency limits.
Start with your machine.
Capacity is the reported total for this setup. Multi-GPU and unified-memory layouts can behave differently.
A sales professional dropped every frontier subscription and runs work agents on two DGX Sparks, averaging about 200 million tokens per week.
After six months of tinkering, Qwen 3.8 27B justified a 7900XTX plus two DGX Sparks, taking over a job the estimating team was spending dozens of hours a week on.
Beyond personal daily use, this person stood up a Spark cluster for a client so their employees get an internal, compliance-friendly LLM platform.
Referral triage that extracts details, standardizes them, and suggests categories and tests — with the final call always left to a doctor, and nothing ever leaving the network.