A four-DGX-Spark cluster under vLLM serves several concurrent agentic coding and research sessions, with unquantized weights and honest notes on concurrency limits.
Capacity range
Exact reported VRAM
Practice
Model
Quantization
Runtime
Tools / workflow (harness)
Record origin
EXPLORE BY CAPACITY
Start with your machine.
Capacity is the reported total for this setup. Multi-GPU and unified-memory layouts can behave differently.
8 GB or lessLaptops, iGPUs and modest cards9–12 GBEntry-level local inference13–16 GBCurrent mid-range cards17–20 GBRX 7900 XT and compact pro cards21–24 GB3090, 4090 and 7900 XTX class25–32 GB5090 and workstation cards33–48 GBWorkstation and multi-GPU setups49–96 GBLarge multi-GPU and datacenter setups97 GB+Unified memory, servers and clusters
3 resultsEvery record keeps its origin visible.
Hybrid orchestrationGenesis archive2026-09
Four DGX Sparks running agentic work all dayHomelab opsGenesis archive2026-09
Migrating an agent between machines over text messageUsing Hermes to patch DeepSeek to a vision release and then move the whole agent, containers, and environment to another computer — remotely, from a phone.
CodingGenesis archive2026-09
98% local, on a self-extended Pi Agent 'software factory'DeepSeek V4 Flash Vision on two DGX Sparks, wrapped in a heavily extended Pi Agent setup, handles nearly all of this person's hands-off work.