A sales professional dropped every frontier subscription and runs work agents on two DGX Sparks, averaging about 200 million tokens per week.
Capacity range
Exact reported VRAM
Practice
Model
Quantization
Runtime
Tools / workflow (harness)
Record origin
EXPLORE BY CAPACITY
Start with your machine.
Capacity is the reported total for this setup. Multi-GPU and unified-memory layouts can behave differently.
8 GB or lessLaptops, iGPUs and modest cards9–12 GBEntry-level local inference13–16 GBCurrent mid-range cards17–20 GBRX 7900 XT and compact pro cards21–24 GB3090, 4090 and 7900 XTX class25–32 GB5090 and workstation cards33–48 GBWorkstation and multi-GPU setups49–96 GBLarge multi-GPU and datacenter setups97 GB+Unified memory, servers and clusters
3 resultsEvery record keeps its origin visible.
Business automationGenesis archive2026-09
Running sales work on GLM 5.3 Flash across two DGX Sparks, 200M tokens a weekPersonal assistantGenesis archive2026-09
A family finance tracker built with Hermes and an MCPA weekend vibe-coded web app plus a Hermes agent tracks family finances, with an MCP for adding expenses and inputs arriving from forwarded emails, Telegram or Discord.
CodingGenesis archive2026-09
Agentic coding on a 512GB Mac Studio with oMLXA 512GB M3 Ultra Mac Studio running models through oMLX handles most agentic coding work, with measured speeds across several models listed by the owner.