This setup uses four Nvidia DGX Sparks running vLLM with a Deepseek harness and models including DeepSeek V4 Flash, Qwen 3.8 Flash Next and DeepSeek V4 Flash Vision. The owner describes the use case as 'literally everything I can think of'.
Reported single-stream speeds are roughly 55-65 tok/s for DeepSeek 4.1 Flash, 70-90 tok/s for V4 Flash, about 45 tok/s for Qwen and about 35 tok/s for GLM. They can run roughly four concurrent agentic sessions before throughput drops below about 20 tok/s and feels too slow. For context, they note their workflow is faster than they can read and write, which makes it genuinely usable despite cloud APIs being far faster. No quantization is used, so output quality is expected to match the API weights.
Honest limitations: concurrency is poor, it was a large financial investment, it takes real effort to maintain, and getting a new model running can consume nearly a full day.
Reported anonymously by an r/LocalLLM contributor · score 1