Back to directory
Hybrid orchestration · 2026-09-14

Four DGX Sparks running agentic work all day

A four-DGX-Spark cluster under vLLM serves several concurrent agentic coding and research sessions, with unquantized weights and honest notes on concurrency limits.

This setup uses four Nvidia DGX Sparks running vLLM with a Deepseek harness and models including DeepSeek V4 Flash, Qwen 3.8 Flash Next and DeepSeek V4 Flash Vision. The owner describes the use case as 'literally everything I can think of'.

Reported single-stream speeds are roughly 55-65 tok/s for DeepSeek 4.1 Flash, 70-90 tok/s for V4 Flash, about 45 tok/s for Qwen and about 35 tok/s for GLM. They can run roughly four concurrent agentic sessions before throughput drops below about 20 tok/s and feels too slow. For context, they note their workflow is faster than they can read and write, which makes it genuinely usable despite cloud APIs being far faster. No quantization is used, so output quality is expected to match the API weights.

Honest limitations: concurrency is poor, it was a large financial investment, it takes real effort to maintain, and getting a new model running can consume nearly a full day.

VERIFIABLE SOURCE

Reported anonymously by an r/LocalLLM contributor · score 1

View source
SETUP HISTORY

Snapshots over time

1 version
CURRENT2026-09-14Four DGX Sparks running agentic work all day

A four-DGX-Spark cluster under vLLM serves several concurrent agentic coding and research sessions, with unquantized weights and honest notes on concurrency limits.

Hybrid orchestration4 machinesDeepSeek V4 Flash, Qwen 3.8 Flash Next, DeepSeek V4 Flash VisionvLLM

This is the currently published snapshot.