The client is a 50-person company across departments including HR, marketing and finance, and the rig is used for local RAG over internal documents, lightweight LLM chat, localized diffusion runs for visual assets and agentic coding, with bursty peaks of 3 to 5 concurrent requests rather than 50 simultaneous streams. The machine holds 4x RTX 5090 (128GB total) on an ASUS W790E-SAGE SE with an Intel Xeon w5-3423, 192GB of ECC DDR5 and 2TB NVMe, each card on its own PCIe 5.0 x16 riser and cooled by 18 server fans. The lab runs DeepSeek V4 Flash, Qwen 3.8 and Gemma 4 32B. The published throughput numbers (24 concurrent at 52 tok/s and about 2.4s to first token in short context, 15.3 tok/s at 25k to 50k context, and 4.3 tok/s at 100k for 8 concurrent) were measured on their 2x RTX 5090 box with Qwen 3.6 27B, not on this 4-card machine, which was not benchmarked after its models were updated.
Reported anonymously by an r/LocalLLM contributor · 315 upvotes at capture
Useful references
Community-provided links related to this setup, workflow or measurements.