This office runs two Lenovo Legion machines, each with an RTX 5090, hosting Qwen 3.8 27B in NVFP4 under Ubuntu and vLLM. They started with a round-robin setup, then dedicated one machine to helpdesk and one to ops, routed through LiteLLM. The model is served from a sharded directory with an fp8 KV cache at a 262K context.
The economics are the driver: they report saving roughly $2,000 per month in Anthropic API consumption, with staff work and automations running at effectively zero marginal cost.
Reported anonymously by an r/LocalLLM contributor · score 1