Back to directory
Business automation · 2026-09-13

Two RTX 5090s running helpdesk and ops for a small office

An office replaced roughly $2,000 a month of API consumption with two RTX 5090 machines running Qwen 3.8 27B under vLLM, one for helpdesk and one for ops.

This office runs two Lenovo Legion machines, each with an RTX 5090, hosting Qwen 3.8 27B in NVFP4 under Ubuntu and vLLM. They started with a round-robin setup, then dedicated one machine to helpdesk and one to ops, routed through LiteLLM. The model is served from a sharded directory with an fp8 KV cache at a 262K context.

The economics are the driver: they report saving roughly $2,000 per month in Anthropic API consumption, with staff work and automations running at effectively zero marginal cost.

VERIFIABLE SOURCE

Reported anonymously by an r/LocalLLM contributor · score 1

View source
SETUP HISTORY

Snapshots over time

1 version
CURRENT2026-09-13Two RTX 5090s running helpdesk and ops for a small office

An office replaced roughly $2,000 a month of API consumption with two RTX 5090 machines running Qwen 3.8 27B under vLLM, one for helpdesk and one for ops.

Business automation2 machinesQwen 3.8 27BvLLM

This is the currently published snapshot.