Back to directory
Coding · 2026-09-14

A cloud-free workstation: RTX Pro 6000, vLLM and a DeepSeek harness

A custom Linux workstation with an RTX Pro 6000 and 48GB of RAM runs programming, copy editing and summarisation through vLLM, with no cloud use all year.

This person runs a custom Linux workstation with an RTX Pro 6000 (96GB), an i9 and 48GB of RAM, using vLLM and vLLM-Moet with a DeepSeek harness and Hermes. They use it for programming, copy editing and summaries.

Reported performance is about 2000-4000 tok/s prefill and 50-90 tok/s output, with context from 256K up to 1,000,000. They have not used the cloud all year and have never paid for online inference, saying the setup is now better than what free tiers offered nine months ago. The only honest limitation they give is that the model sometimes falls into a loop, and they consider themselves the main bottleneck otherwise.

VERIFIABLE SOURCE

Reported anonymously by an r/LocalLLM contributor · score 1

View source
SETUP HISTORY

Snapshots over time

1 version
CURRENT2026-09-14A cloud-free workstation: RTX Pro 6000, vLLM and a DeepSeek harness

A custom Linux workstation with an RTX Pro 6000 and 48GB of RAM runs programming, copy editing and summarisation through vLLM, with no cloud use all year.

Coding1 machineQwen 3.8 Flash Next, DeepSeek V4 FlashvLLM

This is the currently published snapshot.