This person runs a custom Linux workstation with an RTX Pro 6000 (96GB), an i9 and 48GB of RAM, using vLLM and vLLM-Moet with a DeepSeek harness and Hermes. They use it for programming, copy editing and summaries.
Reported performance is about 2000-4000 tok/s prefill and 50-90 tok/s output, with context from 256K up to 1,000,000. They have not used the cloud all year and have never paid for online inference, saying the setup is now better than what free tiers offered nine months ago. The only honest limitation they give is that the model sometimes falls into a loop, and they consider themselves the main bottleneck otherwise.
Reported anonymously by an r/LocalLLM contributor · score 1