Back to directory
Coding · 2026-10-03

Coding locally with Qwen 3.8 27B Q6 on an RTX 5090: 160k context, ~80 tok/s, Hermes harness

The author runs Qwen 3.8 27B at Q6 with a 160k context and vision on an RTX 5090, using the Hermes agent harness and SearxNG in Docker, and reports around 80 tok/s but about 15 minutes for a one-shot website.

The author runs Qwen 3.8 27B at Q6 with a 160k context window, medium thinking and vision enabled on an RTX 5090, driven by the Hermes agent harness, and reports around 80 tok/s on real-world coding tasks. SearxNG runs in Docker for web browsing. A one-shot website took about 15 minutes, which the author finds slow compared with an estimated 3 to 5 minutes for Claude, and they question whether Qwen 3.8 27B overthinks for coding.

VERIFIABLE SOURCE

Reported anonymously by an r/LocalLLM contributor · 1 upvotes at capture

View source
SETUP HISTORY

Snapshots over time

1 version
CURRENT2026-10-03Coding locally with Qwen 3.8 27B Q6 on an RTX 5090: 160k context, ~80 tok/s, Hermes harness

The author runs Qwen 3.8 27B at Q6 with a 160k context and vision on an RTX 5090, using the Hermes agent harness and SearxNG in Docker, and reports around 80 tok/s but about 15 minutes for a one-shot website.

Coding1 machineQwen 3.8 27BRuntime unspecified

This is the currently published snapshot.