The author runs Qwen 3.8 27B at Q6 with a 160k context window, medium thinking and vision enabled on an RTX 5090, driven by the Hermes agent harness, and reports around 80 tok/s on real-world coding tasks. SearxNG runs in Docker for web browsing. A one-shot website took about 15 minutes, which the author finds slow compared with an estimated 3 to 5 minutes for Claude, and they question whether Qwen 3.8 27B overthinks for coding.
Reported anonymously by an r/LocalLLM contributor · 1 upvotes at capture