Back to directory
Coding · 2026-09-11

The first local model this person actually trusts with real work

On a 12GB 5070, a heavily quantized Qwen 3.8 27B through a custom Hermes backend is, in their words, the first local model that's felt confidently productive.

Running Qwen 3.8 27B at a UD-Q2 quant through Hermes on a custom backend, on just a 12GB RTX 5070, this person calls it the first local model they've felt they could confidently use for productive work. It can take a while to work through a problem, but they describe it as thorough, genuinely able to code and debug rather than just producing plausible-looking output.

VERIFIABLE SOURCE

Reported anonymously by an r/LocalLLM contributor · score 2

View comment