This developer uses Qwen 3.8 27B as their primary coding model at work on a setup combining an RTX 3090 and RTX 5060. They describe the model as solid and trustworthy, better than many cloud models, but slower on their hardware. Frontier cloud inference remains a deliberate fallback for tasks requiring deeper expertise or a quick turnaround, making the division of labor explicit rather than treating local and cloud as competing all-or-nothing choices.
VERIFIABLE SOURCE
View comment Reported anonymously by an r/LocalLLM contributor · score 1