After testing Qwen 3.8 27B and Flash Next through Qwen's own API — their own hardware can't run them — this person concluded that for harder thinking and planning tasks they'd still lean on cheap cloud subscriptions rather than build a dedicated local rig, reasoning that the hardware cost alone could buy a decade of cheap cloud inference at current prices. What local they do run happens on their laptop's integrated Radeon 780M graphics, using system RAM rather than a dedicated card, at a Q5 quant, for lighter jobs: summarization, document review, OCR, and transcription, wired up with Open WebUI and an offline Wikipedia copy for privacy.
Reported anonymously by an r/LocalLLM contributor · score 2