Back to directory
Documents · 2026-09-11

Kept paying for cheap cloud coding subs; local stays on an iGPU for light tasks

Weighed building a dedicated inference box against just paying for cheap cloud models — for now, cloud wins for anything demanding, while a laptop's integrated graphics handles the light stuff.

After testing Qwen 3.8 27B and Flash Next through Qwen's own API — their own hardware can't run them — this person concluded that for harder thinking and planning tasks they'd still lean on cheap cloud subscriptions rather than build a dedicated local rig, reasoning that the hardware cost alone could buy a decade of cheap cloud inference at current prices. What local they do run happens on their laptop's integrated Radeon 780M graphics, using system RAM rather than a dedicated card, at a Q5 quant, for lighter jobs: summarization, document review, OCR, and transcription, wired up with Open WebUI and an offline Wikipedia copy for privacy.

VERIFIABLE SOURCE

Reported anonymously by an r/LocalLLM contributor · score 2

View comment