At work, this person runs an end-of-day job that transcribes all customer calls and summarizes them into a database, entirely locally, and says it works great — audio transcription through cloud providers is, for reasons they find odd, more expensive than the equivalent in text tokens, so local saves real money there. For coding, though, they're unambiguous and explicitly reject local: Qwen 27B isn't close to what people claim, and even scraping together enough VRAM to run Qwen 3.8 Flash Next doesn't close the gap — it's still noticeably behind and runs much slower besides, so that part of their work stays on cloud models.
Reported anonymously by an r/LocalLLM contributor · score 2