This person is candid that, for them, local is just for playing around and learning about AI — real work goes through Claude, which they consider their actual workhorse. They run about 90% of their local sessions through Hermes tied to a $20 ChatGPT plan, and have Claude configured to offload what it can to the local GPU and then validate and correct that output, rather than trusting the local model's work directly. Limited by a 16GB 5070 Ti today, they're weighing a 5090 or a DGX purchase, but expect dense ~27B models to stay the sweet spot for consumer hardware for years.
Reported anonymously by an r/LocalLLM contributor · score 2