This person uses GLM 5.3 Flash with Hermes for work in sales, running on two DGX Sparks, and reports decent token speed. They average about 200 million tokens per week just on work tasks and have cancelled all frontier subscriptions, saying it has worked flawlessly.
Their opinion is that if you want the best results after walking away from closed frontier models, you need at least 256GB of VRAM, which is roughly where they feel quality starts to improve.
Reported anonymously by an r/LocalLLM contributor · score 1