This person routes coding to a local Qwen 3.8 model served on nInfer, which they describe as very fast, while using a frontier model to orchestrate and verify the work. The frontier model also verifies the result afterwards.
The combination saves a lot of tokens, and because Claude does most of the planning they typically run the local model with thinking disabled.
VERIFIABLE SOURCE
View source Reported anonymously by an r/LocalLLM contributor · score 3