This person set up a context-aware router in front of compiled llama.cpp servers, running Qwen 3.8 27B through LM Studio and OpenCode with downloaded skills — and says it worked correctly almost immediately, with the remaining effort going into extra features and optimization rather than getting it functional at all. They expect it to eventually handle most of what they currently use Opus 4.8 for at work, just at a slower pace, and say it's already usable for that today.
VERIFIABLE SOURCE
View comment Reported anonymously by an r/LocalLLM contributor · score 2