Back to directory
Coding · 2026-09-11

A context-aware router working on the first try

A router in front of compiled llama.cpp servers, feeding Qwen 3.8 27B through OpenCode with downloaded skills, worked essentially out of the box.

This person set up a context-aware router in front of compiled llama.cpp servers, running Qwen 3.8 27B through LM Studio and OpenCode with downloaded skills — and says it worked correctly almost immediately, with the remaining effort going into extra features and optimization rather than getting it functional at all. They expect it to eventually handle most of what they currently use Opus 4.8 for at work, just at a slower pace, and say it's already usable for that today.

VERIFIABLE SOURCE

Reported anonymously by an r/LocalLLM contributor · score 2

View comment