This is a two-layer agent setup: an outer coding agent is building a separate 'inner' agent, and local models — mainly Qwen 3.8 27B on two Radeon Pro AI R9700s, with Muse Glimmer also proving capable when substituted in — handle the inner agent's specific tasks, like procedural generation, database maintenance, and user interaction. The outer agent designs scenarios and tests for the inner one, which then stress-tests itself, logs friction points, and gets tuned against them. The eventual goal is an agent harness that isn't tied to a specific model or a 'pro-tier' capability level, with some smaller tasks eventually offloaded to 1–2B models.
Reported anonymously by an r/LocalLLM contributor · score 2