This person runs Qwen 3.8 27B with the context pushed to 524K and a 160K maximum output, saying it handles everything they throw at it. Current projects include a Minecraft-style voxel clone in Godot with C# for their nephews, a React wallpaper site with automated image-tier generation, and a Vitest physics simulator module.
They report a concrete agent run: one turn, 73 steps, 15m48s of model time and 1m45s of tool calls, with 1.5s average time-to-first-token, 118 tok/s and a 99% cache hit rate over 13.8M input tokens and 99.4K output tokens. They stress that a large context is what makes the harness work: below 262K, cache hits stayed near 55%, whereas general system work only needs 32K-64K.
Reported anonymously by an r/LocalLLM contributor · score 2