With GPT-6.1 Sol in the cloud splitting and reviewing the work and a single RTX 3090 24GB running Qwen 3.8 27B at Q4 through llama.cpp, three small 3D games (Pool, Bowling, Foosball) took 43.4 minutes and $0.17 total, versus 114.2 minutes for the 3090 alone and 6.6 minutes and $0.75 for Sol alone. The card was the bottleneck throughout, and the author notes 24GB is roughly the floor for a 27B worker, since a 12GB card was too slow on an earlier run. Software was the open-source Atomic Agent in its Fusion mode, where the cloud orchestrator is locked out of writing and the managed llama.cpp server sizes parallel slots from available VRAM. This is a hybrid workflow rather than a replacement: cloud orchestration stays in the loop, power is excluded from the dollar column, and each result is one run per game.
Reported anonymously by an r/LocalLLM contributor · 12 upvotes at capture
Useful references
Community-provided links related to this setup, workflow or measurements.