Back to directory
Hybrid orchestration · 2026-09-30

GPT-6.1 Sol as cloud orchestrator with one RTX 3090 worker: wall clock and cost for three small builds

A single RTX 3090 running Qwen 3.8 27B Q4 via llama.cpp, orchestrated by GPT-6.1 Sol in the cloud, finished three small 3D games in 43.4 minutes and $0.17 total.

With GPT-6.1 Sol in the cloud splitting and reviewing the work and a single RTX 3090 24GB running Qwen 3.8 27B at Q4 through llama.cpp, three small 3D games (Pool, Bowling, Foosball) took 43.4 minutes and $0.17 total, versus 114.2 minutes for the 3090 alone and 6.6 minutes and $0.75 for Sol alone. The card was the bottleneck throughout, and the author notes 24GB is roughly the floor for a 27B worker, since a 12GB card was too slow on an earlier run. Software was the open-source Atomic Agent in its Fusion mode, where the cloud orchestrator is locked out of writing and the managed llama.cpp server sizes parallel slots from available VRAM. This is a hybrid workflow rather than a replacement: cloud orchestration stays in the loop, power is excluded from the dollar column, and each result is one run per game.

VERIFIABLE SOURCE

Reported anonymously by an r/LocalLLM contributor · 12 upvotes at capture

View source

Useful references

Community-provided links related to this setup, workflow or measurements.

SETUP HISTORY

Snapshots over time

1 version
CURRENT2026-09-30GPT-6.1 Sol as cloud orchestrator with one RTX 3090 worker: wall clock and cost for three small builds

A single RTX 3090 running Qwen 3.8 27B Q4 via llama.cpp, orchestrated by GPT-6.1 Sol in the cloud, finished three small 3D games in 43.4 minutes and $0.17 total.

Hybrid orchestration1 machineQwen 3.8 27Bllama.cpp

This is the currently published snapshot.