This contributor runs Hermes alongside Codex, with a local server built from RTX 3080, RTX 5080, and RTX 3060 Ti cards serving a Qwen 27B model at Q4. The local model is used as a subagent rather than the primary assistant. A cron job inventories unfinished work from the day and lets the local system complete it overnight, while a cloud model reviews the result. The contributor describes the local path as somewhat slow but valuable because the marginal token cost is effectively zero once the hardware is running.
VERIFIABLE SOURCE
View comment Reported anonymously by an r/LocalLLM contributor · score 1