A daily local coder on an RTX 5060 8GB and 32GB of RAM finds responsiveness under memory pressure matters more than benchmark speed.
Start with your machine.
Capacity is the reported total for this setup. Multi-GPU and unified-memory layouts can behave differently.
A raw CSV export from a sensor goes into a local model that produces a report, a pattern the owner treats as a deterministic automation overseen by the model.
Hermes uses a multi-GPU local server as a coding subagent, with cron inventorying unfinished work overnight and a cloud model reviewing the result.
A short but clean division of labor: a frontier model decomposes work into tickets, and a local Qwen model implements them on production code, daily.
After trying Claude Code with task-farming and a Qwen-as-orchestrator experiment, this person settled into two parallel workflows — Orca as the local workhorse, and OpenCode or Bionic for solo agent development.