A user relies on local models for multi-hour process orchestration, video workflows, voice tasks and file operations while reserving frontier tools for coding.
Start with your machine.
Capacity is the reported total for this setup. Multi-GPU and unified-memory layouts can behave differently.
A contributor experiments with agents, tools and MCP on an 8GB laptop, using lightweight Gemma models to turn constrained hardware into a practical coding playground.
A security professional runs uncensored Qwen models for client work on a 64GB MacBook and a 128GB Strix Halo, using a containerised OpenCode setup with ported code-review skills.
On an RTX 4070 and 128GB of RAM, a Hermes agent is being set up to query a private SQL database of hobby equipment and watch prices on the owner's behalf.
A router in front of compiled llama.cpp servers, feeding Qwen 3.8 27B through OpenCode with downloaded skills, worked essentially out of the box.