A site reliability engineer builds custom apps and a theming engine inside Dynatrace with local Qwen, using a per-repo manifest so the model only reads the code it actually needs.
Start with your machine.
Capacity is the reported total for this setup. Multi-GPU and unified-memory layouts can behave differently.
An experienced software engineer finds sub-30 tok/s local inference on a 128GB Strix Halo more than adequate for deep coding work.
A garage server with two Radeon Pro AI R9700 cards supports coding, homelab maintenance and an evolving Cline/Open WebUI workflow.
A freelance worker keeps restricted client work local on a large multi-machine setup that also doubles as a Blender and data-processing workstation.
A local Hermes setup handles document review, cron jobs and controlled computer actions with a NAS mirror as a safety net.
A full personal-operations stack runs on one 128GB Strix Halo machine, with several named agents reachable through a single Telegram chat and each handling a different part of daily life.
A security professional runs uncensored Qwen models for client work on a 64GB MacBook and a 128GB Strix Halo, using a containerised OpenCode setup with ported code-review skills.
A researcher runs Qwen 3.8 Flash Next and Gemma 4 31B on a 128GB Strix Halo, splitting strategy and summarisation across models and advising against 64GB machines.
A 512GB M3 Ultra Mac Studio running models through oMLX handles most agentic coding work, with measured speeds across several models listed by the owner.
A contributor uses a quiet Qwen 3.5 9B assistant on an RTX 5060 laptop for planning, documentation, research and private questions, while a Radeon desktop handles heavier work.
Client confidentiality rules out cloud APIs for this freelancer, who spreads different model sizes across three separate machines.
Cron jobs and document review through Hermes, with files kept on a NAS behind a one-way mirror so a bad run can't take real data with it.