This developer works under NDAs strict enough that a data leak could end their career, so sensitive work cannot be sent to cloud APIs. They use Qwen 3.6 35B for small, well-defined coding tasks rather than asking a local agent to plan and execute an entire project autonomously. The workflow requires going step by step and tuning memory, KV cache, mmap, temperature, and model selection to avoid loops on limited hardware. They consider that operational knowledge and the surrounding workflow as important as the parameter count.
Reported anonymously by an r/LocalLLM contributor · score 2