This person uses local models every day on an M4 Pro with 24GB of unified memory, treating them as part of a normal workflow rather than a chatbot: searching and asking questions over local documents, summarising and translating web pages, grammar and writing correction in Chrome, and occasionally using Gmail or Drive as extra context.
They switch between GGUF and MLX models depending on the task, finding Gemma 4 12B MLX 6-bit a good compromise when output quality matters and smaller models preferable when latency matters. Wiring runtimes, RAG and browser extensions together by hand became annoying, so they built a desktop app around the workflow. Even so, they are clear that local does not replace Claude or OpenAI for difficult reasoning or coding; its value is repetitive everyday tasks, private data, and integration with the machine.
Reported anonymously by an r/LocalLLM contributor · score 1