Rather than chatting with a local model, this user runs local document Q&A, summarisation, translation and grammar correction on an M4 Pro, then built a desktop app to tie the pieces together.
Start with your machine.
Capacity is the reported total for this setup. Multi-GPU and unified-memory layouts can behave differently.
A researcher in applied mathematics is assembling a fully local pipeline to process roughly 20,000 PDFs, comparing PDF extractors, chunking strategies and small summarisation models.
After experimenting, this user argues a real computer-manipulating agent needs at least a Strix Halo plus an oculink R9700, or a DGX Spark plus a second GPU for auxiliary models.
A local deep-research pipeline using GPT-OSS 120B and Gemma 4 12B produces results on par with cloud deep research at a fraction of the token cost.
Rather than one big general model, this setup uses several small, fast ones for narrowly scoped, deterministic jobs — starting with OCR and metadata for a self-hosted Paperless document server.
One small model does double duty: correcting speech-to-text output and acting as the decision-making 'brain' for home automation triggers.