Daily OCR of math- and table-heavy PDFs into LaTeX-correct Markdown and Word, using Qwen vision models on llama.cpp instead of a smaller model that misreads numbers.
Start with your machine.
Capacity is the reported total for this setup. Multi-GPU and unified-memory layouts can behave differently.
A product team developed their agentic tooling against local inference on a Strix Halo and midrange GPUs, then retargeted the same harness to cloud inference once it worked.
A site reliability engineer builds custom apps and a theming engine inside Dynatrace with local Qwen, using a per-repo manifest so the model only reads the code it actually needs.
A self-built server with 1TB of ECC RAM and a mix of used 3090s and P100s produced internal tools whose economic impact is reported at ten times the hardware cost.
A writer runs Qwen 3.8 27B locally on two RTX 5060 Ti cards with a model card fine-tuned only to edit books and detect plot holes, not to write.
A researcher in applied mathematics is assembling a fully local pipeline to process roughly 20,000 PDFs, comparing PDF extractors, chunking strategies and small summarisation models.
A developer uses Qwen 27B as a carefully supervised coding partner inside Cline, with explicit planning and changelog-based context management.
A local Qwen model maintains a website, answers email and runs recurring jobs on a dedicated Radeon workstation.
A garage server with two Radeon Pro AI R9700 cards supports coding, homelab maintenance and an evolving Cline/Open WebUI workflow.
A local Hermes setup handles document review, cron jobs and controlled computer actions with a NAS mirror as a safety net.
A privacy-focused workflow combines local models on an M2 Ultra and M3 Max with GCP and OpenRouter for RAG and security work.
A developer uses Qwen 27B for structured feature design and background implementation on the same Mac used for a SaaS project.