Daily OCR of math- and table-heavy PDFs into LaTeX-correct Markdown and Word, using Qwen vision models on llama.cpp instead of a smaller model that misreads numbers.
Start with your machine.
Capacity is the reported total for this setup. Multi-GPU and unified-memory layouts can behave differently.
A four-DGX-Spark cluster under vLLM serves several concurrent agentic coding and research sessions, with unquantized weights and honest notes on concurrency limits.
A self-built server with 1TB of ECC RAM and a mix of used 3090s and P100s produced internal tools whose economic impact is reported at ten times the hardware cost.
Qwen 3.8 Flash Next on an RTX Pro 6000 takes implementation and review tasks delegated from Codex or Claude, replacing Sonnet and Opus for that slice.
With access to both Claude Fable and high-end local models, this developer chooses local because a less autonomous model keeps him involved in debugging and research work.
A developer runs Qwen 3.8 Flash Next on a Strix Halo through a Pi harness and reports 40-50 tok/s decode plus 1300 tok/s prefill while maintaining three substantial projects.
A software engineer uses Qwen locally for personal projects and compares it directly with Claude used professionally.
A self-hosted AI ecosystem spans two Linux servers, assigning different GPUs to long-context coding, inference, speech, RAG and media workloads.
A laptop workflow combines Qwen Flash Next, llama.cpp and a DeepSeek harness, using system memory and SSD-backed loading to make a MoE model practical on 12GB of VRAM.
A custom Linux workstation with an RTX Pro 6000 and 48GB of RAM runs programming, copy editing and summarisation through vLLM, with no cloud use all year.
A security professional runs uncensored Qwen models for client work on a 64GB MacBook and a 128GB Strix Halo, using a containerised OpenCode setup with ported code-review skills.
After experimenting, this user argues a real computer-manipulating agent needs at least a Strix Halo plus an oculink R9700, or a DGX Spark plus a second GPU for auxiliary models.