With 13–16 GB,
here is what really runs.
Experiences reported by people using their machines every day. The figure is the reported setup capacity; multi-GPU and unified-memory layouts can behave differently.
in this range
Local is for tinkering; Claude still does the real work
On a VRAM-constrained 5070 Ti, this person keeps local strictly in the 'learning and playing around' lane, and routes actual work through Claude with local models as a validation step at most.
Background coding at 10-15 tok/s on an Intel Arc laptop GPU
A NUC12 with a 16GB Intel Arc A770M handles a 262K-context Qwen build entirely in VRAM, coding in the background while the owner does something else.
Building a personal academic RAG pipeline on a 16GB 5060 Ti
A researcher in applied mathematics is assembling a fully local pipeline to process roughly 20,000 PDFs, comparing PDF extractors, chunking strategies and small summarisation models.
Iterative coding with Qwen 27B on a 4070 Ti Super
A developer uses Qwen 27B as a carefully supervised coding partner inside Cline, with explicit planning and changelog-based context management.
Local models as checker and tracker for a text adventure engine
A game developer on an RTX 5080 uses Gemma 4 26B for background passes in a text adventure engine, while larger models handle knowledge-heavy generation.
Every business app points at a local model on an RX 9060 XT
To keep email and business data away from cloud providers, this owner runs an abliterated Qwen 27B on an RX 9060 XT with a vision model on an old GTX 1650.
Going fully local on a home server for under $3,000
Using a gifted 600GB server and a cheap used P100, this owner reports that a capable fully local setup is doable for under $3,000 today.