Full directory
GUIDE BY CAPACITY

With 9–12 GB,
here is what really runs.

Experiences reported by people using their machines every day. The figure is the reported setup capacity; multi-GPU and unified-memory layouts can behave differently.

6documented uses
in this range
Documents

A fully local OCR and RAG pipeline for a self-hosted document server

Rather than one big general model, this setup uses several small, fast ones for narrowly scoped, deterministic jobs — starting with OCR and metadata for a self-hosted Paperless document server.

Personal assistant

An agent that plans a wedding from your inbox

The same low-key homelab setup also runs a semi-agentic wedding planner that reads email, keeps a running database and document trail, and answers questions about where things stand — plus a separate loop that just deletes marketing mail.

Coding

The first local model this person actually trusts with real work

On a 12GB 5070, a heavily quantized Qwen 3.8 27B through a custom Hermes backend is, in their words, the first local model that's felt confidently productive.

Coding

Good enough for 80% of coding, honest about the other 20%

Daily local coding on a 4070 is explicitly not as good as Claude for complex refactoring — but the privacy and iteration speed still make it worth keeping.

Personal assistant

A Hermes agent coordinating hobbies, inventories and price tracking

On an RTX 4070 and 128GB of RAM, a Hermes agent is being set up to query a private SQL database of hobby equipment and watch prices on the owner's behalf.

Coding

Qwen Flash Next on a 12GB RTX 3500 Ada laptop

A laptop workflow combines Qwen Flash Next, llama.cpp and a DeepSeek harness, using system memory and SSD-backed loading to make a MoE model practical on 12GB of VRAM.