Full directory
GUIDE BY CAPACITY

With 25–32 GB,
here is what really runs.

Experiences reported by people using their machines every day. The figure is the reported setup capacity; multi-GPU and unified-memory layouts can behave differently.

12documented uses
in this range
Business automation

Company email automation kept off the cloud for legal reasons

This company believes uploading user data to a cloud model would be illegal for their use case, so email and user-text automation runs on a Gemma 4 build on a single Radeon Pro AI R9700.

Coding

A senior engineer with no need for cloud on a 64GB Strix Halo

An experienced software engineer runs Qwen3-Coder 30B on a 64GB Strix Halo and finds local working better than newer, larger models for well-scoped tasks.

Homelab ops

A daily-refreshed RAG advisor built from your own infrastructure

A Python collector pulls read-only configs, firewall rules, and network health data every day and feeds it into a local RAG pipeline, so the advisor always knows the current state of the homelab without ever being able to touch it.

Hybrid orchestration

A local model and Codex reviewing each other's work, with a hard stop on disagreement

The same Qwen 3.8 27B instance that drives this person's Hermes agent also sits as an MCP tool for Codex, reviewing its proposals and its final output — and the task simply stops if the two can't agree.

Media

A dedicated local editing model for novels, running on two 5060 Ti cards

A writer runs Qwen 3.8 27B locally on two RTX 5060 Ti cards with a model card fine-tuned only to edit books and detect plot holes, not to write.

Coding

One R9700 with card-specific kernels for local coding

A single Radeon Pro AI R9700 runs coding work using Lucebox and a recommended 4-bit quant, relying on custom optimized kernels for the card.

Business automation

Daily agent workflows on an RTX 5090

An RTX 5090 runs recurring agents for project maintenance, site contributions and personalized news monitoring.

Regulated work

Legal work: a cloud planner driving local anonymisation, OCR and RAG

A legal user has Claude build task plans while a local Qwen 27B executes them, with anonymisation, OCR for difficult documents and a RAG pipeline using nomic embeddings.

Business automation

Two RTX 5090s running helpdesk and ops for a small office

An office replaced roughly $2,000 a month of API consumption with two RTX 5090 machines running Qwen 3.8 27B under vLLM, one for helpdesk and one for ops.

Hybrid orchestration

An orchestrator dispatching to five fresh-context subagents

A custom harness layers a written skill and an orchestrator agent that dispatches tasks to five subagents, each starting from a clean context with hard test gates.

Coding

Three concurrent coding sessions at 150 tok/s on one 5090

A developer with 20 years of experience runs three concurrent coding sessions at about 150 tok/s each on a single RTX 5090 and has been cloud-free for three months.

Coding

Two Radeon RX 9070 XT cards as a 32GB budget rig

A compact configuration note: two RX 9070 XT cards give 32GB total VRAM and 50-60 tok/s at a 200K context on Qwen 3.8 27B in MXFP4.