START FROM YOUR MACHINE

Can I run
this locally?

Pick the capacity your machine reports and what you want to do. You get the setups people already run within that budget. Nothing is estimated: a setup without a reported capacity never appears here, and we never invent a number.

Reported capacity at or below 97 GB+. 65 setups match.

24 GBDocuments
Grading 500 student exams with a local model

A teacher runs an entire exam-correction pipeline on a single 3090, from generating answer sheets to scoring hundreds of students automatically.

Qwen 3.8 Flash Next
20 GBBusiness automation
Running a small e-commerce business on Qwen 3.8 27B

A cloned five-store e-commerce operation — cart, mailings, WhatsApp login, payment gateways — runs day to day on a single 7900XT, turning a real monthly profit.

Qwen 3.8 27B
24 GBCoding
Fitting Qwen 27B and a 96K context onto a single 3090

A 3090 owner runs Qwen 3.8 27B daily in Cline at a 96K context by quantising the KV cache to q8_0, reporting about 43 tok/s once the model stops thinking.

Qwen 3.8 27B · Ollama
32 GBCoding
Two Radeon RX 9070 XT cards as a 32GB budget rig

A compact configuration note: two RX 9070 XT cards give 32GB total VRAM and 50-60 tok/s at a 200K context on Qwen 3.8 27B in MXFP4.

Qwen 3.8 27B
96 GBRegulated work
Three rigs for freelance work an NDA won't let leave the building

Client confidentiality rules out cloud APIs for this freelancer, who spreads different model sizes across three separate machines.

Kimi K3
96 GBRegulated work
Privacy-constrained freelance work on a four-3090 workstation

A freelance worker keeps restricted client work local on a large multi-machine setup that also doubles as a Blender and data-processing workstation.

Kimi K3
16 GBCoding
Iterative coding with Qwen 27B on a 4070 Ti Super

A developer uses Qwen 27B as a carefully supervised coding partner inside Cline, with explicit planning and changelog-based context management.

Qwen 3.8 27B
24 GBBusiness automation
Running a website with Qwen 27B on a 7900XTX

A local Qwen model maintains a website, answers email and runs recurring jobs on a dedicated Radeon workstation.

Qwen 3.8 27B
64 GBCoding
Coding a SaaS product on an M1 Max, one well-scoped feature at a time

Running Qwen 3.8 27B on the same Mac they work on, this developer credits per-feature domain glossaries and tight scoping — not the model — for making local coding reliable.

Qwen 3.8 27B
24 GBBatch utility
Video cropping, UI bug fixes, and news roundups on a 24GB card

A detailed personal checklist for getting a local setup to a genuinely reliable state — quant choice, inference engine by OS, and only then, real tasks.

Qwen 3.8 27B
12 GBDocuments
A fully local OCR and RAG pipeline for a self-hosted document server

Rather than one big general model, this setup uses several small, fast ones for narrowly scoped, deterministic jobs — starting with OCR and metadata for a self-hosted Paperless document server.

Gemma 4 12B
12 GBPersonal assistant
An agent that plans a wedding from your inbox

The same low-key homelab setup also runs a semi-agentic wedding planner that reads email, keeps a running database and document trail, and answers questions about where things stand — plus a separate loop that just deletes marketing mail.

Qwen 3.6 35B-A3B
24 GBCoding
Full app development on a single 3090 with OpenCode

A four-bit Qwen 3.8 quant at 200K context handles complete app builds through OpenCode, left running until it hits a decision point.

Qwen 3.8 27B
48 GBCoding
Qwen 27B coding and infrastructure on two 3090s

A two-3090 local setup runs Qwen 27B for coding and infrastructure work, including bug fixing and quality-of-life automation.

Qwen 3.8 27B
32 GBRegulated work
Legal work: a cloud planner driving local anonymisation, OCR and RAG

A legal user has Claude build task plans while a local Qwen 27B executes them, with anonymisation, OCR for difficult documents and a RAG pipeline using nomic embeddings.

Qwen 3.8 27B
32 GBCoding
Three concurrent coding sessions at 150 tok/s on one 5090

A developer with 20 years of experience runs three concurrent coding sessions at about 150 tok/s each on a single RTX 5090 and has been cloud-free for three months.

Stack not specified
24 GBBusiness automation
PR review and support triage on two budget RTX 3060s

Two 12GB 3060s handle both code review and customer-support lookups, reasoning read-only over live account data and internal docs.

Qwen 3.6 35B-A3B
32 GBBusiness automation
Company email automation kept off the cloud for legal reasons

This company believes uploading user data to a cloud model would be illegal for their use case, so email and user-text automation runs on a Gemma 4 build on a single Radeon Pro AI R9700.

Gemma 4
96 GBCoding
An RTX 6000 Pro and an M4 Max, benchmarked against an active Claude Max plan

This person kept paying for Claude Max specifically so they'd have a real comparison point for their own coding hardware — and admits the math still doesn't favor local.

Stack not specified
8 GBMedia
AI tagging for a media library on an 8GB card

Proving out AI tagging and logging for a media archive on one of the smallest cards in the thread, then building toward a self-hosted digital asset manager.

Gemma 4 8B
24 GBBusiness automation
A local-first portfolio tracker fed private financial documents all day

Three GPUs spread across machines run Qwen 3.8 27B through a custom stateful agent layer to build a personal finance and portfolio tracker without shipping documents to the cloud.

Qwen 3.8 27B · nInfer
24 GBBusiness automation
Construction estimating work moved onto a 7900XTX and two DGX Sparks

After six months of tinkering, Qwen 3.8 27B justified a 7900XTX plus two DGX Sparks, taking over a job the estimating team was spending dozens of hours a week on.

Qwen 3.8 27B
16 GBPersonal assistant
Local is for tinkering; Claude still does the real work

On a VRAM-constrained 5070 Ti, this person keeps local strictly in the 'learning and playing around' lane, and routes actual work through Claude with local models as a validation step at most.

Stack not specified
640 GBCoding
Eight A100s in tensor-parallel, inside a company with a hundred more

A large enterprise deployment: eight A100 80GBs running Qwen 3.8 under SGLang with tensor parallelism, while the company brings newer B200s online for bigger models.

Qwen 3.8 27B · SGLang
128 GBRegulated work
A HIPAA-style internal LLM platform, built on a cluster of DGX Sparks

Beyond personal daily use, this person stood up a Spark cluster for a client so their employees get an internal, compliance-friendly LLM platform.

Stack not specified
12 GBCoding
The first local model this person actually trusts with real work

On a 12GB 5070, a heavily quantized Qwen 3.8 27B through a custom Hermes backend is, in their words, the first local model that's felt confidently productive.

Qwen 3.8 27B
64 GBCoding
Using a local model to stress-test the agent a coding agent is building

An unusual layered setup: the local model isn't the coding agent — it's the thing an outer coding agent uses to build and stress-test a separate 'inner' agent.

Qwen 3.8 27B
16 GBCoding
Background coding at 10-15 tok/s on an Intel Arc laptop GPU

A NUC12 with a 16GB Intel Arc A770M handles a 262K-context Qwen build entirely in VRAM, coding in the background while the owner does something else.

Qwen 3.6 35B-A3B · llama.cpp
32 GBHomelab ops
A daily-refreshed RAG advisor built from your own infrastructure

A Python collector pulls read-only configs, firewall rules, and network health data every day and feeds it into a local RAG pipeline, so the advisor always knows the current state of the homelab without ever being able to touch it.

Qwen 3.8 27B
24 GBBatch utility
Processing thousands of in-game documents to build roleplay casefiles

Playing a detective in a large GTA V roleplay server, this person used a local model to turn thousands of in-game arrest reports and applications into structured criminal profiles.

Stack not specified
8 GBMedia
Describing security camera events on an 8GB card

A vision model on a modest 3070 turns raw detection events from a home-built surveillance system into plain-language descriptions.

Qwen3-VL-4B
32 GBHybrid orchestration
A local model and Codex reviewing each other's work, with a hard stop on disagreement

The same Qwen 3.8 27B instance that drives this person's Hermes agent also sits as an MCP tool for Codex, reviewing its proposals and its final output — and the task simply stops if the two can't agree.

Qwen 3.8 27B
128 GBRegulated work
A cardiologist's referral system, fully air-gapped

Referral triage that extracts details, standardizes them, and suggests categories and tests — with the final call always left to a doctor, and nothing ever leaving the network.

Stack not specified
12 GBPersonal assistant
A Hermes agent coordinating hobbies, inventories and price tracking

On an RTX 4070 and 128GB of RAM, a Hermes agent is being set up to query a private SQL database of hobby equipment and watch prices on the owner's behalf.

LM Studio
8 GBPersonal assistant
A quiet local personal assistant on a two-machine setup

A contributor uses a quiet Qwen 3.5 9B assistant on an RTX 5060 laptop for planning, documentation, research and private questions, while a Radeon desktop handles heavier work.

Qwen 3.5 9B
64 GBCoding
Building a game iteratively with a local Pi harness

A non-professional developer uses a local model in small, iterative steps to build a long-term personal game project.

Stack not specified
64 GBHomelab ops
Coding and homelab maintenance on a dual-R9700 server

A garage server with two Radeon Pro AI R9700 cards supports coding, homelab maintenance and an evolving Cline/Open WebUI workflow.

Qwen 3.6 35B-A3B · Open WebUI
32 GBBusiness automation
Daily agent workflows on an RTX 5090

An RTX 5090 runs recurring agents for project maintenance, site contributions and personalized news monitoring.

Qwen 3.8 27B
24 GBRegulated work
Unit tests and reviews for NDA-bound client code on a single 3090

A contractor whose clients forbid commercial providers from reading their source code runs unit tests and code reviews locally on a 3090, with a long-context llama.cpp setup.

Qwen 3.8 27B · llama.cpp
16 GBMedia
Local models as checker and tracker for a text adventure engine

A game developer on an RTX 5080 uses Gemma 4 26B for background passes in a text adventure engine, while larger models handle knowledge-heavy generation.

Gemma 4 26B
16 GBBusiness automation
Every business app points at a local model on an RX 9060 XT

To keep email and business data away from cloud providers, this owner runs an abliterated Qwen 27B on an RX 9060 XT with a vision model on an old GTX 1650.

Qwen 3.8 27B
32 GBHybrid orchestration
An orchestrator dispatching to five fresh-context subagents

A custom harness layers a written skill and an orchestrator agent that dispatches tasks to five subagents, each starting from a clean context with hard test gates.

llama.cpp
128 GBHybrid orchestration
Four DGX Sparks running agentic work all day

A four-DGX-Spark cluster under vLLM serves several concurrent agentic coding and research sessions, with unquantized weights and honest notes on concurrency limits.

DeepSeek V4 Flash · vLLM
8 GBDocuments
Auto-filling varied lab reports from a multimodal RAG on an RTX 4060

A small research group built a deterministic local pipeline that chunks project data, finds it again with a multimodal RAG and matches it into each project's own report structure.

Qwen 3.8 9B
24 GBDocuments
A daily document and writing workflow on a 24GB Mac, wrapped in a desktop app

Rather than chatting with a local model, this user runs local document Q&A, summarisation, translation and grammar correction on an M4 Pro, then built a desktop app to tie the pieces together.

Gemma 4 12B · MLX
12 GBCoding
Good enough for 80% of coding, honest about the other 20%

Daily local coding on a 4070 is explicitly not as good as Claude for complex refactoring — but the privacy and iteration speed still make it worth keeping.

Qwen 3.8 27B
24 GBBusiness automation
Internal company tools on a scavenged 1TB ECC server with mixed GPUs

A self-built server with 1TB of ECC RAM and a mix of used 3090s and P100s produced internal tools whose economic impact is reported at ten times the hardware cost.

Qwen 3.8 27B
32 GBMedia
A dedicated local editing model for novels, running on two 5060 Ti cards

A writer runs Qwen 3.8 27B locally on two RTX 5060 Ti cards with a model card fine-tuned only to edit books and detect plot holes, not to write.

Qwen 3.8 27B
8 GBCoding
Coding on an 8GB 5060, where tuning the GPU/CPU split matters more than tok/s

A daily local coder on an RTX 5060 8GB and 32GB of RAM finds responsiveness under memory pressure matters more than benchmark speed.

Qwen · Ollama
32 GBCoding
One R9700 with card-specific kernels for local coding

A single Radeon Pro AI R9700 runs coding work using Lucebox and a recommended 4-bit quant, relying on custom optimized kernels for the card.

Lucebox
96 GBCoding
A workstation GPU as a drop-in subagent for Codex and Claude

Qwen 3.8 Flash Next on an RTX Pro 6000 takes implementation and review tasks delegated from Codex or Claude, replacing Sonnet and Opus for that slice.

Qwen 3.8 Flash Next · DeepSeek Harness
16 GBDocuments
Building a personal academic RAG pipeline on a 16GB 5060 Ti

A researcher in applied mathematics is assembling a fully local pipeline to process roughly 20,000 PDFs, comparing PDF extractors, chunking strategies and small summarisation models.

Gemma 4 12B
96 GBCoding
A cloud-free workstation: RTX Pro 6000, vLLM and a DeepSeek harness

A custom Linux workstation with an RTX Pro 6000 and 48GB of RAM runs programming, copy editing and summarisation through vLLM, with no cloud use all year.

Qwen 3.8 Flash Next · vLLM
128 GBBusiness automation
Running sales work on GLM 5.3 Flash across two DGX Sparks, 200M tokens a week

A sales professional dropped every frontier subscription and runs work agents on two DGX Sparks, averaging about 200 million tokens per week.

GLM 5.3 Flash
32 GBBusiness automation
Two RTX 5090s running helpdesk and ops for a small office

An office replaced roughly $2,000 a month of API consumption with two RTX 5090 machines running Qwen 3.8 27B under vLLM, one for helpdesk and one for ops.

Qwen 3.8 27B · vLLM
24 GBCoding
A llama.cpp fork reaching 90+ tok/s through a 100K context on a 3090

A custom llama.cpp fork for Ampere cards pushes Qwen 3.8 27B to 90+ tok/s through a 100K context, aimed at caching-heavy agentic work.

Qwen 3.8 27B · llama.cpp (llamAmpere fork)
24 GBBatch utility
Generating 60,000 helpdesk tickets overnight instead of burning API credits

A 3080 churned for about 18 hours to produce roughly 60,000 synthetic helpdesk tickets, complete with email threads and time entries, for testing a ticketing system.

Stack not specified
16 GBHomelab ops
Going fully local on a home server for under $3,000

Using a gifted 600GB server and a cheap used P100, this owner reports that a capable fully local setup is doable for under $3,000 today.

Qwen 3.8 27B
16 GBDocuments
Turning math-heavy PDFs into clean Markdown and Word with Qwen vision

Daily OCR of math- and table-heavy PDFs into LaTeX-correct Markdown and Word, using Qwen vision models on llama.cpp instead of a smaller model that misreads numbers.

Qwen 3.8 27B · llama.cpp
8 GBHybrid orchestration
A three-layer AMD ROCm agent architecture for 24/7 local automation

A Radeon 6650 XT runs quantized Qwen models through vLLM and Ollama, while a custom three-layer harness handles private, deterministic automation around the clock.

Qwen 2.5 7B · vLLM (ROCm 6.x), Ollama fallback
364 GBHybrid orchestration
Building a two-server local AI ecosystem across 17 GPUs

A self-hosted AI ecosystem spans two Linux servers, assigning different GPUs to long-context coding, inference, speech, RAG and media workloads.

Qwen 3.8 27B · llama.cpp; vLLM
8 GBHomelab ops
Making an 8GB laptop useful through DIY agent workflows

A contributor experiments with agents, tools and MCP on an 8GB laptop, using lightweight Gemma models to turn constrained hardware into a practical coding playground.

Gemma 4 · LM Studio
12 GBCoding
Qwen Flash Next on a 12GB RTX 3500 Ada laptop

A laptop workflow combines Qwen Flash Next, llama.cpp and a DeepSeek harness, using system memory and SSD-backed loading to make a MoE model practical on 12GB of VRAM.

Qwen 3.8 Flash Next · llama.cpp
44 GBHybrid orchestration
27B reasoning model on 2x RTX 2080 Ti (NVFP4, vLLM) for agentic orchestration

Qwen3.8-27B (NVFP4 W4A16 + FP8 KV) served on 2x RTX 2080 Ti 22GB over NVLink via a vLLM TP2 fork. 196K context cap, MTP K=3, prefix caching on. Feeds a 1-orchestrator + 7-worker harness: ~1227 tok/s prefill and ~29 tok/s decode solo at 65K, ~192 tok/s aggregate at 8-way. Rejected AWQ/FP8 weights (72% EOS failure), ExLlamaV3 (3-6x slower), and PP=2 (infeasible on 2 GPUs).

unsloth_Qwen3.8-27B-NVFP4 · vLLM
8 GBPersonal assistant
Creative writing and a Discord bot on an 8GB RTX 4060

A hybrid personal-assistant setup for creative writing, quick searches and Discord bot experiments on an 8GB RTX 4060, where 16GB of system RAM is the real ceiling

Gemma 4 26B A4B QAT · Unsloth Studio

Bands follow reported capacity: a multi-GPU or unified-memory layout can behave differently. See every capacity guide or the full directory.