Can I run
this locally?
Pick the capacity your machine reports and what you want to do. You get the setups people already run within that budget. Nothing is estimated: a setup without a reported capacity never appears here, and we never invent a number.
Reported capacity at or below 97 GB+. 65 setups match.
A teacher runs an entire exam-correction pipeline on a single 3090, from generating answer sheets to scoring hundreds of students automatically.
Qwen 3.8 Flash NextA cloned five-store e-commerce operation — cart, mailings, WhatsApp login, payment gateways — runs day to day on a single 7900XT, turning a real monthly profit.
Qwen 3.8 27BA 3090 owner runs Qwen 3.8 27B daily in Cline at a 96K context by quantising the KV cache to q8_0, reporting about 43 tok/s once the model stops thinking.
Qwen 3.8 27B · OllamaA compact configuration note: two RX 9070 XT cards give 32GB total VRAM and 50-60 tok/s at a 200K context on Qwen 3.8 27B in MXFP4.
Qwen 3.8 27BClient confidentiality rules out cloud APIs for this freelancer, who spreads different model sizes across three separate machines.
Kimi K3A freelance worker keeps restricted client work local on a large multi-machine setup that also doubles as a Blender and data-processing workstation.
Kimi K3A developer uses Qwen 27B as a carefully supervised coding partner inside Cline, with explicit planning and changelog-based context management.
Qwen 3.8 27BA local Qwen model maintains a website, answers email and runs recurring jobs on a dedicated Radeon workstation.
Qwen 3.8 27BRunning Qwen 3.8 27B on the same Mac they work on, this developer credits per-feature domain glossaries and tight scoping — not the model — for making local coding reliable.
Qwen 3.8 27BA detailed personal checklist for getting a local setup to a genuinely reliable state — quant choice, inference engine by OS, and only then, real tasks.
Qwen 3.8 27BRather than one big general model, this setup uses several small, fast ones for narrowly scoped, deterministic jobs — starting with OCR and metadata for a self-hosted Paperless document server.
Gemma 4 12BThe same low-key homelab setup also runs a semi-agentic wedding planner that reads email, keeps a running database and document trail, and answers questions about where things stand — plus a separate loop that just deletes marketing mail.
Qwen 3.6 35B-A3BA four-bit Qwen 3.8 quant at 200K context handles complete app builds through OpenCode, left running until it hits a decision point.
Qwen 3.8 27BA two-3090 local setup runs Qwen 27B for coding and infrastructure work, including bug fixing and quality-of-life automation.
Qwen 3.8 27BA legal user has Claude build task plans while a local Qwen 27B executes them, with anonymisation, OCR for difficult documents and a RAG pipeline using nomic embeddings.
Qwen 3.8 27BA developer with 20 years of experience runs three concurrent coding sessions at about 150 tok/s each on a single RTX 5090 and has been cloud-free for three months.
Stack not specifiedTwo 12GB 3060s handle both code review and customer-support lookups, reasoning read-only over live account data and internal docs.
Qwen 3.6 35B-A3BThis company believes uploading user data to a cloud model would be illegal for their use case, so email and user-text automation runs on a Gemma 4 build on a single Radeon Pro AI R9700.
Gemma 4This person kept paying for Claude Max specifically so they'd have a real comparison point for their own coding hardware — and admits the math still doesn't favor local.
Stack not specifiedProving out AI tagging and logging for a media archive on one of the smallest cards in the thread, then building toward a self-hosted digital asset manager.
Gemma 4 8BThree GPUs spread across machines run Qwen 3.8 27B through a custom stateful agent layer to build a personal finance and portfolio tracker without shipping documents to the cloud.
Qwen 3.8 27B · nInferAfter six months of tinkering, Qwen 3.8 27B justified a 7900XTX plus two DGX Sparks, taking over a job the estimating team was spending dozens of hours a week on.
Qwen 3.8 27BOn a VRAM-constrained 5070 Ti, this person keeps local strictly in the 'learning and playing around' lane, and routes actual work through Claude with local models as a validation step at most.
Stack not specifiedA large enterprise deployment: eight A100 80GBs running Qwen 3.8 under SGLang with tensor parallelism, while the company brings newer B200s online for bigger models.
Qwen 3.8 27B · SGLangBeyond personal daily use, this person stood up a Spark cluster for a client so their employees get an internal, compliance-friendly LLM platform.
Stack not specifiedOn a 12GB 5070, a heavily quantized Qwen 3.8 27B through a custom Hermes backend is, in their words, the first local model that's felt confidently productive.
Qwen 3.8 27BAn unusual layered setup: the local model isn't the coding agent — it's the thing an outer coding agent uses to build and stress-test a separate 'inner' agent.
Qwen 3.8 27BA NUC12 with a 16GB Intel Arc A770M handles a 262K-context Qwen build entirely in VRAM, coding in the background while the owner does something else.
Qwen 3.6 35B-A3B · llama.cppA Python collector pulls read-only configs, firewall rules, and network health data every day and feeds it into a local RAG pipeline, so the advisor always knows the current state of the homelab without ever being able to touch it.
Qwen 3.8 27BPlaying a detective in a large GTA V roleplay server, this person used a local model to turn thousands of in-game arrest reports and applications into structured criminal profiles.
Stack not specifiedA vision model on a modest 3070 turns raw detection events from a home-built surveillance system into plain-language descriptions.
Qwen3-VL-4BThe same Qwen 3.8 27B instance that drives this person's Hermes agent also sits as an MCP tool for Codex, reviewing its proposals and its final output — and the task simply stops if the two can't agree.
Qwen 3.8 27BReferral triage that extracts details, standardizes them, and suggests categories and tests — with the final call always left to a doctor, and nothing ever leaving the network.
Stack not specifiedOn an RTX 4070 and 128GB of RAM, a Hermes agent is being set up to query a private SQL database of hobby equipment and watch prices on the owner's behalf.
LM StudioA contributor uses a quiet Qwen 3.5 9B assistant on an RTX 5060 laptop for planning, documentation, research and private questions, while a Radeon desktop handles heavier work.
Qwen 3.5 9BA non-professional developer uses a local model in small, iterative steps to build a long-term personal game project.
Stack not specifiedA garage server with two Radeon Pro AI R9700 cards supports coding, homelab maintenance and an evolving Cline/Open WebUI workflow.
Qwen 3.6 35B-A3B · Open WebUIAn RTX 5090 runs recurring agents for project maintenance, site contributions and personalized news monitoring.
Qwen 3.8 27BA contractor whose clients forbid commercial providers from reading their source code runs unit tests and code reviews locally on a 3090, with a long-context llama.cpp setup.
Qwen 3.8 27B · llama.cppA game developer on an RTX 5080 uses Gemma 4 26B for background passes in a text adventure engine, while larger models handle knowledge-heavy generation.
Gemma 4 26BTo keep email and business data away from cloud providers, this owner runs an abliterated Qwen 27B on an RX 9060 XT with a vision model on an old GTX 1650.
Qwen 3.8 27BA custom harness layers a written skill and an orchestrator agent that dispatches tasks to five subagents, each starting from a clean context with hard test gates.
llama.cppA four-DGX-Spark cluster under vLLM serves several concurrent agentic coding and research sessions, with unquantized weights and honest notes on concurrency limits.
DeepSeek V4 Flash · vLLMA small research group built a deterministic local pipeline that chunks project data, finds it again with a multimodal RAG and matches it into each project's own report structure.
Qwen 3.8 9BRather than chatting with a local model, this user runs local document Q&A, summarisation, translation and grammar correction on an M4 Pro, then built a desktop app to tie the pieces together.
Gemma 4 12B · MLXDaily local coding on a 4070 is explicitly not as good as Claude for complex refactoring — but the privacy and iteration speed still make it worth keeping.
Qwen 3.8 27BA self-built server with 1TB of ECC RAM and a mix of used 3090s and P100s produced internal tools whose economic impact is reported at ten times the hardware cost.
Qwen 3.8 27BA writer runs Qwen 3.8 27B locally on two RTX 5060 Ti cards with a model card fine-tuned only to edit books and detect plot holes, not to write.
Qwen 3.8 27BA daily local coder on an RTX 5060 8GB and 32GB of RAM finds responsiveness under memory pressure matters more than benchmark speed.
Qwen · OllamaA single Radeon Pro AI R9700 runs coding work using Lucebox and a recommended 4-bit quant, relying on custom optimized kernels for the card.
LuceboxQwen 3.8 Flash Next on an RTX Pro 6000 takes implementation and review tasks delegated from Codex or Claude, replacing Sonnet and Opus for that slice.
Qwen 3.8 Flash Next · DeepSeek HarnessA researcher in applied mathematics is assembling a fully local pipeline to process roughly 20,000 PDFs, comparing PDF extractors, chunking strategies and small summarisation models.
Gemma 4 12BA custom Linux workstation with an RTX Pro 6000 and 48GB of RAM runs programming, copy editing and summarisation through vLLM, with no cloud use all year.
Qwen 3.8 Flash Next · vLLMA sales professional dropped every frontier subscription and runs work agents on two DGX Sparks, averaging about 200 million tokens per week.
GLM 5.3 FlashAn office replaced roughly $2,000 a month of API consumption with two RTX 5090 machines running Qwen 3.8 27B under vLLM, one for helpdesk and one for ops.
Qwen 3.8 27B · vLLMA custom llama.cpp fork for Ampere cards pushes Qwen 3.8 27B to 90+ tok/s through a 100K context, aimed at caching-heavy agentic work.
Qwen 3.8 27B · llama.cpp (llamAmpere fork)A 3080 churned for about 18 hours to produce roughly 60,000 synthetic helpdesk tickets, complete with email threads and time entries, for testing a ticketing system.
Stack not specifiedUsing a gifted 600GB server and a cheap used P100, this owner reports that a capable fully local setup is doable for under $3,000 today.
Qwen 3.8 27BDaily OCR of math- and table-heavy PDFs into LaTeX-correct Markdown and Word, using Qwen vision models on llama.cpp instead of a smaller model that misreads numbers.
Qwen 3.8 27B · llama.cppA Radeon 6650 XT runs quantized Qwen models through vLLM and Ollama, while a custom three-layer harness handles private, deterministic automation around the clock.
Qwen 2.5 7B · vLLM (ROCm 6.x), Ollama fallbackA self-hosted AI ecosystem spans two Linux servers, assigning different GPUs to long-context coding, inference, speech, RAG and media workloads.
Qwen 3.8 27B · llama.cpp; vLLMA contributor experiments with agents, tools and MCP on an 8GB laptop, using lightweight Gemma models to turn constrained hardware into a practical coding playground.
Gemma 4 · LM StudioA laptop workflow combines Qwen Flash Next, llama.cpp and a DeepSeek harness, using system memory and SSD-backed loading to make a MoE model practical on 12GB of VRAM.
Qwen 3.8 Flash Next · llama.cppQwen3.8-27B (NVFP4 W4A16 + FP8 KV) served on 2x RTX 2080 Ti 22GB over NVLink via a vLLM TP2 fork. 196K context cap, MTP K=3, prefix caching on. Feeds a 1-orchestrator + 7-worker harness: ~1227 tok/s prefill and ~29 tok/s decode solo at 65K, ~192 tok/s aggregate at 8-way. Rejected AWQ/FP8 weights (72% EOS failure), ExLlamaV3 (3-6x slower), and PP=2 (infeasible on 2 GPUs).
unsloth_Qwen3.8-27B-NVFP4 · vLLMA hybrid personal-assistant setup for creative writing, quick searches and Discord bot experiments on an 8GB RTX 4060, where 16GB of system RAM is the real ceiling
Gemma 4 26B A4B QAT · Unsloth StudioBands follow reported capacity: a multi-GPU or unified-memory layout can behave differently. See every capacity guide or the full directory.