With 21–24 GB,
here is what really runs.
Experiences reported by people using their machines every day. The figure is the reported setup capacity; multi-GPU and unified-memory layouts can behave differently.
in this range
Grading 500 student exams with a local model
A teacher runs an entire exam-correction pipeline on a single 3090, from generating answer sheets to scoring hundreds of students automatically.
Video cropping, UI bug fixes, and news roundups on a 24GB card
A detailed personal checklist for getting a local setup to a genuinely reliable state — quant choice, inference engine by OS, and only then, real tasks.
Full app development on a single 3090 with OpenCode
A four-bit Qwen 3.8 quant at 200K context handles complete app builds through OpenCode, left running until it hits a decision point.
PR review and support triage on two budget RTX 3060s
Two 12GB 3060s handle both code review and customer-support lookups, reasoning read-only over live account data and internal docs.
A daily document and writing workflow on a 24GB Mac, wrapped in a desktop app
Rather than chatting with a local model, this user runs local document Q&A, summarisation, translation and grammar correction on an M4 Pro, then built a desktop app to tie the pieces together.
A local-first portfolio tracker fed private financial documents all day
Three GPUs spread across machines run Qwen 3.8 27B through a custom stateful agent layer to build a personal finance and portfolio tracker without shipping documents to the cloud.
Construction estimating work moved onto a 7900XTX and two DGX Sparks
After six months of tinkering, Qwen 3.8 27B justified a 7900XTX plus two DGX Sparks, taking over a job the estimating team was spending dozens of hours a week on.
Processing thousands of in-game documents to build roleplay casefiles
Playing a detective in a large GTA V roleplay server, this person used a local model to turn thousands of in-game arrest reports and applications into structured criminal profiles.
Internal company tools on a scavenged 1TB ECC server with mixed GPUs
A self-built server with 1TB of ECC RAM and a mix of used 3090s and P100s produced internal tools whose economic impact is reported at ten times the hardware cost.
Running a website with Qwen 27B on a 7900XTX
A local Qwen model maintains a website, answers email and runs recurring jobs on a dedicated Radeon workstation.
Unit tests and reviews for NDA-bound client code on a single 3090
A contractor whose clients forbid commercial providers from reading their source code runs unit tests and code reviews locally on a 3090, with a long-context llama.cpp setup.
Fitting Qwen 27B and a 96K context onto a single 3090
A 3090 owner runs Qwen 3.8 27B daily in Cline at a 96K context by quantising the KV cache to q8_0, reporting about 43 tok/s once the model stops thinking.
A llama.cpp fork reaching 90+ tok/s through a 100K context on a 3090
A custom llama.cpp fork for Ampere cards pushes Qwen 3.8 27B to 90+ tok/s through a 100K context, aimed at caching-heavy agentic work.
Generating 60,000 helpdesk tickets overnight instead of burning API credits
A 3080 churned for about 18 hours to produce roughly 60,000 synthetic helpdesk tickets, complete with email threads and time entries, for testing a ticketing system.