This contributor uses a single desktop, an RTX 4060 with 8GB of VRAM, an i5-12400F and 16GB of DDR5, for everyday local work: creative writing, quick searches, and Discord bot experiments. Two Gemma 4 variants, a larger 26B A4B and a smaller E4B, both in QAT form, run through Unsloth Studio at roughly 40 tokens per second with up to a 64K context. The honest assessment is that local beats ChatGPT for creative writing and roughly matches Gemini depending on the task, and it handles searches well. The real limit is memory: 16GB of system RAM is not enough to keep the model loaded and use the desktop comfortably at the same time.
Submitted through the vram.wiki contributor dashboard · reviewed before publication.