This person lays out the groundwork they think people skip: checking VRAM and RAM speed, picking a model that actually fits (they land on Qwen 3.8 27B as the practical floor), matching the inference engine to the OS — llama.cpp under 40GB VRAM on Windows, vLLM or SGLang on Linux with more — and only then tuning sampling parameters by hand rather than trusting a cloud model's defaults. With that groundwork done on a 24GB card at an IQ4 quant, they've used it to crop video by having it work out the right pixel offsets and write the script itself, fix UI bugs before a maintainer gets to them, write PowerShell scripts, and compile personal news digests from sources they choose.
Reported anonymously by an r/LocalLLM contributor · score 3