Running Gemma 4 12B within about 10GB of VRAM on an older GPU, this person uses it all day as both a speech-to-text corrector and the reasoning layer behind their home automation — it gets triggered by events and decides what actions to take. They mention working on several other use cases too, some of which they expect could run on an even smaller model than this one.
VERIFIABLE SOURCE
View comment Reported anonymously by an r/LocalLLM contributor · score 1