This contributor runs sub-10B models on an RTX 2070 Ti Super and argues that small, specialized models can be more useful than a single large general model when surrounded by good skillware and memory layers. Their custom system exposes more than 30 capabilities through natural-language triggers, covering email, monitoring, smart-contract vetting, finance, DNA literature comparison, stock and supply-chain analysis, generative media pipelines, and social posting. The setup is explicitly local and offline, with Ollama as the model runtime. The most important part of the system is not the model alone but the surrounding semantic layers, vector databases, embeddings, and deterministic workflows.
Reported anonymously by an r/LocalLLM contributor · score 2