This person works with two persona profiles. For general tasks and scripting they use gpt-oss:120b; for coding they also use gpt-oss:120b plus deepseek-coder-v2:16b for Swift, with llama3.2:latest for fast initial prompt handling.
Their central claim is that how a local AI is configured matters far more than tokens per second. The foundation, they argue, is a precise description of identity, well-structured knowledge modules, a clear link between identity and personas, a sensible folder and tag structure for prompts, and correct parameters such as chunk size, overlap and temperature. Setting that up was the time-consuming part; downloading models and a GUI was easy. Speed has been more than sufficient even on large documents, and they suspect people fixated on tok/s are not using local models for real daily work.
Reported anonymously by an r/LocalLLM contributor · score 1