Running two RTX 3060s for a combined 24GB of VRAM, this team uses a local model for pull-request review as well as internal support triage: when a support ticket comes in about a client's account, the model can fetch and reason over that account's data read-only, plus search a vector database of guides and other reference PDFs, to give the support team a starting point. They're currently on Qwen 3.6 35B-A3B and are considering moving up to Qwen 3.8 27B once they add two 3090s.
VERIFIABLE SOURCE
View comment Reported anonymously by an r/LocalLLM contributor · score 2