This workflow has a local model read through reports from 1,500 companies and produce financial summaries, which then feed a smarter hosted model for the analysis that actually needs stronger reasoning — cutting the token cost of the hosted step considerably. They've also built a small service that exposes their local model to other apps online, for cases where they want an LLM backend without a conversation being retained indefinitely by a cloud provider.
VERIFIABLE SOURCE
View comment Reported anonymously by an r/LocalLLM contributor · score 2