This person built a local version of the deep-research feature offered by Claude, ChatGPT and Gemini. Their pipeline runs GPT-OSS 120B through the Brave API, which costs about $5 a month with a credit included, and uses Gemma 4 12B QAT for summarisation, which they found performed just as well as 26B and 31B models for that task.
A run takes five to ten minutes, and when they tested the same research prompts against cloud deep research the results were on par. Their motivation is cost: they describe the token consumption of cloud deep-research runs as a rip-off.
Reported anonymously by an r/LocalLLM contributor · score 2