Back to directory
Hybrid orchestration · 2026-09-14

Building a two-server local AI ecosystem across 17 GPUs

A self-hosted AI ecosystem spans two Linux servers, assigning different GPUs to long-context coding, inference, speech, RAG and media workloads.

This contributor uses local AI when cloud tokens run out, when work needs to remain private, and for everyday knowledge lookup, ideation and sensitive questions. Local models also control a Kubernetes cluster and assist with application development, depending on the task and the availability of cloud tokens. Their longer-term goal is a complete personal AI ecosystem with model management, remote llama.cpp and vLLM control, an OpenAI-compatible router with server-side MCP, RAG and memory for documents, chats, rules and code repositories, plus text-to-speech, speech-to-text, chat, image generation and ComfyUI services. The system is still evolving, but the contributor already runs distinct GPU roles for long-context coding, multi-conversation inference, TTS/STT, embeddings, reranking, OCR and media experiments.

VERIFIABLE SOURCE

Reported anonymously by an r/LocalLLM contributor · score 0

View comment