Hybrid orchestrationGenesis archive2026-09
Building a two-server local AI ecosystem across 17 GPUsA self-hosted AI ecosystem spans two Linux servers, assigning different GPUs to long-context coding, inference, speech, RAG and media workloads.
Every published setup that reports this GPU, with its models, quantization, runtime and reported results. Use the directory filters below to narrow by capacity, model or runtime.
Capacity is the reported total for this setup. Multi-GPU and unified-memory layouts can behave differently.
A self-hosted AI ecosystem spans two Linux servers, assigning different GPUs to long-context coding, inference, speech, RAG and media workloads.
A 3080 churned for about 18 hours to produce roughly 60,000 synthetic helpdesk tickets, complete with email threads and time entries, for testing a ticketing system.
Hermes uses a multi-GPU local server as a coding subagent, with cron inventorying unfinished work overnight and a cloud model reviewing the result.