Back to directory
Coding · 2026-09-11

Eight A100s in tensor-parallel, inside a company with a hundred more

A large enterprise deployment: eight A100 80GBs running Qwen 3.8 under SGLang with tensor parallelism, while the company brings newer B200s online for bigger models.

Inside a large company with more than a hundred 80GB A100 GPUs, this team runs Qwen 3.8 across eight of them using SGLang with tensor parallelism set to 8, primarily for coding. They're now working on deploying larger models on newly arrived B200 GPUs, suggesting the local footprint here is still growing rather than settled.

VERIFIABLE SOURCE

Reported anonymously by an r/LocalLLM contributor · score 2

View comment