Inside a large company with more than a hundred 80GB A100 GPUs, this team runs Qwen 3.8 across eight of them using SGLang with tensor parallelism set to 8, primarily for coding. They're now working on deploying larger models on newly arrived B200 GPUs, suggesting the local footprint here is still growing rather than settled.
VERIFIABLE SOURCE
View comment Reported anonymously by an r/LocalLLM contributor · score 2