CodingGenesis archive2026-09
Qwen 27B on SGLang at NVFP4 with a 180K contextA short but specific configuration note: Qwen 3.8 27B served with SGLang in NVFP4, deliberately capped at a 180K context to stay out of the degradation zone.
Every published setup that reports this runtime, with its machines, models, quantization and reported results. Use the directory filters below to narrow by capacity, model or quantization.
Capacity is the reported total for this setup. Multi-GPU and unified-memory layouts can behave differently.
A short but specific configuration note: Qwen 3.8 27B served with SGLang in NVFP4, deliberately capped at a 180K context to stay out of the degradation zone.
A large enterprise deployment: eight A100 80GBs running Qwen 3.8 under SGLang with tensor parallelism, while the company brings newer B200s online for bigger models.