A Tesla V100 32GB and an RTX 3090 on Strata: Qwen 3.8 Flash Next at 262K context for local business management
Two single-GPU Ubuntu servers, one with a blower-modded Tesla V100 32GB and one with an RTX 3090, run Strata with Qwen 3.8 Flash Next at 262K context for local business management automation.
Two single-GPU Ubuntu VMs each run Strata at 262K context: an RTX 3090 capped at 220W in a Dell R740 runs Qwen 3.8 Flash Next at IQ3_S, and a blower-modded Tesla V100 32GB capped at 180W in an EPYC server runs an uncensored Qwen 3.8 Flash Next at IQ4_XS. The author reports 60 to 65 tok/s decode and 1200 to 1300 tok/s prefill on the 3090, and 50 to 55 tok/s decode with 850 to 900 tok/s prefill on the V100, dropping to roughly 50 to 55 and 40 to 45 tok/s respectively on general tasks through Hermes Agent. The workload is local business management: database and memory vaults, orchestrating processes and software, and Computer Use automation, where Flash Next is preferred over Qwen 3.8 27B because it completes tasks around 5x faster at xhigh despite a slightly lower raw decode rate. An IQ3_XXS test on the V100 matched 3090 speed but triggered frequent reasoning loops, and the author notes the V100 needs Ubuntu and Strata modifications for Volta.
VERIFIABLE SOURCE
Reported anonymously by an r/LocalLLM contributor · 1 upvotes at capture
CURRENT2026-10-04A Tesla V100 32GB and an RTX 3090 on Strata: Qwen 3.8 Flash Next at 262K context for local business management
Two single-GPU Ubuntu servers, one with a blower-modded Tesla V100 32GB and one with an RTX 3090, run Strata with Qwen 3.8 Flash Next at 262K context for local business management automation.
Business automation2 machinesQwen 3.8 Flash Next, Qwen 3.8 27BStrata