A 16GB RTX 3080 Laptop with 32GB RAM and an SSD runs Qwen 3.8 Flash Next 176B (UD-IQ1_M) through TensorSharp, the author's engine that schedules MoE experts across VRAM, system RAM and SSD. Against Strata, decode is close at 11.09 vs 10.24 tok/s, but whole-process time is 16.54s vs 62.15s. TensorSharp peaks at 14,832.5 MiB device-wide GPU and a 19.74 GiB OS working set versus 15,729 MiB and 18.51 GiB for Strata. The machine is a laptop whose RTX 3080 is capped at 95W and throttles when hot.
Reported anonymously by an r/LocalLLM contributor · 1 upvotes at capture
Useful references
Community-provided links related to this setup, workflow or measurements.