Back to directory
Research & analysis · 2026-10-03

Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop with 32GB RAM and SSD

TensorSharp runs the 176B Qwen 3.8 Flash Next MoE on a 16GB RTX 3080 Laptop with 32GB RAM, matching Strata on decode while finishing the benchmark process far faster.

A 16GB RTX 3080 Laptop with 32GB RAM and an SSD runs Qwen 3.8 Flash Next 176B (UD-IQ1_M) through TensorSharp, the author's engine that schedules MoE experts across VRAM, system RAM and SSD. Against Strata, decode is close at 11.09 vs 10.24 tok/s, but whole-process time is 16.54s vs 62.15s. TensorSharp peaks at 14,832.5 MiB device-wide GPU and a 19.74 GiB OS working set versus 15,729 MiB and 18.51 GiB for Strata. The machine is a laptop whose RTX 3080 is capped at 95W and throttles when hot.

VERIFIABLE SOURCE

Reported anonymously by an r/LocalLLM contributor · 1 upvotes at capture

View source

Useful references

Community-provided links related to this setup, workflow or measurements.

SETUP HISTORY

Snapshots over time

1 version
CURRENT2026-10-03Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop with 32GB RAM and SSD

TensorSharp runs the 176B Qwen 3.8 Flash Next MoE on a 16GB RTX 3080 Laptop with 32GB RAM, matching Strata on decode while finishing the benchmark process far faster.

Research & analysis1 machineQwen 3.8 Flash NextTensorSharp

This is the currently published snapshot.