Qwen 3.8 Flash Next on an RTX Pro 6000 takes implementation and review tasks delegated from Codex or Claude, replacing Sonnet and Opus for that slice.
Capacity range
Exact reported VRAM
Practice
Model
Quantization
Context length
Reported speed
Runtime
Tools / workflow (harness)
Record origin
EXPLORE BY CAPACITY
Start with your machine.
Capacity is the reported total for this setup. Multi-GPU and unified-memory layouts can behave differently.
8 GB or lessLaptops, iGPUs and modest cards9–12 GBEntry-level local inference13–16 GBCurrent mid-range cards17–20 GBRX 7900 XT and compact pro cards21–24 GB3090, 4090 and 7900 XTX class25–32 GB5090 and workstation cards33–48 GBWorkstation and multi-GPU setups49–96 GBLarge multi-GPU and datacenter setups97 GB+Unified memory, servers and clusters
3 resultsEvery record keeps its origin visible.
CodingGenesis archive2026-09
A workstation GPU as a drop-in subagent for Codex and ClaudeCodingGenesis archive2026-09
A cloud-free workstation: RTX Pro 6000, vLLM and a DeepSeek harnessA custom Linux workstation with an RTX Pro 6000 and 48GB of RAM runs programming, copy editing and summarisation through vLLM, with no cloud use all year.
CodingGenesis archive2026-09
An RTX 6000 Pro and an M4 Max, benchmarked against an active Claude Max planThis person kept paying for Claude Max specifically so they'd have a real comparison point for their own coding hardware — and admits the math still doesn't favor local.