Back to directory
Research & analysis · 2026-10-03

Dual Radeon Instinct MI50 benchmarks: dense models at 12 to 18 tok/s, MoE models at 47 to 61 tok/s

Two Radeon Instinct MI50 cards limited to 145W each are benchmarked with llama.cpp Vulkan across dense and MoE models, with MoE models far ahead of dense ones.

Two Radeon Instinct MI50 cards (16GB HBM2 each, 32GB total) are power-limited to 145W each and benchmarked with a pre-built Ubuntu Vulkan build of llama.cpp (b11325), with HBM2 clocked at 1000MHz and overclockable to 1200MHz at 1.02 TB/s of bandwidth. Dense models ran slowly (Qwen 3.8 27B at Q6_K around 17.5 to 18 tok/s decode, Gemma 4 31B around 12 to 15.2 tok/s) while MoE models were much faster (Nemotron-3.5-30B-A3B at 60.6 tok/s, Qwen 3.6 35B-A3B around 46.9 to 49.2 tok/s), with prefill from 122 to 983 tok/s. The author uses a mix of dense and MoE GGUF models between 16GB and 32GB total and had no good cooling solution yet, and notes the third MI50 sits beside a Radeon RX 7900 GRE rather than in this dual-card benchmark.

VERIFIABLE SOURCE

Reported anonymously by an r/LocalLLM contributor · 1 upvotes at capture

View source
SETUP HISTORY

Snapshots over time

1 version
CURRENT2026-10-03Dual Radeon Instinct MI50 benchmarks: dense models at 12 to 18 tok/s, MoE models at 47 to 61 tok/s

Two Radeon Instinct MI50 cards limited to 145W each are benchmarked with llama.cpp Vulkan across dense and MoE models, with MoE models far ahead of dense ones.

Research & analysis1 machineQwen 3.8 27B, Gemma 4 31B, Qwen 3.6 35B-A3B, Nemotron-3.5-30B-A3B, Laguna-XS-2.1, Agents-A1llama.cpp

This is the currently published snapshot.