Two Radeon Instinct MI50 cards (16GB HBM2 each, 32GB total) are power-limited to 145W each and benchmarked with a pre-built Ubuntu Vulkan build of llama.cpp (b11325), with HBM2 clocked at 1000MHz and overclockable to 1200MHz at 1.02 TB/s of bandwidth. Dense models ran slowly (Qwen 3.8 27B at Q6_K around 17.5 to 18 tok/s decode, Gemma 4 31B around 12 to 15.2 tok/s) while MoE models were much faster (Nemotron-3.5-30B-A3B at 60.6 tok/s, Qwen 3.6 35B-A3B around 46.9 to 49.2 tok/s), with prefill from 122 to 983 tok/s. The author uses a mix of dense and MoE GGUF models between 16GB and 32GB total and had no good cooling solution yet, and notes the third MI50 sits beside a Radeon RX 7900 GRE rather than in this dual-card benchmark.
Reported anonymously by an r/LocalLLM contributor · 1 upvotes at capture