Kyojin is a new engine built on ExLlamaV3 for Strix Halo (gfx1151, ROCm) that runs two 300B-class MoE models on a single Ryzen AI Max+ 395 mini PC with 128 GB unified memory. GLM 5.3 Flash (99.7 GB, mixed ~2.5 bpw EXL3) prefills at 580 tok/s at 3.5K and 546 at 64K, decoding 26-30 tok/s with MTP; MiMo-V2.6-Flash (105 GB) prefills about 650 tok/s at 4K and decodes 32 tok/s on prose, 35 on chat and 44 on code with speculative decoding, 29 plain. Reported KLD against official FP8 is 0.151 for GLM and 0.0713 for MiMo, with 89.3% and 92.0% top-1 agreement respectively. Limits: task-suite scores for MiMo, GLM at 128K context, and any GPU other than gfx1151 were not measured, and the author says he had no real coding benchmark for either model.
Reported anonymously by an r/LocalLLM contributor · 1 upvotes at capture
Useful references
Community-provided links related to this setup, workflow or measurements.