Back to directory
Research & analysis · 2026-10-03

GLM 5.3 Flash and MiMo-V2.6-Flash on one 128 GB Strix Halo mini PC with the Kyojin engine

Release benchmark of an ExLlamaV3-based ROCm engine running two 300B-class MoE models, each on a single 128 GB Ryzen AI Max+ 395 mini PC.

Kyojin is a new engine built on ExLlamaV3 for Strix Halo (gfx1151, ROCm) that runs two 300B-class MoE models on a single Ryzen AI Max+ 395 mini PC with 128 GB unified memory. GLM 5.3 Flash (99.7 GB, mixed ~2.5 bpw EXL3) prefills at 580 tok/s at 3.5K and 546 at 64K, decoding 26-30 tok/s with MTP; MiMo-V2.6-Flash (105 GB) prefills about 650 tok/s at 4K and decodes 32 tok/s on prose, 35 on chat and 44 on code with speculative decoding, 29 plain. Reported KLD against official FP8 is 0.151 for GLM and 0.0713 for MiMo, with 89.3% and 92.0% top-1 agreement respectively. Limits: task-suite scores for MiMo, GLM at 128K context, and any GPU other than gfx1151 were not measured, and the author says he had no real coding benchmark for either model.

VERIFIABLE SOURCE

Reported anonymously by an r/LocalLLM contributor · 1 upvotes at capture

View source

Useful references

Community-provided links related to this setup, workflow or measurements.

SETUP HISTORY

Snapshots over time

1 version
CURRENT2026-10-03GLM 5.3 Flash and MiMo-V2.6-Flash on one 128 GB Strix Halo mini PC with the Kyojin engine

Release benchmark of an ExLlamaV3-based ROCm engine running two 300B-class MoE models, each on a single 128 GB Ryzen AI Max+ 395 mini PC.

Research & analysis1 machineGLM 5.3 Flash, MiMo-V2.6-FlashKyojin

This is the currently published snapshot.