This person uses oMLX to run models for agentic coding on a 512GB M3 Ultra Mac Studio. They say it has handled most problems well since Minimax M2.5, and that Qwen 3.8 Flash Next has been as good as cloud AI, just slower. Cloud is only used when in a real hurry or when the Mac Studio is broken by their IT department.
For their workflow, which includes agentic coding of electron charge density analysis software and quantum chemistry post-processing, they report rough single-model speeds: Qwen 3.6 35B-A3B with MTP at 95-105 tok/s, Qwen 3.8 27B at 20 tok/s, Qwen 3.8 Flash Next at 35 tok/s, GLM 5.3 Flash at 20 tok/s and DeepSeek V4 Flash at 25 tok/s. They highlight oMLX's continuous batching over RAM and SSD, which lets them run parallel inference with full cache utilisation.
Reported anonymously by an r/LocalLLM contributor · score 2