Back to directory
Coding · 2026-09-13

Agentic coding on a 512GB Mac Studio with oMLX

A 512GB M3 Ultra Mac Studio running models through oMLX handles most agentic coding work, with measured speeds across several models listed by the owner.

This person uses oMLX to run models for agentic coding on a 512GB M3 Ultra Mac Studio. They say it has handled most problems well since Minimax M2.5, and that Qwen 3.8 Flash Next has been as good as cloud AI, just slower. Cloud is only used when in a real hurry or when the Mac Studio is broken by their IT department.

For their workflow, which includes agentic coding of electron charge density analysis software and quantum chemistry post-processing, they report rough single-model speeds: Qwen 3.6 35B-A3B with MTP at 95-105 tok/s, Qwen 3.8 27B at 20 tok/s, Qwen 3.8 Flash Next at 35 tok/s, GLM 5.3 Flash at 20 tok/s and DeepSeek V4 Flash at 25 tok/s. They highlight oMLX's continuous batching over RAM and SSD, which lets them run parallel inference with full cache utilisation.

VERIFIABLE SOURCE

Reported anonymously by an r/LocalLLM contributor · score 2

View source
SETUP HISTORY

Snapshots over time

1 version
CURRENT2026-09-13Agentic coding on a 512GB Mac Studio with oMLX

A 512GB M3 Ultra Mac Studio running models through oMLX handles most agentic coding work, with measured speeds across several models listed by the owner.

Coding1 machineQwen 3.8 Flash Next, Qwen 3.6 35B-A3B, Qwen 3.8 27B, GLM 5.3 Flash, DeepSeek V4 FlashoMLX

This is the currently published snapshot.