Tested Qwen 3.8 27B (NVFP4/MLX, 128K configured context) through Ollama 0.34.4 on a MacBook Pro M4 Max with 36 GB unified memory for agentic coding. Warm decode was about 40-41 tok/s with TTFT near 3.5 s and prefill around 180-190 tok/s at small context, but at ~32K active input the wait before generation rose to 3m 29s, and at ~115K it reached 21m 01s with decode down to 14 tok/s, even though 4/4 retrieval markers still passed. The Q4_K_M/GGUF build managed only ~10.7 tok/s against ~40-41 for NVFP4/MLX (a 3.8x gap), and at 96K memory fell to 24% available with 2.9 GB swap. Limits: one machine, one request at a time (OLLAMA_NUM_PARALLEL=1), and large contexts stayed functional but were not pleasant to use.
Reported anonymously by an r/LocalLLM contributor · 98 upvotes at capture
Useful references
Community-provided links related to this setup, workflow or measurements.