Back to directory
Coding · 2026-10-02

Qwen 3.8 27B on a MacBook Pro M4 Max 36GB: context scaling, TTFT and decode through Ollama

Context-scaling benchmark of Qwen 3.8 27B in NVFP4/MLX through Ollama on a 36 GB M4 Max for agentic coding, measuring TTFT, prefill and decode from 4K to 115K.

Tested Qwen 3.8 27B (NVFP4/MLX, 128K configured context) through Ollama 0.34.4 on a MacBook Pro M4 Max with 36 GB unified memory for agentic coding. Warm decode was about 40-41 tok/s with TTFT near 3.5 s and prefill around 180-190 tok/s at small context, but at ~32K active input the wait before generation rose to 3m 29s, and at ~115K it reached 21m 01s with decode down to 14 tok/s, even though 4/4 retrieval markers still passed. The Q4_K_M/GGUF build managed only ~10.7 tok/s against ~40-41 for NVFP4/MLX (a 3.8x gap), and at 96K memory fell to 24% available with 2.9 GB swap. Limits: one machine, one request at a time (OLLAMA_NUM_PARALLEL=1), and large contexts stayed functional but were not pleasant to use.

VERIFIABLE SOURCE

Reported anonymously by an r/LocalLLM contributor · 98 upvotes at capture

View source

Useful references

Community-provided links related to this setup, workflow or measurements.

SETUP HISTORY

Snapshots over time

1 version
CURRENT2026-10-02Qwen 3.8 27B on a MacBook Pro M4 Max 36GB: context scaling, TTFT and decode through Ollama

Context-scaling benchmark of Qwen 3.8 27B in NVFP4/MLX through Ollama on a 36 GB M4 Max for agentic coding, measuring TTFT, prefill and decode from 4K to 115K.

Coding1 machineQwen 3.8 27BOllama

This is the currently published snapshot.