Back to directory
Coding · 2026-09-14

A llama.cpp fork reaching 90+ tok/s through a 100K context on a 3090

A custom llama.cpp fork for Ampere cards pushes Qwen 3.8 27B to 90+ tok/s through a 100K context, aimed at caching-heavy agentic work.

This person maintains a fork of llama.cpp, llamAmpere, tuned for Ampere-generation Nvidia cards. They report 90+ tok/s through a 100K context and 240K generation with Qwen 3.8 27B, which they describe as much faster than API speed.

They recommend it for work that needs heavy caching, such as scraping, researching and monitoring agents, and suggest pairing it with a specific IQ4_XS quant they provide. They note many of the fixes are not Ampere-specific, but the gain is largest on Ampere, and a 3090 is the natural target.

VERIFIABLE SOURCE

Reported anonymously by an r/LocalLLM contributor · score 1

View source
SETUP HISTORY

Snapshots over time

1 version
CURRENT2026-09-14A llama.cpp fork reaching 90+ tok/s through a 100K context on a 3090

A custom llama.cpp fork for Ampere cards pushes Qwen 3.8 27B to 90+ tok/s through a 100K context, aimed at caching-heavy agentic work.

Coding1 machineQwen 3.8 27Bllama.cpp (llamAmpere fork)

This is the currently published snapshot.