This person maintains a fork of llama.cpp, llamAmpere, tuned for Ampere-generation Nvidia cards. They report 90+ tok/s through a 100K context and 240K generation with Qwen 3.8 27B, which they describe as much faster than API speed.
They recommend it for work that needs heavy caching, such as scraping, researching and monitoring agents, and suggest pairing it with a specific IQ4_XS quant they provide. They note many of the fixes are not Ampere-specific, but the gain is largest on Ampere, and a 3090 is the natural target.
Reported anonymously by an r/LocalLLM contributor · score 1