This person has more than 20 years of software engineering experience and works daily in C, Rust, Python, TypeScript and Go. On a single RTX 5090 they run three concurrent sessions averaging about 150 tok/s each.
They have not used cloud models for the past three months, saying local models do everything they need and more.
Reported anonymously by an r/LocalLLM contributor · score 3