Back to directory
Regulated work · 2026-09-13

Unit tests and reviews for NDA-bound client code on a single 3090

A contractor whose clients forbid commercial providers from reading their source code runs unit tests and code reviews locally on a 3090, with a long-context llama.cpp setup.

This person's customers do not allow their source code to be read by commercial LLM providers, so the work runs locally on an RTX 3090 with 96GB of RAM. The main uses are generating unit tests and performing code reviews.

Their working configuration is Qwen 3.8 27B GGUF on llama.cpp with OpenCode, using a 163840-token context and q4_0 quantisation for both the K and V caches, which they found gave enough performance for that context size. They were surprised at how well it works.

VERIFIABLE SOURCE

Reported anonymously by an r/LocalLLM contributor · score 2

View source
SETUP HISTORY

Snapshots over time

1 version
CURRENT2026-09-13Unit tests and reviews for NDA-bound client code on a single 3090

A contractor whose clients forbid commercial providers from reading their source code runs unit tests and code reviews locally on a 3090, with a long-context llama.cpp setup.

Regulated work1 machineQwen 3.8 27Bllama.cpp

This is the currently published snapshot.