Benchmarking Strata n-gram table placement on a modded RTX 4090 48GB with 128GB DDR5
A measured benchmark on a modded 48GB RTX 4090 shows Strata's n-gram table belongs on SSD while conversation parking delivers the large speedup.
On a modded RTX 4090 with 48GB VRAM, 128GB DDR5-6000 and a 262K Strata config, holding the 28.8GB n-gram table entirely in RAM gained only +0.65% prefill and +1.2% decode for +28.4 GiB RAM, while the default 90MB row cache already absorbed 92% of reusable traffic. Conversation parking cut the return to a 91,836-token conversation from 18.8s to 537ms for 3.1GB of RAM, about 35x, and preview auto 32768 gave +9.7% prefill and +13.6% decode at a 78.7K prompt. The calibrate pass changed only pool workers from 23 to 12 for +3.2%. The author reports an 11.5% session-to-session spread and did not test output quality.
VERIFIABLE SOURCE
Reported anonymously by an r/LocalLLM contributor · 1 upvotes at capture