Full directory
DISCOVER BY RUNTIME

Setups running
vLLM.

Every published setup that reports this runtime, with its machines, models, quantization and reported results. Use the directory filters below to narrow by capacity, model or quantization.

5documented uses
with this runtime
Sort by
Compare
Capacity range
Exact reported VRAM
Practice
Model
Quantization
Context length
Reported speed
Runtime
Tools / workflow (harness)
Record origin
EXPLORE BY CAPACITY

Start with your machine.

Capacity is the reported total for this setup. Multi-GPU and unified-memory layouts can behave differently.

5 resultsEvery record keeps its origin visible.
Hybrid orchestrationCommunity snapshot2026-09
27B reasoning model on 2x RTX 2080 Ti (NVFP4, vLLM) for agentic orchestration

Qwen3.8-27B (NVFP4 W4A16 + FP8 KV) served on 2x RTX 2080 Ti 22GB over NVLink via a vLLM TP2 fork. 196K context cap, MTP K=3, prefix caching on. Feeds a 1-orchestrator + 7-worker harness: ~1227 tok/s prefill and ~29 tok/s decode solo at 65K, ~192 tok/s aggregate at 8-way. Rejected AWQ/FP8 weights (72% EOS failure), ExLlamaV3 (3-6x slower), and PP=2 (infeasible on 2 GPUs).

33–48 GBunsloth_Qwen3.8-27B-NVFP4vLLMRTX 2080 Ti