Full directory
DISCOVER BY GPU

Setups running on
RTX 2080 Ti.

Every published setup that reports this GPU, with its models, quantization, runtime and reported results. Use the directory filters below to narrow by capacity, model or runtime.

1documented use
on this GPU
Sort by
Compare
Capacity range
Exact reported VRAM
Practice
Model
Quantization
Context length
Reported speed
Runtime
Tools / workflow (harness)
Record origin
EXPLORE BY CAPACITY

Start with your machine.

Capacity is the reported total for this setup. Multi-GPU and unified-memory layouts can behave differently.

1 resultsEvery record keeps its origin visible.
Hybrid orchestrationCommunity snapshot2026-09
27B reasoning model on 2x RTX 2080 Ti (NVFP4, vLLM) for agentic orchestration

Qwen3.8-27B (NVFP4 W4A16 + FP8 KV) served on 2x RTX 2080 Ti 22GB over NVLink via a vLLM TP2 fork. 196K context cap, MTP K=3, prefix caching on. Feeds a 1-orchestrator + 7-worker harness: ~1227 tok/s prefill and ~29 tok/s decode solo at 65K, ~192 tok/s aggregate at 8-way. Rejected AWQ/FP8 weights (72% EOS failure), ExLlamaV3 (3-6x slower), and PP=2 (infeasible on 2 GPUs).

33–48 GBunsloth_Qwen3.8-27B-NVFP4vLLMRTX 2080 Ti