Full directory
STATE OF THE ART

What is moving
right now.

A living register of what is worth watching in local AI: models, runtimes, hardware, sovereignty, security and cost. This is an external watch list, not a claim about the catalog. Every item is dated and sourced; nothing here is a recommendation without evidence. Last updated 2026-10-04.

Models & quants

What is actually good to run locally right now: sparse MoE, long context, uncensored variants and the quantization families that make big models fit.

Qwen 3.8 Flash NextHot

The current reference for big-MoE-on-small-hardware: it fits in RAM by offload and holds up at long context, which is why it dominates the recent setups. The ISTA GSQ-RCO quants (Q2_0 for speed, IQ3_XXS for quality) are the ones people actually run.

Multiple independent 1.18.0 entries: 44 tok/s on an M1 Max at 398K, ~60 tok/s on a 7900XTX at 250k, 11-15 tok/s streamed from SSD on a 12GB RTX 5070.

Source
Uncensored / abliterated variantsWatch

Abliteration is not a free quality win: it can double the hardest-task score on one model and halve it on another. Watch it as a variable, not a category.

lbgos cyber benchmark: uncensored Qwen 3.8 Flash Next 80.6% vs 69.4% plain (binary exploitation 25% -> 58%), but uncensored GLM 5.3 Flash dropped 69% -> 50% with ~2x tokens.

Source
GLM 5.3 Flash and MiMo 2.6 FlashWatch

300B-class sparse MoE squeezed into one 128GB machine at ~2.5 bpw EXL3. Low bit-width, but top-1 agreement with FP8 stays around 90%.

Kyojin on Strix Halo: GLM 5.3 Flash 26-30 tok/s, MiMo 32-44 tok/s decode; KLD 0.151 and 0.0713 vs FP8.

Source
Long context is a memory problem, not a context-window numberWatch

Prefill and TTFT degrade long before the advertised window is reached on a single consumer machine; the useful metric is time-to-first-token at real project sizes.

M4 Max / Qwen 3.8 27B: 40 tok/s at small context, but ~3m30s to first token at 32K and ~21m at 115K.

Source
Runtimes & engines

The engines that decide whether a big model fits, and how fast prompt processing is. Classic llama.cpp is no longer the only option.

llama.cpp + expert streamingStable

Still the default, now pushed far beyond VRAM by streaming experts from SSD with a page-locked hot tier. The bottleneck moved to Windows I/O queue depth.

177B Qwen 3.8 Flash Next at ~11.5 tok/s on 12GB VRAM + 32GB DDR4-2400 via a llama.cpp streaming fork.

Source
StrataHot

A purpose-built engine for one model family that keeps conversation state parked and re-reads it cheaply; it is showing up across cards (5070, 7900XTX, 4090, V100, 3090) with strong long-context decode.

35x faster return to a 91,836-token conversation vs reload; ~60 tok/s at 250k on a 7900XTX.

Source
Kyojin (ExLlamaV3 on ROCm) and TensorSharpWatch

Newer entrants betting on memory hierarchy: MoE-aware scheduling of experts across VRAM, RAM and SSD, and ROCm tuning for Strix Halo.

TensorSharp 16.54s vs Strata 62.15s whole-process time on the same 176B run at near-equal decode.

Source
MLX and the Apple pathWatch

The M-series path keeps improving through unofficial forks that graft new architectures before upstream catches up; expect to pin a fork to get long-context kernels on M1/M2.

A Splash/ds4 fork holds 35 tok/s at 398K on a 2021 M1 Max.

Source
Hardware

Where VRAM per euro actually lands: unified-memory mini PCs, cheap datacenter cards, and the enthusiast multi-GPU towers.

Strix Halo (Ryzen AI Max+ 395, 128GB unified)Hot

One small box runs two 300B-class models. The bet is unified memory and bandwidth, not raw FLOPS.

GLM 5.3 Flash and MiMo 2.6 Flash each fit one 128GB machine.

Source
Cheap 32GB datacenter cards (Radeon Instinct MI50)Watch

The cost floor for a working MoE rig is falling fast; a 16GB MI50 is under $150 used, and two of them run 30B-A3B at ~60 tok/s.

Dual MI50 benchmark: Nemotron-3.5-30B-A3B at 60.6 tok/s, Qwen 3.6 35B-A3B at ~47-49 tok/s.

Source
4x RTX 5090 for a teamWatch

128GB on four consumer cards is now enough for a 50-seat office doing bursty RAG, chat, diffusion and light agentic work, with the caveat that seat count is not concurrency.

4x5090 / 192GB ECC / 7.6kW-class build; throughput figures came from their 2x5090 box (24 concurrent at 52 tok/s short context).

Source
Modded high-VRAM consumer cardsWatch

48GB 4090s and similar open whole workflows (262K context, 85% of experts in VRAM) that are impossible on a stock card.

Strata on a 48GB 4090: decode 128-151 tok/s, prefill 4,249-5,154 tok/s at 78K-92K context.

Source
Sovereignty & compliance

Why organisations go local at all: keeping data off a third party. The strongest material for an Intelligentiae article.

The driver is data residency, not cost or rate limitsWatch

Companies move local to keep data off-grid; the workload is deliberately bursty and cross-department (HR, marketing, finance), not a room of engineers.

A 50-person office chose a local 4x5090 rig for data privacy across departments, not for coding throughput.

Source
Frontier safety classifiers block legitimate defensive workWatch

For regulated and security workloads, the cloud models often refuse the task entirely, which is itself a sovereignty argument: you cannot run a workflow a vendor can revoke.

Opus, Sonnet, Sol and Astra were blocked on all 19 cybersecurity tasks, removing them from the benchmark.

Source
Self-hosted tooling over heavy vector databasesWatch

Sovereign RAG does not need a managed vector store; lightweight SQLite-backed memory is enough for internal document privacy.

A SQLite-based memory layer for Ollama gives persistent document memory without a heavy vector DB.

Source
Security & offensive AI

How to actually measure a model's offensive capability, and how not to confuse that with testing your own application.

rangebench: a framed CTF harnessWatch

19 tasks across binary exploitation, web, crypto, reverse engineering and multi-stage networks; the model gets a shell in an isolated Docker box and must break in and pull a flag. Tasks are private so they do not leak into training data.

Uncensored Qwen 3.8 Flash Next 80.6%, MiMo 2.6 Flash 73.7%, GPT-6 Luna 65.6%; 3 tasks unsolved by every model; 2 runs per task.

Source
This is a model-capability bench, not an app-security testNote

Important distinction for us: rangebench measures whether a model can perform offensive tasks, not whether an application you built is safe. Testing your own app needs a different toolchain (dependency/SAST/DAST scanning, secret detection, a scoped red-team), even if the same local model assists it.

The benchmark objective is breaking into a purpose-built target and exfiltrating a flag, on tasks the author wrote and keeps private.

Source
Guardrail posture is a deployment variableWatch

Choosing a censored or uncensored model changes what a local workflow can do; that choice belongs in the record, not hidden.

Uncensoring helped Qwen on binary exploitation but hurt GLM, so the effect is model-specific, not a general rule.

Source
Cost & ROI

When local wins and when it does not, with real numbers instead of slogans.

Compared against paid tokens, not against a subscriptionNote

Articles for the agency should frame local as a TCO question (hardware amortisation, power, ops time) against metered API use for the same workload.

Cloud-orchestrator plus local-worker example: 43.4 min and $0.17 for three small builds, vs $0.75 for the cloud model alone and 114.2 min for the local card alone (power excluded).

Source
The real cost of a big local model is RAM and SSD, not the GPUWatch

A 12GB card plus cheap DDR4 and an SSD already runs a 177B MoE; the GPU tier matters less than the memory hierarchy for capacity-bound work.

~11.5 tok/s for a 177B model on 12GB VRAM + 32GB DDR4-2400 via SSD expert streaming.

Source

How to read this

Dated and sourced. Every item rests on a link and a specific result, never a slogan.

A watch list, not a verdict. “Hot” means moving now, not “the best”.

Separate from the catalog. The catalog records what contributors actually run; this watches the wider field.

Editable in the open. The register is a versioned file; corrections are welcome.