A living register of what is worth watching in local AI: models, runtimes, hardware, sovereignty, security and cost. This is an external watch list, not a claim about the catalog. Every item is dated and sourced; nothing here is a recommendation without evidence. Last updated 2026-10-04.
Models & quants
What is actually good to run locally right now: sparse MoE, long context, uncensored variants and the quantization families that make big models fit.
Qwen 3.8 Flash NextHot
The current reference for big-MoE-on-small-hardware: it fits in RAM by offload and holds up at long context, which is why it dominates the recent setups. The ISTA GSQ-RCO quants (Q2_0 for speed, IQ3_XXS for quality) are the ones people actually run.
Multiple independent 1.18.0 entries: 44 tok/s on an M1 Max at 398K, ~60 tok/s on a 7900XTX at 250k, 11-15 tok/s streamed from SSD on a 12GB RTX 5070.
Abliteration is not a free quality win: it can double the hardest-task score on one model and halve it on another. Watch it as a variable, not a category.
lbgos cyber benchmark: uncensored Qwen 3.8 Flash Next 80.6% vs 69.4% plain (binary exploitation 25% -> 58%), but uncensored GLM 5.3 Flash dropped 69% -> 50% with ~2x tokens.
Long context is a memory problem, not a context-window numberWatch
Prefill and TTFT degrade long before the advertised window is reached on a single consumer machine; the useful metric is time-to-first-token at real project sizes.
M4 Max / Qwen 3.8 27B: 40 tok/s at small context, but ~3m30s to first token at 32K and ~21m at 115K.
The engines that decide whether a big model fits, and how fast prompt processing is. Classic llama.cpp is no longer the only option.
llama.cpp + expert streamingStable
Still the default, now pushed far beyond VRAM by streaming experts from SSD with a page-locked hot tier. The bottleneck moved to Windows I/O queue depth.
177B Qwen 3.8 Flash Next at ~11.5 tok/s on 12GB VRAM + 32GB DDR4-2400 via a llama.cpp streaming fork.
A purpose-built engine for one model family that keeps conversation state parked and re-reads it cheaply; it is showing up across cards (5070, 7900XTX, 4090, V100, 3090) with strong long-context decode.
35x faster return to a 91,836-token conversation vs reload; ~60 tok/s at 250k on a 7900XTX.
The M-series path keeps improving through unofficial forks that graft new architectures before upstream catches up; expect to pin a fork to get long-context kernels on M1/M2.
A Splash/ds4 fork holds 35 tok/s at 398K on a 2021 M1 Max.
128GB on four consumer cards is now enough for a 50-seat office doing bursty RAG, chat, diffusion and light agentic work, with the caveat that seat count is not concurrency.
4x5090 / 192GB ECC / 7.6kW-class build; throughput figures came from their 2x5090 box (24 concurrent at 52 tok/s short context).
Why organisations go local at all: keeping data off a third party. The strongest material for an Intelligentiae article.
The driver is data residency, not cost or rate limitsWatch
Companies move local to keep data off-grid; the workload is deliberately bursty and cross-department (HR, marketing, finance), not a room of engineers.
A 50-person office chose a local 4x5090 rig for data privacy across departments, not for coding throughput.
For regulated and security workloads, the cloud models often refuse the task entirely, which is itself a sovereignty argument: you cannot run a workflow a vendor can revoke.
Opus, Sonnet, Sol and Astra were blocked on all 19 cybersecurity tasks, removing them from the benchmark.
How to actually measure a model's offensive capability, and how not to confuse that with testing your own application.
rangebench: a framed CTF harnessWatch
19 tasks across binary exploitation, web, crypto, reverse engineering and multi-stage networks; the model gets a shell in an isolated Docker box and must break in and pull a flag. Tasks are private so they do not leak into training data.
Uncensored Qwen 3.8 Flash Next 80.6%, MiMo 2.6 Flash 73.7%, GPT-6 Luna 65.6%; 3 tasks unsolved by every model; 2 runs per task.
This is a model-capability bench, not an app-security testNote
Important distinction for us: rangebench measures whether a model can perform offensive tasks, not whether an application you built is safe. Testing your own app needs a different toolchain (dependency/SAST/DAST scanning, secret detection, a scoped red-team), even if the same local model assists it.
The benchmark objective is breaking into a purpose-built target and exfiltrating a flag, on tasks the author wrote and keeps private.
When local wins and when it does not, with real numbers instead of slogans.
Compared against paid tokens, not against a subscriptionNote
Articles for the agency should frame local as a TCO question (hardware amortisation, power, ops time) against metered API use for the same workload.
Cloud-orchestrator plus local-worker example: 43.4 min and $0.17 for three small builds, vs $0.75 for the cloud model alone and 114.2 min for the local card alone (power excluded).