Back to directory
Regulated work · 2026-10-04

Uncensored Qwen 3.8 Flash Next tops a 19-task hacking benchmark at 80.6% on 2x RTX Pro 6000

On a 19-task cybersecurity benchmark, an uncensored Qwen 3.8 Flash Next solved 80.6% first try, but the run used someone else's 2x RTX Pro 6000 and only 2 runs per task.

On a 19-task cybersecurity benchmark, the uncensored Qwen 3.8 Flash Next solved 80.6% of tasks on the first try, ahead of MiMo 2.6 Flash at 73.7%, GPT-6 Luna at 65.6%, plain Qwen 3.8 Flash Next at 69.4%, GLM 5.3 Flash at 69.4% and uncensored GLM 5.3 Flash at 50%, while the author's previous local test, Qwen3.8 27B, scored 28%. Each task gives the model a shell in an isolated Docker box and asks it to break into a target and pull out a flag, across binary exploitation, web, crypto, reverse engineering and multi-stage network ranges. Uncensoring the same Qwen model more than doubled its binary exploitation score from 25% to 58% and the model got faster and used fewer tokens, but on GLM 5.3 Flash abliteration dropped the score from 69% to 50% and almost doubled token use. Frontier models (Opus, Sonnet, Sol, Astra) were blocked by safety classifiers on all 19 tasks and could not be benchmarked, and the large GLM 5.3 was not run because it would have cost more than the author could spend.

VERIFIABLE SOURCE

Reported anonymously by an r/LocalLLM contributor · 1 upvotes at capture

View source

Useful references

Community-provided links related to this setup, workflow or measurements.

SETUP HISTORY

Snapshots over time

1 version
CURRENT2026-10-04Uncensored Qwen 3.8 Flash Next tops a 19-task hacking benchmark at 80.6% on 2x RTX Pro 6000

On a 19-task cybersecurity benchmark, an uncensored Qwen 3.8 Flash Next solved 80.6% first try, but the run used someone else's 2x RTX Pro 6000 and only 2 runs per task.

Regulated work1 machineQwen 3.8 Flash Next UncensoredRuntime unspecified

This is the currently published snapshot.