On a 19-task cybersecurity benchmark, the uncensored Qwen 3.8 Flash Next solved 80.6% of tasks on the first try, ahead of MiMo 2.6 Flash at 73.7%, GPT-6 Luna at 65.6%, plain Qwen 3.8 Flash Next at 69.4%, GLM 5.3 Flash at 69.4% and uncensored GLM 5.3 Flash at 50%, while the author's previous local test, Qwen3.8 27B, scored 28%. Each task gives the model a shell in an isolated Docker box and asks it to break into a target and pull out a flag, across binary exploitation, web, crypto, reverse engineering and multi-stage network ranges. Uncensoring the same Qwen model more than doubled its binary exploitation score from 25% to 58% and the model got faster and used fewer tokens, but on GLM 5.3 Flash abliteration dropped the score from 69% to 50% and almost doubled token use. Frontier models (Opus, Sonnet, Sol, Astra) were blocked by safety classifiers on all 19 tasks and could not be benchmarked, and the large GLM 5.3 was not run because it would have cost more than the author could spend.
Reported anonymously by an r/LocalLLM contributor · 1 upvotes at capture
Useful references
Community-provided links related to this setup, workflow or measurements.