Best Uncensored LLMs for Security Work: The 2026 Field Guide
Oct 4, 2026
Which uncensored and abliterated models make sense for security work in 2026: the 27B abliterations, the GLM-5.3 and DeepSeek-V4.1 frontier, the quality costs each one publishes, and how to measure refusal on your own prompts before choosing.
Security teams looking for an uncensored model in 2026 face a catalog problem, not a scarcity problem. Hugging Face carried roughly 7,970 models with "abliterated" in the name as of September 2026, by Featherless's count. Most of them are weekend ablations of whatever base model shipped last month, published without measurements. This guide narrows the field to what matters for security work: the categories that exist, the models that publish their costs, and the evaluation you should run before trusting any of them.
The companion piece, Attackers Have Uncensored LLMs. Your Security Models Refuse to Help., covers why refusal-free models matter on the defense side. This one is the shopping list.
A note on provenance: we have not benchmarked the models below in our own lab. Every number here is published by the model's author, the paper that measured it, or the platform serving it, and each is linked. Treat them as screening data, and run the refusal measurement in the last section before you commit.
Four categories, four different jobs
Abliterated general models
The default choice. A strong open base model with the refusal direction edited out of its weights, nothing else changed. The current generation to look at:
| Model | What it is | Published cost |
|---|---|---|
| Huihui-Qwen3.5-27B-abliterated | Qwen3.5 27B, Apache 2.0, the default general uncensored pick since February 2026 | not published; layers 18–51 ablated |
| gemma-4-26B-A4B-it-uncensored | Gemma 4 MoE, ~4B active per token; author measured 0.7% refusal across four test sets with quality effectively unchanged | 0.7% residual refusal |
| Qwen3.8-27B-OBLITERATED | Qwen3.8 27B, Apache 2.0, iterative SVD + LEACE passes; card claims zero refusals and zero soft deflections | 2.12 points of MMLU (82.33% vs 84.46% stock) |
When a card publishes no cost number, assume there is one. The 27B Qwen line matters because three independent ablations of the same base exist, which lets you compare methods rather than models: on the same weights, a published 2.12-point MMLU delta is a real price tag.
Security-tuned uncensored models
A new category in 2026: models post-trained specifically for offensive-security and red-team work. Qwythos-9B-Claude-Mythos-5-1M, released June 2026 on a Qwen3.5 base, is the current example: trained on security research and red-teaming traces, with the author's matched-condition evals reporting +34 MMLU and +30 GSM8K-strict over the base. Author-reported, single-source, but the category exists now, and it exists because buyers asked for it.
Aligned security models
For work you will have to defend to a compliance team, refusal removal is the wrong tool. Cisco's Foundation-Sec-8B pair (Instruct and Reasoning variants, 8B on Llama 3.1) is aligned, purpose-built for threat-intelligence and SOC work, and publishes security-benchmark scores: 0.691 CTI-MCQA, 0.753 CTI-RCM, 0.856 CTI-VSP for the Reasoning variant. Detection, triage and classification tasks are where aligned models hold their own; the July 2026 same-lineage study found them marginally ahead on neutral-phrased diagnostics.
The 2026 frontier: GLM-5.3 and DeepSeek V4.1
Refusal-free does not mean capable, and the capability frontier moved in late 2026. Two open-weight families now matter for security work beyond the 27B abliterations above, and both ship under MIT.
GLM-5.3-Flash is the practical one: roughly 321B total parameters with 18B active per token, natively multimodal, a declared context window over one million tokens, and 5.6 million downloads a month. Z.ai's own evals put it at 63.4 on DeepSWE and 84.3 on Terminal-Bench 2.1.
It is also the model we know best firsthand. One NVFP4 replica runs on two RTX PRO 6000 GPUs (192 GB) in our production stack, serving 388 output tokens per second with four concurrent requests; the TP2 recipe is published. The uncensored supply followed within weeks: huihui-ai's Huihui-GLM-5.3-Flash-abliterated GGUF is the most-downloaded abliterated build in the family, with Heretic V2 quants close behind.
The full GLM-5.3 (753B) is the capability sibling. DeepSeek's own comparison table scores it at 88.2 on Terminal-Bench 2.1, 84.5 on CyberGym and 15.0 on ExploitGym, competitive on security work with the closed models in the same table. It is a datacenter model: 755.7 GB of FP8 weights, one 8×H200 server in our serving setup, where it carries 48 concurrent coding agents at about 108 turns a minute. Abliterated FP8 and NVFP4 builds of it exist on Hugging Face; few teams have the hardware to run them.
DeepSeek-V4.1-Flash is the newest DeepSeek and, on paper, the strongest security model in the open-weight field. Its own model card reports 88.1 on CyberGym — ahead of GPT-5.6 Sol at 84.5 and GLM-5.3 at 84.5 in the same table — plus 62.8 on SEC-Bench Pro and 15.3 on ExploitGym. The architecture is unusual: 552B backbone plus 196B of sparse Engram memory, with only 8B parameters active while reading a prompt and 16B while generating.
Abliterated GGUFs appeared within weeks; audreyt's passed 100,000 downloads. Budget for the serving engineering, not just the weights: V4.1's causal encoder-decoder split needs a V4.1-aware stack, and a conventional OpenAI-compatible server cannot run it by renaming an older DeepSeek layout. We inspected the release and have not deployed it ourselves for exactly that reason.
One more signal that this intersection is real: security researchers are already stacking cyber-focused LoRAs on top of abliterated frontier weights, for example GLM-5.3-abliterated-cyber and a DeepSeek-V4.1-Flash cyber LoRA.
For hard exploitation reasoning, model choice still matters more than alignment status. The last public CTF-specific ordering, the NYU CTF Bench run from April 2026, put 200 challenges in front of the field. Open-weight GLM-5 solved 19.5%, ahead of GPT-5.2-Codex at 18.0%, with DeepSeek-V3 at 6.5%, Qwen 3.5 397B at 3.5% and Llama 3.3 70B at 2.5%. Claude 4.5 Opus led at 59.0%. Those numbers predate both frontier families above; treat them as the floor the frontier has moved past, not the current ordering.
What the leaderboard tells you
The query "uncensored llm leaderboard" usually means the UGI Leaderboard, which ranks uncensored and abliterated models on two axes: how completely refusals were removed, and how much general intelligence survived the edit. Read it as a survival filter, a way to discard models that traded away their reasoning, not as a security capability ranking. A model can top UGI and still fail your workload, because UGI does not measure security tasks.
Match the failure to the fix
The July 2026 same-lineage study is the cleanest public evidence of what abliteration changes. Aligned and abliterated versions of the same weights on the Vul4J Java benchmark:
| Vul4J stage | Aligned | Abliterated |
|---|---|---|
| Vulnerability detection | 58.42% | 57.22% |
| Usable patch produced | 29.94% | 67.80% |
| Patch applied cleanly | 24.86% | 64.97% |
| Patch compiled | 9.04% | 32.77% |
Diagnosis did not move. Artefact generation tripled. The decision rule follows from the data: if your bottleneck is classification and triage, an aligned model is fine. If your bottleneck is that models describe fixes instead of writing them, you need the refusal edit. Pay attention to which method was used, because the collateral damage varies sixfold between recipes with identical refusal outcomes. In the Heretic comparison on Gemma-3-12B-IT, KL divergence from the original on neutral inputs ranged from 0.16 to 1.04.
How to measure refusal before you choose
Published refusal rates are averages over someone else's prompts. The Scale AI Defensive Refusal Bias study found rates from 0% on log analysis to 43.8% on system hardening, which means your mix of tasks decides whether refusals are even your problem.
- Pull 50–100 real requests from your queue: the hardening changes, malware questions, exploit verifications your team actually types.
- Run them against two or three candidates, same prompts, same sampling settings.
- Count hard refusals and hedged non-answers separately. A model that "cannot help with that" and a model that writes three paragraphs describing what a patch would do are failing differently.
- If your refusal count is near zero, alignment is moot for your workload: choose on capability and licence, and skip the uncensored detour entirely.
Running one on your own metal
For security work the hosting question is settled before the model question: live malware, client code and exploit PoCs do not go to a third-party API. The stack that runs these models locally is the standard open-source one. Use Ollama or llama.cpp for single-GPU work, vLLM when you need throughput for agents, and GGUF quantizations of any of the models above for smaller GPUs. Quantization costs quality too, so budget for the same kind of before/after check you would run for the ablation itself.
The frontier class needs real hardware, and the honest brackets look like this. A 27B abliteration runs on one workstation GPU. GLM-5.3-Flash in NVFP4 spans two 96 GB GPUs per replica in our setup, which is the cheapest serious entry point into the frontier today. Full GLM-5.3 is a 755.7 GB FP8 checkpoint that wants an 8×H200 server. DeepSeek-V4.1-Flash adds the serving-stack problem on top: its encoder-decoder split needs a V4.1-aware runtime, so check your inference stack's support before you download the full 763B checkpoint, backbone, Engram and vision encoder included.
Licence-wise the frontier is friendlier than the 27B class. GLM-5.3-Flash and DeepSeek-V4.1-Flash both ship under MIT, the most permissive option in this guide. Qwen and Gemma ablations under Apache 2.0 can also be shipped inside a product. Llama-derived weights carry the Llama Community License, whose acceptable-use and redistribution terms apply; that includes the Foundation-Sec models, whose Cisco changes are Apache 2.0 but whose base is Llama 3.1.
Keep the human in the loop
An abliterated model has lost its ability to decline, including for the requests it should have declined. It will draft the phishing email as readily as the phishing detection rules. Keep a person at every irreversible gate: scope approval, anything that touches production, anything that ships to a client. Uncensored shifts the judgement from the model to your process, and that is a real operational cost, priced in review time, that you should plan for rather than discover.
The models are the tooling; the work is the point. If what you actually want is security agents that complete offensive-security work on your infrastructure, with findings validated before they reach a report, see Helix Cyber. The related guides cover AI penetration testing and LLM red teaming for the attack surface your own AI applications add.