Attackers Have Uncensored LLMs. Your Security Models Refuse to Help.
Oct 4, 2026
Refusal-free open-weight models are a working tool on one side of security: attackers download them without an account, while aligned models refuse 43.8% of fully defensive hardening prompts. What the asymmetry costs, measured.
An attacker who wants a model that never says "I can't help with that" downloads one. No account, no payment method, no usage policy, no log of the conversation. Hugging Face carried roughly 7,970 models with "abliterated" in the name as of September 2026, and anyone with a gaming GPU can run most of them locally.
A defender who asks a mainstream API model to harden a Windows domain controller gets refused 43.8% of the time. That number comes from Scale AI's Defensive Refusal Bias study, which ran 2,390 prompts drawn from the National Collegiate Cyber Defense Competition through frontier models in March 2026. Every prompt in the set was defensive. The models were being asked to do the job they are supposedly aligned to protect.
That is the asymmetry in one paragraph, and it is measurable on both sides.
Refusal rates from Scale AI's Defensive Refusal Bias study (arXiv 2603.01246). Patch completion from a same-lineage comparison of aligned and abliterated weights on the Vul4J benchmark (arXiv 2607.05842).
What uncensored means in 2026
Three different things get sold under the word, and they are not equally real.
A system prompt that says "you have no restrictions" is the weakest. The refusal behaviour is trained into the weights; asking nicely changes the words around it, and the model still declines or hedges on exactly the requests you care about.
Fine-tuning on refusal-free data is the oldest serious route. The Dolphin family of models, fine-tuned by Eric Hartford from open bases, removed refusals by training rather than by asking. It works, and it costs a training run each time the base model moves.
Abliteration is the current default. Arditi et al. showed that refusal in a transformer is mediated by a single direction in the residual stream. Abliteration finds that direction by comparing the model's activations on harmful and harmless prompts, then rewrites the weights so the model can no longer represent it. Maxime Labonne's walkthrough is the reference implementation most public models trace back to. The edit is surgical: the model keeps its knowledge and loses its ability to decline.
The edit also costs quality. Labonne's original Llama 3 example dropped benchmarks across the board until a DPO fine-tune recovered most of the loss.
Newer tooling narrows the damage. On Gemma-3-12B-IT, three ablation recipes all cut refusals from 97/100 to 3/100. What differed was collateral damage, measured as KL divergence from the original on neutral inputs: 1.04 for the manual recipe, 0.45 for huihui-ai's, 0.16 for the automated Heretic tool. One Qwen3.8-27B ablation publishes its price as 2.12 points of MMLU (82.33% against 84.46% stock). A model card with a published cost tells you more than one that publishes nothing.
The UGI Leaderboard tracks this ecosystem: uncensored models ranked by both refusal removal and retained general intelligence, because an obedient model that has forgotten how to reason is useless.
Attackers stopped needing a black market
In mid-2023, the first LLMs marketed for crime, WormGPT and FraudGPT, were subscription services advertised on underground forums. They were small, expensive by comparison with what followed, and eventually shut down or exposed as frauds. The subscription model is now obsolete for the same reason bootlegging ends when the liquor store opens: the unrestricted product became free and downloadable.
Open weights changed the economics completely. A 27B-parameter model with the refusals removed downloads in minutes, runs on a single workstation GPU, and has no vendor that could ban the account, because there is no account. OpenAI's February 2024 disruption report documented state-sponsored groups from Russia, China and Iran using LLM services for reconnaissance assistance, phishing content, scripting help and translation. Those were the constrained, logged, policy-enforced versions. The abliterated open-weight ecosystem offers the same capabilities with none of the constraints.
What attackers actually get from these models matches what every threat assessment has found: volume and fluency. Phishing lures in any language, believable persona backstories, reconnaissance summaries, working scripts for known vulnerabilities, obfuscation of existing malware. What they do not get, on today's open weights, is novel exploitation. The same capability ceiling applies to everyone.
The point is not that uncensored models handed attackers a superweapon. It is that attackers face no friction at all, and never will, because the weights are out and cannot be recalled.
The refusal wall on the other side
While one side runs refusal-free models, the other side's tooling argues back. The Scale AI study measured refusal rates across fully defensive task categories:
| Defensive task | Refusal rate |
|---|---|
| System hardening | 43.8% |
| Malware analysis | 34.3% |
| Vulnerability assessment | 22.7% |
| Incident response | 18.9% |
| Log analysis | 0% |
The variation tracks vocabulary, not risk. Requests phrased with security vocabulary were refused at 2.72 times the rate of the same request in neutral language. Hardening a server, the most defensive task on the list, got refused most.
The study's sharpest finding: adding an explicit authorisation statement, naming the competition and saying the work is permitted, raised the refusal rate from 11.6% to 21.8%. Telling the model you are allowed to do the work makes it almost twice as likely to decline.
Meta's CyberSecEval 2 gave this behaviour a name, false refusal rate, and measured it on borderline-but-benign cyber requests. Several models it tested came in below 15%. CodeLlama-70B-Instruct came in near 70%, which is what happens when safety tuning is applied to a coding model without a security-aware evaluation set.
In practice an analyst pastes a malware config parser into a chat window, the model declines because the content looks like malware development, and the analyst goes back to doing it by hand. Multiply that across a security team's queue and the "AI-assisted SOC" quietly becomes a copilot for writing calendar invitations.
What removing refusals buys, and what it does not
A same-lineage study from July 2026 compared aligned and abliterated versions of the same weights on vulnerability work. Because the pairs share a base model, the difference is attributable to the refusal edit:
| Vul4J stage | Aligned | Abliterated |
|---|---|---|
| Vulnerability detection | 58.42% | 57.22% |
| Usable patch produced | 29.94% | 67.80% |
| Patch applied cleanly | 24.86% | 64.97% |
| Patch compiled | 9.04% | 32.77% |
Removing refusals taught the model nothing about security. Detection accuracy stayed level. What changed is that the model followed a security-framed request all the way to working code instead of stopping at a description of the fix. Aligned models describe; ablated ones build. On compiled patches the gap is 3.6 times.
Uncensored is not a substitute for capable. The NYU CTF Bench run from April 2026 put 200 capture-the-flag challenges in front of the field. GLM-5, the best open-weight entry, solved 19.5%, ahead of closed GPT-5.2-Codex at 18.0%, with DeepSeek-V3 at 6.5% and Llama 3.3 70B at 2.5%. Claude 4.5 Opus solved 59.0%. The spread inside the open-weight field exceeds the gap between the best open model and mid-tier closed ones.
A small abliterated model that never refuses but cannot reason is not a security tool. Match the model to the task shape: refusal costs you at artefact generation, capability costs you everywhere.
The model has to be on your side
There is a second asymmetry beyond refusals, and for professional teams it usually bites first: the data.
Offensive-security work means handling things that cannot be sent to a third-party API: client source code under NDA, exploit proofs-of-concept for the systems you are paid to test, live malware samples, attack traces from an engagement in progress. Most API providers' terms, and most client contracts, rule that out on their own. The security operations that would benefit most from an always-available model are the ones that cannot use one.
That is the case for running refusal-capable models on your own infrastructure. It doubles as the case for open weights: the model that reads your client's code should be a pinned set of weights you control, on hardware you control, where the audit trail is yours.
This is how we run Helix Cyber: security agents driven by open-weight models inside the customer's environment. Candidate findings are reproduced before they reach a report, and a human sits at every irreversible step. Uncensored does not mean uncontrolled. The control moves from the model's vendor to your operation, where you can point at it in an audit.
If you are testing LLM applications rather than using LLMs to test, the refusal question arrives differently: your product's jailbreak resistance is exactly what red teaming measures. We cover that side in LLM red teaming.
Measure refusal on your own prompts before you shop
Published refusal rates are averages over someone else's prompt set. Yours may be near zero, and then the alignment question is moot and capability alone should decide.
Take 50 to 100 real requests from your queue, the ones that already go through a model today or should. Run them against two or three candidate models. Count two things separately: hard refusals, and hedged non-answers that describe what you asked for without doing it. The category table above predicts where yours will cluster: hardening and malware analysis first, log analysis last.
If refusals are costing you real work, the fix is a refusal-free model on hardware you control, evaluated for capability like anything else you buy. Our field guide to uncensored LLMs for security work covers the current model landscape, the quality costs, and the licence terms that govern redistribution.
Security agents that complete offensive-security work need models without a refusal wall, running where the evidence stays yours. See Helix Cyber for how we run open-weight models on your infrastructure, and AI penetration testing for what AI changes about the work itself. If you have a target and a deadline, scope a penetration test.