Ask a mainstream AI model to write a working exploit and it will usually stop you cold. That refusal is by design. It is also, for a small but consequential group of professionals, a growing problem.
TechCrunch reported on July 23, 2026 that several offensive cybersecurity researchers — the people who hunt for unknown vulnerabilities and build the tools to exploit them — say the safety guardrails baked into models from OpenAI and Anthropic are getting in the way of legitimate work. These are not hobbyists poking at systems for fun. They are the researchers whose findings feed vendor patches, harden critical software, and keep real attackers a step behind. And they are increasingly bumping into the same walls the models put up to block criminals.
When the safety net catches the wrong people
The logic behind the guardrails is sound. Companies like OpenAI and Anthropic have every reason to stop their systems from generating functional malware, phishing kits, or ready-made intrusion code on demand. A model that cheerfully writes ransomware for anyone who asks is a liability with a marketing budget. So the labs train their systems to recognize the shape of an offensive request and decline.
The trouble is that offensive security research looks, at the level of text and code, almost identical to offensive security crime. A penetration tester probing a client’s network and a criminal breaking into that same network are doing much of the same work. Both write code that finds flaws. Both craft payloads. Both study how to slip past defenses. A model cannot always read intent, so it errs toward refusal, and the refusal lands on defender and attacker alike. Researchers told TechCrunch that this friction is now a routine part of their day.
The cost is slower defense, not blocked crime
Here is the uncomfortable irony. Determined criminals have options. They can turn to unaligned open-weight models, jailbroken systems, or purpose-built underground tools that carry no guardrails at all. The commercial safety layer that OpenAI and Anthropic maintain is not the thing standing between a serious attacker and their goal.
Legitimate researchers, by contrast, tend to play by the rules. They use the sanctioned, well-supported commercial models because those are the best tools available and because their employers expect them to stay inside terms of service. When the guardrails slow them down, the people who actually lose time are the ones working to close holes before the bad actors find them. The defender pays the tax; the attacker routes around it.
That dynamic reshapes how researchers work. Some rephrase requests, split a task into innocuous-looking fragments, or strip out the context that trips the filter, which wastes time and strips away exactly the information that would let a model help intelligently. Others give up on the frontier tools for certain tasks and fall back to manual methods or local models that are less capable but more compliant. Every one of those detours is friction that a criminal, working outside the rules, simply never encounters.
What the labs are actually balancing
None of this makes the guardrails a mistake. OpenAI and Anthropic are managing genuine risk, and the reputational and legal stakes of a model that mass-produces attack tooling are enormous. The hard part is precision. A blunt filter that treats every mention of exploitation as malicious is easy to build and easy to defend in a press statement. A nuanced one that can tell a red-team engagement from a real intrusion is far harder, and it requires trusting signals about who is asking and why.
Some of that trust can be engineered. Verified accounts, enterprise agreements, and vetted research programs give the labs a way to extend more latitude to users who have shown they are legitimate. The question is whether that access scales to the independent researchers and small shops who do a large share of the field’s most valuable work and who rarely have an enterprise contract to wave.
Where this goes next
The tension between safety and capability is not going to resolve on its own, and the security community is unlikely to accept a permanent handicap on its best tools. Watch for the labs to build more granular access tiers for vetted security professionals, and watch for researchers to keep drifting toward open models they can run without asking permission. If the guardrails stay blunt, the people they inconvenience most will be the defenders, while the attackers keep operating in the one place no guardrail reaches.
The open question is whether OpenAI, Anthropic, and the security field can agree on what a trusted researcher looks like before the friction pushes serious work off the platforms entirely.
For more coverage of AI and cybersecurity, visit Mylistingo.
Source: Original Article







