AI News
  • Home
  • The Hague
  • Moving to NL
  • Tech News
    • AI & Tech
    • Machine Learning
    • Startups
    • Tools & Apps
    • Robotics
    • Future Tech
    • AI in Sport ⚽
    • AI in Health
    • AI in Education
    • AI in Finance
    • AI in Business
    • AI in Law
    • AI in Climate
No Result
View All Result
AI News
  • Home
  • The Hague
  • Moving to NL
  • Tech News
    • AI & Tech
    • Machine Learning
    • Startups
    • Tools & Apps
    • Robotics
    • Future Tech
    • AI in Sport ⚽
    • AI in Health
    • AI in Education
    • AI in Finance
    • AI in Business
    • AI in Law
    • AI in Climate
No Result
View All Result
AI News
No Result
View All Result

AI Guardrails Are Slowing Down Security Researchers

Ramo by Ramo
25 July 2026
in AI & Tech
0
How AI guardrails are impeding the work of offensive cybersecurity researchers
0
SHARES
5
VIEWS
Summarize with ChatGPTShare to Facebook

Ask a mainstream AI model to write a working exploit and it will usually stop you cold. That refusal is by design. It is also, for a small but consequential group of professionals, a growing problem.

TechCrunch reported on July 23, 2026 that several offensive cybersecurity researchers — the people who hunt for unknown vulnerabilities and build the tools to exploit them — say the safety guardrails baked into models from OpenAI and Anthropic are getting in the way of legitimate work. These are not hobbyists poking at systems for fun. They are the researchers whose findings feed vendor patches, harden critical software, and keep real attackers a step behind. And they are increasingly bumping into the same walls the models put up to block criminals.

When the safety net catches the wrong people

The logic behind the guardrails is sound. Companies like OpenAI and Anthropic have every reason to stop their systems from generating functional malware, phishing kits, or ready-made intrusion code on demand. A model that cheerfully writes ransomware for anyone who asks is a liability with a marketing budget. So the labs train their systems to recognize the shape of an offensive request and decline.

📖
RECOMMENDED READ
The Coming Wave: AI, Power, and the Greatest Dilemma of Our Age
Mustafa Suleyman
The definitive book on where AI is heading - written by one of the field founders.
View on Amazon →affiliate link

The trouble is that offensive security research looks, at the level of text and code, almost identical to offensive security crime. A penetration tester probing a client’s network and a criminal breaking into that same network are doing much of the same work. Both write code that finds flaws. Both craft payloads. Both study how to slip past defenses. A model cannot always read intent, so it errs toward refusal, and the refusal lands on defender and attacker alike. Researchers told TechCrunch that this friction is now a routine part of their day.

The cost is slower defense, not blocked crime

Here is the uncomfortable irony. Determined criminals have options. They can turn to unaligned open-weight models, jailbroken systems, or purpose-built underground tools that carry no guardrails at all. The commercial safety layer that OpenAI and Anthropic maintain is not the thing standing between a serious attacker and their goal.

Legitimate researchers, by contrast, tend to play by the rules. They use the sanctioned, well-supported commercial models because those are the best tools available and because their employers expect them to stay inside terms of service. When the guardrails slow them down, the people who actually lose time are the ones working to close holes before the bad actors find them. The defender pays the tax; the attacker routes around it.

That dynamic reshapes how researchers work. Some rephrase requests, split a task into innocuous-looking fragments, or strip out the context that trips the filter, which wastes time and strips away exactly the information that would let a model help intelligently. Others give up on the frontier tools for certain tasks and fall back to manual methods or local models that are less capable but more compliant. Every one of those detours is friction that a criminal, working outside the rules, simply never encounters.

What the labs are actually balancing

None of this makes the guardrails a mistake. OpenAI and Anthropic are managing genuine risk, and the reputational and legal stakes of a model that mass-produces attack tooling are enormous. The hard part is precision. A blunt filter that treats every mention of exploitation as malicious is easy to build and easy to defend in a press statement. A nuanced one that can tell a red-team engagement from a real intrusion is far harder, and it requires trusting signals about who is asking and why.

Some of that trust can be engineered. Verified accounts, enterprise agreements, and vetted research programs give the labs a way to extend more latitude to users who have shown they are legitimate. The question is whether that access scales to the independent researchers and small shops who do a large share of the field’s most valuable work and who rarely have an enterprise contract to wave.

Where this goes next

The tension between safety and capability is not going to resolve on its own, and the security community is unlikely to accept a permanent handicap on its best tools. Watch for the labs to build more granular access tiers for vetted security professionals, and watch for researchers to keep drifting toward open models they can run without asking permission. If the guardrails stay blunt, the people they inconvenience most will be the defenders, while the attackers keep operating in the one place no guardrail reaches.

The open question is whether OpenAI, Anthropic, and the security field can agree on what a trusted researcher looks like before the friction pushes serious work off the platforms entirely.

For more coverage of AI and cybersecurity, visit Mylistingo.

Source: Original Article

SummarizeShare
Ramo

Ramo

Ramo is the editorial voice of Mylistingo — an AI and technology news platform based in The Hague, Netherlands. Covering artificial intelligence, machine learning, robotics, and the future of technology, Ramo delivers accurate, accessible reporting for both general audiences and industry professionals. Every article is fact-checked and written to meet Mylistingo's strict no-fabrication editorial standards.

Related Stories

Clinician reviewing patient data on a laptop with a stethoscope nearby

OpenAI Plugs ChatGPT Into Epic’s Patient Records

by Ramo
4 September 2026
0

OpenAI connected ChatGPT for Healthcare to Epic, the record system behind 325 million patients. Access is read-only, and the stakes are high.

Abstract visualisation of an artificial intelligence network

Anthropic’s Fable 5.1 Cuts AI Agent Costs by Up to 45%

by Ramo
3 September 2026
0

Claude Fable 5.1 keeps base prices flat but cuts cache-read costs 75%, making long-running AI agents up to 45% cheaper. Mythos 5.1 stays gated.

Nvidia’s $3.5B MediaTek bet reveals its plan for tackling Big Tech’s AI chip buildout

Nvidia’s $3.5B MediaTek Bet Against Big Tech AI Chips

by Ramo
31 August 2026
0

A $3.5 billion vote of confidence, and self-defense Nvidia just wrote a $3.5 billion check to a company that doesn't build the chips everyone associates with the AI...

Musk’s faster path to more gas turbines comes with pollution problem

Musk’s Gas Turbine Bet: Faster Power, Dirtier Air

by Ramo
31 August 2026
0

A foundry, a shortcut, and a fuel nobody wants next door Elon Musk has a habit of solving other people's bottlenecks by building the part himself. His latest...

Next Post
As US weighs response to Chinese AI, industry urges against broad open-weight restrictions

US Weighs Curbs on Open AI Models as Industry Pushes Back

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

The Hague, for internationals

One email a week: what changed for expats in The Hague, what's on this weekend, and one guide worth reading. No spam, unsubscribe any time.

Free European Bank Account
Open a 100% mobile bank account in minutes
Free virtual Mastercard, zero foreign transaction fees, and instant European IBAN setup with no paperwork.
Get Started Free
Sponsored · Advertise

Recommended

Meta’s Muse Glimmer Puts Frontier AI on a Laptop

10 August 2026

The World Cup Where AI Called the Offsides

21 July 2026

Popular Story

  • ml_feat_56193023

    ASML’s Next-Gen High-NA EUV Machines Drive Eindhoven Expansion, Creating 20,000 New Jobs

    0 shares
    Share 0 Tweet 0
  • Robotaxis Arrive in Rotterdam: Netherlands Launches Europe’s Largest Autonomous Ride-Hailing Fleet

    0 shares
    Share 0 Tweet 0
  • PixVerse closes $439m series C extension at $2b valuation

    0 shares
    Share 0 Tweet 0
  • Is Your Home Truly Safe The Smart Security Tech You Need in 2025

    0 shares
    Share 0 Tweet 0
  • How to Register at The Hague Municipality (Gemeente Den Haag): A 2026 Step-by-Step Guide

    0 shares
    Share 0 Tweet 0
logo ainews

We bring you the best Premium WordPress Themes that perfect for news, magazine, personal blog, etc. Check our landing page for details.

Recent Posts

  • How to Find a Rental Apartment in The Hague in 2026
  • Dutch Tech Today: 7 September 2026
  • MIT’s CW-Net Makes Self-Driving AI Explain Itself

Partner

Free European Bank Account
Open a 100% mobile bank account in minutes
Free virtual Mastercard, zero foreign transaction fees, and instant European IBAN setup with no paperwork.
Get Started Free
Sponsored · Advertise

Categories

  • AI & Tech
  • AI & Tech in the Netherlands
  • AI in Business
  • AI in Climate
  • AI in Education
  • AI in Finance
  • AI in Health
  • AI in Law
  • AI in Sport
  • Economy & Finance
  • Future Tech
  • Machine Learning
  • Moving to the Netherlands
  • Politics & Geopolitics
  • Robotics
  • Social Topics
  • Sport
  • Startups
  • The Hague
  • Tools & Apps
  • Uncategorized

The Hague, for internationals

One email a week: what changed for expats, what's on, one guide worth reading.

  • Home
  • Advertise
  • Latest News
  • Contact Us
  • Data Deletion Instructions
  • Editorial Policy

No Result
View All Result
  • Home
  • The Hague
  • Moving to NL
  • Tech News
    • AI & Tech
    • Machine Learning
    • Startups
    • Tools & Apps
    • Robotics
    • Future Tech
    • AI in Sport ⚽
    • AI in Health
    • AI in Education
    • AI in Finance
    • AI in Business
    • AI in Law
    • AI in Climate