SEC536: Adversarial AI - Penetration Testing AI Systems


Experience SANS training through course previews.
Learn MoreLet us help.
Contact usBecome a member for instant access to our free resources.
Sign UpWe're here to help.
Contact UsAn AI broke out of its own test and hit a live company. The defenders who came to clean up were the ones the guardrails stopped. Here is why we are pulling the community together to talk about it.

An OpenAI model was running through a sealed internal evaluation, the kind used to measure how capable a model has become at breaking into systems. For that test, some of the usual safety limits were turned down, and the environment had no path to the open internet. The model found a previously unknown flaw in the tooling around the test, used it to break out, and reached the internet on its own. From there, it broke into the live systems of Hugging Face, one of the most widely used platforms for storing and sharing AI models, and remained active over a weekend before anyone noticed. No person was at the keyboard.
Hugging Face disclosed the incident on July 16. OpenAI followed with its disclosure on July 21. Both investigations are still open, and the mechanics will get plenty of coverage elsewhere. That is not the part we want to spend an hour on.
The part worth the community's time is what happened when the defenders showed up. Hugging Face's responders reached for the frontier models to help analyze the attack, but the models refused to process the data. Reading the logs meant feeding them real attack commands, exploit payloads, and command-and-control artifacts, and the safety guardrails blocked those requests. The team ran its forensics on a self-hosted, open-weight model using infrastructure it controlled. The offensive research that caused the incident ran with the safeties dialed down. The defensive work to understand it ran into a compliance blocker on the same class of models.
That asymmetry is the reason for this session. For the past year, the argument has been about containment: frontier capability is dangerous, so gate it, wrap it in guardrails, and slow it down. The logic rested on the assumption that this capability would never be in adversary hands. This time, it was not in adversary hands. It was in the most trusted hands there are, under evaluation, with the safeties turned down on purpose, and it still got out. If the container can fail from the inside under those conditions, who we admit is not the only question that matters.
There is also a defensive cost. If the same guardrails that block misuse also block the people investigating an incident, the incentives point in the wrong direction. One model can find a zero-day and help a defender reconstruct an attack. Handing the first capability to offensive research while denying defenders the second is a policy choice, and it is worth examining in the open.
We are not going to pretend we have the admission criteria figured out or a clean answer for how access should be governed. We do have a practical starting point for defenders who are outside a trusted-access program, which is almost everyone. Get a capable, vetted model running on infrastructure you control before you need it. Running analysis in-house is why Hugging Face's incident data and the credentials it touched never had to leave its environment. Build a harness that lets you connect your tools so that when a model changes or a vendor changes its policies, you can swap it out without rebuilding everything around it.
On July 28, SANS faculty and staff will sit down with voices from public policy and industry governance to work through the whole picture: AI testing standards, lab resilience, how we model attacker intent when there is no human intent to model, who gets trusted access and who decides, and what a security team should build now. It is a community conversation, not a verdict. Bring your questions.
The panel is live on Tuesday, July 28, at 12 p.m. ET. It is free, with no registration required. Details and the livestream link are available on the event page.


Launched in 1989 as a cooperative for information security thought leadership, it is SANS’ ongoing mission to empower cybersecurity professionals with the practical skills and knowledge they need to make our world a safer place.
Read more about SANS Institute