Group Purchasing
Group Purchasing

Nine Questions That Tell You Whether Your Team Could Handle an AI-Run Attack

Authored bySANS Institute
SANS Institute

An AI agent recently ran a four-day autonomous attack against a real company. This was an actual intrusion, in which the agent independently conducted reconnaissance, gains access, and moved lateral through a production environment defended by a real security team. That team survived it, then did something rare in this industry: they talked openly about what went wrong. 

One line from the team’s debrief has stuck with everyone who heard it: "There were alerts. They did not rise to the right level. How does the SOC miss this?" 

The honest answer is uncomfortable. The SOC missed it because the detection stack was tuned for human attackers. An agent working slowly across multiple paths never generated a single event severe enough to page anyone. Viewed in isolation, every alert looked like noise. The team was watching the attack unfold the entire time but could not see it. 

SANS took the lesson from that debrief and built something you can use before your own 2 a.m. phone call: a nine-question readiness check drawn directly from the failures and near misses in that incident.  

Work through it here: Is Your Team Ready to Handle an AI-Run Attack? 

Why a Questionnaire and Not Another Threat Report

Because the gap exposed by this incident has nothing to do with awareness. Practitioners already know AI-driven attacks are happening. They have read the headlines. The gap is that most teams have never audited their own environments against how these attacks actually behave, and that behavior breaks many of the assumptions on which detection programs were built. 

Consider three examples, all pulled from the incident itself. 

Alert-severity rules assume a noisy attacker or a critical event. An autonomous agent may give you neither. Readiness means having severity logic designed to detect slow, multipath activity in which no single event appears critical and still ensuring the person on call gets paged at 2 a.m. on a Saturday. 

Your own automation is now camouflage for someone else’s. If you do not have a baseline for what your legitimate automation normally does, hostile automation can blend right in. Readiness means understanding your own machine behavior well enough that a stranger’s behavior stands out immediately. 

Deception works better against agents than against people. A careful human attacker may walk past a canary token. An agent issuing thousands of commands could trip one within the first hour. The defending team in this incident said plainly that it should have had deception in place. Most environments still lack it, and against automated attackers, that means leaving one of the cheapest, highest-signal detection methods on the table. 

There is a fourth question in the readiness check that almost nobody is asking yet: If a self-hosted, open-weight model is part of your incident response fallback plan, have you evaluated it for hidden behavior?  

That model is now part of your security infrastructure, and it deserves the same scrutiny as anything else you depend on during a breach. The affected team acknowledged that it had not performed that evaluation. Odds are good your team has not either. 

How to Use the Readiness Check

Work through the nine questions with your team and be honest about the answers. The exercise is diagnostic, not performative. A flattering self-assessment defeats the purpose. 

The questions map to six capability areas. For each one, the page explains what readiness looks like and identifies the hands-on courses that build the relevant skills, Courses are labeled core when the skill is a central learning outcome and partial when it is one piece of a broader curriculum.  

No single course prepares a team for every aspect of an incident like this, and the page says so up front. That is the kind of candor effective readiness planning requires. 

If all nine of your answers come back clean, congratulations, and please tell the rest of us how you did it.  

If they come back the way most teams’ answers will, you now have a specific, prioritized list of improvements based on a real incident instead of a hypothetical one. 

Run the readiness check, then use your answers to build a training plan for your team.