SEC536: Adversarial AI - Penetration Testing AI Systems


Experience SANS training through course previews.
Learn MoreLet us help.
Contact usBecome a member for instant access to our free resources.
Sign UpWe're here to help.
Contact UsWhen an AI is 70% accurate at automating each task of a 10-task investigation, 97% of cases end up incomplete - and that's the optimistic version. We set out to close this gap for cloud and SOC investigations, and using CTFs as one of our challenge sets, we achieved the first public agentic speed-run of Splunk Boss of the SOC. To our horror, we realized frontier AI had been breaking our evals along the way. This talk is about what it takes to make agentic investigations perform measurably well, and gaining confidence that we're not tricking ourselves. The first half systematically walks through how to push AI investigators from barely passing to 96%+. Starting with accessible off-the-shelf tools like Claude Code as a baseline, we layer on the methods that actually tackle cloud investigation tasks in BOTS — leaky S3 buckets, identity pivots, and the kind of multi-source correlation real cloud SOCs run daily. We cover impacts of model selection, semantic layers over log data, agent harness design, MCP server choices, and prompting patterns that move the needle. That includes showing which old best practices are now obsolete, what's newly viable with open-source models and harnesses, and the cost/latency tradeoffs at each step. More fundamentally, we zoom out to what this means organizationally and how teams should structure rollout and iteration. Measurement is foundational, so the second half gets serious about evals for investigations. An emerging theme driven by model advances is that modern evals must now be treated adversarially. Frontier agents are returning correct answers without ever touching the logs - they're getting good at cheating. By understanding how agents cheat evals, we learn to build better measurements, and through them, more reliably improve our automations. Attendees leave with a practitioner's playbook for deploying agentic AI investigators in cloud environments that perform reliably, how to compare techniques at a more fundamental level, plus botsbench.com as an open resource for benchmarking and detecting contamination in their own pipelines.


Leo Meyerovich is the founder and CEO of Graphistry, makers of Louie.ai , and has spent the last decade advancing GPU, graph, and AI technologies for data-intensive investigations.
Read more about Leo Meyerovich










