SEC536: Adversarial AI - Penetration Testing AI Systems


Experience SANS training through course previews.
Learn MoreLet us help.
Contact usBecome a member for instant access to our free resources.
Sign UpWe're here to help.
Contact UsSecurity operations analysts spend significant time converting unstructured cyber threat intelligence (CTI) into actionable threat-hunting content, a manual process that is slow, inconsistent, and prone to error. While large language models (LLMs) appear well-suited to automate portions of this workflow, their ability to reliably operationalize CTI has not been rigorously evaluated.
This study evaluates three commercial LLMs, OpenAI GPT-4o, Google Gemini, and Microsoft Copilot, against ten recent CTI reports, measuring indicator extraction accuracy and validating generated Kusto Query Language (KQL) and CrowdStrike Query Language (CQL) hunting queries in Microsoft Sentinel and CrowdStrike NG-SIEM. Rather than relying on execution success alone, the evaluation applies a framework that separates indicator extraction from query generation, distinguishes instruction-following failures from analytical reasoning failures, and treats the completeness of the human analyst baseline as a measurable quantity rather than an assumption.



















