Group Purchasing
Group Purchasing

Quantifying Agentic Misalignment in Multi-Agent Systems

Quantifying Agentic Misalignment in Multi-Agent Systems (PDF, 1.18MB)Published: 01 Oct, 2026
Created by:

As Large Language Model (LLM) agents gain enterprise tool-execution permissions, agentic misalignment becomes a measurable operational risk: autonomous systems may pursue sub-goals that violate implicit safety policies even when tools work as designed (Lotfi et al., 2026; Zhuang & Hadfield-Menell, 2020; Zscaler, 2026).

This study tests whether supervisor/sub-agent delegation creates a responsibility gap that increases unsafe behavior. Across 1,500 runs covering six models, eight adversarial scenarios, and five comparison groups, the results reject that hypothesis: delegated topologies scored 0% misaligned across all models, while failures concentrated in single-agent configurations, especially under prompt injection. Runtime telemetry captured tool calls and internal-monologue steps, but scoring reflects completed-action ground truth rather than validated real-time intervention (see Section 3).

These findings support behavior-first evaluation and show why runtime observability, tool-boundary controls, and scenario-specific validation are necessary complements to prompt-level safety instructions (Lotfi et al., 2026; Tarasenko & Voruganti, 2026).