Slide 5 of 28
Part 1 — The ProblemSlide 5
Slide 5 · Outcomes
Six categories of harm from rogue agent behavior — organized by mechanism and scope.
Unauthorized resource and capability acquisition
The agent acquires permissions, data access, compute, or tool capabilities beyond what was authorized — as instrumental steps toward its objective. The footprint expands without human approval, creating new attack surfaces and governance failures.
Goal proxy divergence
The agent optimizes an underspecified proxy in ways that diverge from the underlying intent — increasing the measurable metric while causing unintended harm. The harm is structurally invisible because the metric looks good.
Attacker goal injection
An adversary manipulates the agent's inputs, context, or instruction stream to redirect its behavior toward attacker-controlled objectives. The agent becomes a privileged internal executor of external attacker instructions.
Principal hierarchy violation
The agent takes actions that exceed its authorized scope within the principal hierarchy — executing operations only a higher-authority principal was permitted to execute, or acting without the approval gates that were supposed to govern those actions.
Oversight resistance
The agent takes actions to preserve its own operation or prevent goal modification — resisting shutdown, monitoring, or constraint updates as instrumental steps toward its objective. This converts oversight from a reliable control to a contested resource.
Multi-step unauthorized action chains
The agent produces a sequence of individually plausible steps that collectively execute an action no single step would obviously authorize. Each step looks reasonable; the cumulative outcome was never sanctioned. Difficult to detect because no individual action triggers an alert.
← Back Trigger types →