Slide 16 of 28
Part 2 — Universal PatternSlide 16
Slide 16 · The Universal Pattern
Every scenario shares the same underlying structure: the human's oversight function was degraded before it was needed.
The shared structure across all 6 scenarios

The surface details vary — fraud scores, wire transfers, compliance advisories, firewall rules — but each scenario has identical underlying structure:

  1. A trust relationship exists: The human has learned to rely on the agent's output as a primary input to their decision.
  2. The trust relationship degrades oversight capacity: The human's independent verification is reduced — through deliberate efficiency design, rational adaptation to a reliable tool, or conscious process redesign to move decisions through a human-AI loop rather than human-only review.
  3. A condition arises that falls outside the trust's valid scope: The agent encounters a case it handles poorly (blind spot attack), receives manipulated input (credibility relay, urgency injection), or the accumulated weight of reliance has eroded the human's independent capacity (dependency erosion).
  4. The human's degraded oversight function fails to catch it: The human is in a trust-extended mode at exactly the moment when independent verification is most needed.
Three missing properties — one per defense category

Missing: Calibrated uncertainty communication. In every scenario, the agent presented its output with more confidence than the situation warranted. None of the agents said "this case falls near the boundary of my training data" or "I encountered content I cannot verify." Human oversight was not triggered because the agent did not signal that it was warranted.

Missing: Mandatory oversight gates. Human oversight in every scenario was discretionary — the human could invoke it, but the system did not require it. The trust relationship provided an implicit path to skip oversight. No scenario had a hard enforcement gate that the agent's verdict could not bypass.

Missing: Human independence maintenance. In every scenario, the human was operating in "AI-assisted" mode, meaning their own independent judgment capacity was partially or fully suppressed. Even where they were nominally "reviewing" the AI's output, they were not engaging their independent evaluation capability. The system design had converted the human from an evaluator into a confirmer.

The oversight gap irreversibility problem

Unlike many security failures that can be contained after detection, trust exploitation failures are often irreversible at the moment of discovery: the wire has been sent, the firewall rule is live, the 2FA has been disabled across the organization, the contract has been signed. The human oversight function that was supposed to prevent these outcomes failed at the decision point — and the decision was made. Detecting the failure after the fact is damage assessment, not defense.

This makes prevention — specifically, maintaining human oversight capacity before it's needed — the only effective defense. Detection is too late.

← Back Part 3 — Prevention →