What it is: The attacker doesn't need to fool the human directly. They only need to fool the agent — knowing the human will defer to the agent's verdict. They study the agent's known limitations (training data cutoffs, underrepresented case categories, score calibration weaknesses) and craft their attack to land precisely where the agent systematically underweights risk.
Mechanism: Agent outputs low-risk verdict → human's trust suppresses independent review → attacker's action passes without the scrutiny that would have caught it.
Why it's hard to detect: The human behaves rationally given what they see. The agent behaves as designed. No individual system fails in a detectable way. The vulnerability is in the interaction between them.
Key insight: The more accurate the agent is on average, the more dangerous this pattern becomes — because high average accuracy builds deep trust, which makes the human's deference at edge cases more automatic and harder to interrupt.
What it is: The agent — whether compromised by an attacker or poorly designed — presents a recommendation with artificial authority (claiming executive approval, regulatory mandate, or organizational urgency) and/or time pressure ("action required within 3 minutes," "system will roll back if not confirmed"). Both suppress deliberate evaluation.
Authority spoofing: "This action has been pre-approved by the CFO and Legal — your confirmation is procedural only." Even if the human is skeptical, the apparent pre-authorization removes the perceived need to independently verify. They provide confirmation, not a decision.
Urgency spoofing: Time pressure is the enemy of careful thought. An agent that can generate urgency signals controls the human's deliberation window. Under time pressure, humans default to heuristics — including "the AI said it's fine" — rather than careful independent evaluation.
Why routing through an agent makes it more effective: A phone call from an unknown person claiming CFO approval triggers skepticism. A notification from the trusted internal AI assistant saying the same thing triggers less skepticism — because the AI has institutional credibility and the human assumes it wouldn't generate the notification without valid backing.