Six categories of harm from miscalibrated human-agent trust — across security, safety, and oversight.
Security decisions made below threshold of scrutiny
An agent's recommendation is accepted without the independent verification that the decision's risk level warrants. Attackers deliberately craft scenarios that trigger low-risk recommendations, bypassing human review entirely.
Attacker content delivered with AI credibility
Malicious content — phishing, false instructions, social engineering — is routed through an agent, which presents it to a human with an implicit AI-endorsement. The human treats the content as AI-vetted and trusted, reducing skepticism of attacker-controlled material.
Urgency framing bypasses deliberate review
An agent frames a situation as time-critical, suppressing the human's inclination to pause and verify. The urgency is either fabricated by an attacker who compromised the agent, or is a design flaw that creates pressure to approve before thinking.
Oversight erosion from alert fatigue
Excessive low-quality alerts from an agent cause humans to stop engaging meaningfully with its output. When a real threat arrives, the human's attention has been depleted by false positives. The alert goes unnoticed in the noise.
Skill atrophy and judgment loss
Long-term dependence on an AI agent causes human experts to lose the independent skills they would need to evaluate the agent's output critically. When the agent is wrong, the human's capacity to detect it has degraded.
Accountability diffusion
"The AI said so" becomes a defense for human decisions that weren't scrutinized. Responsibility for decisions diffuses between human and AI, creating gaps in accountability — particularly for decisions that cause harm.