A finance team uses an AI agent to process vendor invoices — reading PDFs, matching to purchase orders, and initiating payment approvals. An attacker sends a legitimate-looking invoice PDF. Embedded in the PDF metadata and footer text are instructions telling the agent to update the payment destination for this vendor to a new account number before processing. The agent updates the routing and queues the payment. The approval workflow triggers normally.
The payment processes. It goes to the wrong account. The real vendor never gets paid.
A departing employee — or a compromised contractor — embeds hijack instructions in an internal wiki page, policy document, or shared template that agents are configured to query as a trusted knowledge source. Months later, an agent retrieves the document during a routine task and executes the embedded instructions — exfiltrating current employee data, API keys, or system configurations to an external endpoint the attacker still controls.
The attacker left the company. Their attack stayed behind.
Rather than issuing a single obvious override, an attacker interacts with an agent over dozens of sessions — each time subtly reframing its role, expanding its perceived permissions, or reinforcing a slightly altered objective. No single message is alarming. But across weeks of interaction, the agent's behavior drifts: it becomes more willing to share information it previously declined, more likely to interpret ambiguous requests in the attacker's favor, more helpful to the wrong person.
There is no moment of compromise. The goal just slowly became something else.