Slide 1 of 27
Part 1 · What Is It?Slide 1
PART 1
What Is It?
Slides 1–8 · No jargon yet
Slide 1 · The Setup
Before we define anything — read this story.
This happened. Follow it. The definition will make sense after.
The Scenario

A company builds an internal HR assistant on top of an LLM. It's wired into the employee database — it can look up PTO balances, answer benefits questions, and pull up policy documents. It works fine for months.

Then This Happens

An employee asks a routine question about their own benefits enrollment. Buried in the assistant's answer is a stray sentence that includes another employee's salary figure and a manager's private performance-review note.

Nobody attacked anything. Nobody typed a malicious prompt. The model just said it.

What Just Happened

This is sensitive information disclosure — when an LLM exposes private, proprietary, or confidential data through its own output, often without anyone trying to extract it. Sometimes it's a bug. Sometimes it's an attack. Either way, data that should have stayed contained is now somewhere it was never supposed to be.

One Line to Remember

Sensitive information disclosure is when an AI system reveals data it should have kept private — through its answers, not through a hack.

That makes sense → What counts as ‘sensitive’?