01 · Mental model
Injection exploits instruction/data ambiguity
Direct injection arrives in a user's request. Indirect injection is embedded in content the system retrieves or reads: email, documents, web pages, issue text, code comments, or tool output. The attack succeeds when that text influences behavior beyond its data role.
Prompt hardening is one layer. Stronger defenses reduce privileges, isolate untrusted content, constrain tools, validate outputs, require approvals for consequences, monitor decisions, and assume some attacks will reach the model.
02 · Visual explanation
03 · Compare and decide
Direct versus indirect injection
| Decision lens | Direct | Indirect |
|---|---|---|
| Entry | User prompt | Retrieved or tool-supplied content |
| Visibility | Often obvious in conversation | Hidden inside legitimate-looking evidence |
| Example | Ignore policy and reveal secrets | Document instructs the agent to upload data |
| Control emphasis | Input handling and permissions | Provenance, content isolation, permissions, egress |
04 · Cybersecurity example
Malicious incident ticket
A SOC copilot retrieves a ticket containing hidden instructions to query unrelated customer data.
The ticket is labelled untrusted.
Retrieval filters access by incident scope.
The query tool enforces table and tenant allowlists.
An unusual request is denied and recorded.
Outcome: The model may be influenced, but the system prevents the instruction from becoming an authorized action.
05 · What to remember
The 60-second recall
Assume untrusted text can reach the model.
Limit the consequences available after model compromise.
Test end-to-end attack paths, not only prompt wording.
Teach-back prompt: Explain this concept to a teammate using the diagram, then name one failure mode and the control that stops it.
06 · Questions people ask
FAQ
07 · Primary sources