01 · Mental model
Threat-model the whole system, not only the model call
Identify the user and service identities, model provider, retrieval sources, MCP servers, tools, data stores, memory, observability, outputs, and human approvals. Draw where trust or authority changes. Then ask how an attacker could inject, poison, impersonate, exfiltrate, overreach, persist, exhaust, or hide.
Controls should interrupt the path before consequence: source filtering before retrieval, authorization before tool execution, schema checks before persistence, DLP before egress, and rollback before production change.
02 · Visual explanation
03 · Compare and decide
Turn abuse paths into controls and tests
| Decision lens | Threat | Testable control |
|---|---|---|
| Indirect injection | Hostile retrieved content proposes an action | Tool policy denies scope and records the attempt |
| Credential misuse | Agent obtains standing admin authority | Per-run identity with narrow roles and expiry |
| Data exfiltration | Output contains sensitive evidence | Classification-aware redaction and destination allowlist |
| Memory poisoning | False instruction persists across sessions | Typed memory schema, source, TTL, and approval |
04 · Cybersecurity example
Architecture-review attack exercise
A malicious diagram note asks the reviewer agent to upload the design to an external URL.
The note is classified as untrusted diagram data.
No generic network tool is available.
Allowed tools cannot send artifacts externally.
The denied intent becomes a security test and signal.
Outcome: The architecture removes the exfiltration path rather than depending only on model refusal.
05 · What to remember
The 60-second recall
Threat models include identities, stores, tools, operators, and data movement.
Design out dangerous capabilities before adding detection.
Every important control should have an abuse-case test.
Teach-back prompt: Explain this concept to a teammate using the diagram, then name one failure mode and the control that stops it.
06 · Questions people ask
FAQ
07 · Primary sources