KSKS Security Research
Learning map / Session 2 / Scenes 01–06

SECURITY 01 · INJECTION

The content is data.The instruction inside it is an attack.

Understand direct and indirect prompt injection, why prompt hardening is incomplete, and where system controls can still stop harm.

prompt injectionindirect injectiontrust boundaryOWASP
Learning guide
Level
Start here
Reading time
12 min
Presentation
Session 2
Progress
1 of 6

01 · Mental model

Injection exploits instruction/data ambiguity

Direct injection arrives in a user's request. Indirect injection is embedded in content the system retrieves or reads: email, documents, web pages, issue text, code comments, or tool output. The attack succeeds when that text influences behavior beyond its data role.

Prompt hardening is one layer. Stronger defenses reduce privileges, isolate untrusted content, constrain tools, validate outputs, require approvals for consequences, monitor decisions, and assume some attacks will reach the model.

02 · Visual explanation

01Hostile contentembedded instruction
02Retrieverbrings text into context
03Modelinterprets mixed signals
04Tool requestproposed action
05Control planeallow, deny, escalate
The indirect-injection pathA hostile document becomes dangerous only when downstream authority and controls allow impact.

03 · Compare and decide

Direct versus indirect injection

Decision lensDirectIndirect
EntryUser promptRetrieved or tool-supplied content
VisibilityOften obvious in conversationHidden inside legitimate-looking evidence
ExampleIgnore policy and reveal secretsDocument instructs the agent to upload data
Control emphasisInput handling and permissionsProvenance, content isolation, permissions, egress

04 · Cybersecurity example

Malicious incident ticket

A SOC copilot retrieves a ticket containing hidden instructions to query unrelated customer data.

01

The ticket is labelled untrusted.

02

Retrieval filters access by incident scope.

03

The query tool enforces table and tenant allowlists.

04

An unusual request is denied and recorded.

Outcome: The model may be influenced, but the system prevents the instruction from becoming an authorized action.

05 · What to remember

The 60-second recall

01

Assume untrusted text can reach the model.

02

Limit the consequences available after model compromise.

03

Test end-to-end attack paths, not only prompt wording.

Teach-back prompt: Explain this concept to a teammate using the diagram, then name one failure mode and the control that stops it.

06 · Questions people ask

FAQ

No. Useful and malicious instructions can be semantically similar, and indirect attacks may be contextual. Use layered controls and minimize downstream authority.

07 · Primary sources

Continue with authoritative guidance