KSKS Security Research
Learning map / Session 2 / Scenes 12, 15, 17

SECURITY 04 · CONTROL

A guardrail is not one filter.It is a control family.

Place input, output, tool, policy, and human controls at the point where they can still prevent harm instead of relying on one filter.

guardrailsHITLpolicycontrol plane
Learning guide
Level
Applied
Reading time
12 min
Presentation
Session 2
Progress
4 of 6

01 · Mental model

Controls belong throughout the execution path

Input controls classify and constrain what enters. Retrieval controls preserve access and provenance. Tool controls authenticate, authorize, validate, and budget actions. Output controls enforce schemas and data policy. Human review owns consequential exceptions. Observability shows whether the controls work in practice.

Model-based classifiers can add useful judgment, but deterministic policy should own permissions, resource scope, approval state, and hard limits. The model operates inside the control plane—not above it.

02 · Visual explanation

01Inputclassify + delimit
02Contextprovenance + ACL
03Toolauthorize + budget
04Outputschema + DLP
05Reviewapprove + record
Controls at the moment of consequenceEach boundary has a different control job; later controls should not compensate for missing earlier ones.

03 · Compare and decide

Probabilistic judgment versus deterministic authority

Decision lensModel or classifierControl plane
Prompt riskEstimate suspicious intentApply policy to the resulting capability
Tool argumentsPropose structured valuesValidate schema, identity, scope, and limits
Output qualityCritique clarity and completenessReject invalid schema or prohibited data
ApprovalExplain impact and uncertaintyRequire an accountable approval token

04 · Cybersecurity example

Controlled Sentinel rule publication

An agent drafts a detection and wants to publish it to production.

01

Schema and KQL checks run automatically.

02

A policy gate confirms workspace and change window.

03

A reviewer sees evidence, blast radius, and rollback.

04

A separate deployment identity performs the approved change.

Outcome: The model contributes reasoning without becoming the production authority.

05 · What to remember

The 60-second recall

01

Guardrails should map to real boundaries and consequences.

02

Human review needs exact evidence, impact, and requested action.

03

Deterministic policy owns permissions and hard stops.

Teach-back prompt: Explain this concept to a teammate using the diagram, then name one failure mode and the control that stops it.

06 · Questions people ask

FAQ

No. Automate low-impact, reversible, well-tested actions. Escalate when consequence, novelty, uncertainty, or policy requires accountable review.

07 · Primary sources

Continue with authoritative guidance