01 · Mental model
Controls belong throughout the execution path
Input controls classify and constrain what enters. Retrieval controls preserve access and provenance. Tool controls authenticate, authorize, validate, and budget actions. Output controls enforce schemas and data policy. Human review owns consequential exceptions. Observability shows whether the controls work in practice.
Model-based classifiers can add useful judgment, but deterministic policy should own permissions, resource scope, approval state, and hard limits. The model operates inside the control plane—not above it.
02 · Visual explanation
03 · Compare and decide
Probabilistic judgment versus deterministic authority
| Decision lens | Model or classifier | Control plane |
|---|---|---|
| Prompt risk | Estimate suspicious intent | Apply policy to the resulting capability |
| Tool arguments | Propose structured values | Validate schema, identity, scope, and limits |
| Output quality | Critique clarity and completeness | Reject invalid schema or prohibited data |
| Approval | Explain impact and uncertainty | Require an accountable approval token |
04 · Cybersecurity example
Controlled Sentinel rule publication
An agent drafts a detection and wants to publish it to production.
Schema and KQL checks run automatically.
A policy gate confirms workspace and change window.
A reviewer sees evidence, blast radius, and rollback.
A separate deployment identity performs the approved change.
Outcome: The model contributes reasoning without becoming the production authority.
05 · What to remember
The 60-second recall
Guardrails should map to real boundaries and consequences.
Human review needs exact evidence, impact, and requested action.
Deterministic policy owns permissions and hard stops.
Teach-back prompt: Explain this concept to a teammate using the diagram, then name one failure mode and the control that stops it.
06 · Questions people ask
FAQ
07 · Primary sources