Governed Agentic AIfor Cybersecurity
Opening
⌂ 00:00
The outcome first

One agentic application. Three cybersecurity deliverables.

Reduce repeated delivery effort by turning one governed request into a security review, delivery pack or validated KQL—while experts keep approval.

UnderstandSecureDemonstrate
Cybersecurity
request
Security reviewfindings · evidence · report
Delivery packHLD · LLD · plan · architecture
KQL studioquery · MITRE · schema check
What we are delivering

Three real workspaces. One governed delivery engine.

Current Cyber Orchestrator Security Review workspace
10 versioned skills available253 tests collected in the repository audit0 live cloud write actions in Phase 1
One answer ≠ repeatable delivery

Better prompts improve the answer. They do not create a repeatable workflow.

Security review requestengagement 01
USER REQUESTReview this Sentinel architecture.
SYSTEM INSTRUCTIONSRole · rules · expected behaviour
REFERENCE CONTEXTArchitecture · standards · requirements
EXPECTED OUTPUTFindings · ratings · validation checklist
LLM responseSTRUCTURED REVIEW DRAFT
WHEN THE NEXT ENGAGEMENT ARRIVES
PROJECT 01Build the method
PROJECT 02Copy + edit it
PROJECT 03Copy + check again
Repeat setupInstructions, references and checks are copied or reconstructed.
Method driftSmall prompt changes can alter how the review is performed.
Manual controlThe user still selects inputs, invokes tools, validates results and manages approval.
The prompt contains the method. The method is not yet a managed capability.
Make repeated work reusable

Move what repeats from the prompt into the system.

Remember the context. Package the method. Control the execution.

Less repeated setup. More consistent delivery. Human accountability remains.
Skill anatomy

A skill packages expert work as a reviewable operating playbook.

1
Match request
2
Load playbook
3
Read needed references
4
Run checks
5
Return typed result
LLM → agent

A model becomes an agent when it gains a controlled loop.

LLMreason + propose
Goal + skillspurpose, method and boundaries
Tools + MCPapproved actions and context
State + memoryprogress, artifacts and evidence
Controlslimits, validation and approval
1
LLMAnswers one prompt and normally stops.
2
Single agentOwns one goal and repeats under runtime controls.
3
Multi-agent systemAn orchestrator coordinates genuinely separate specialist roles.
PLAN→ACT→OBSERVE→REPEAT OR STOP
LLM only: useful draft returned inside the conversation; no controlled external action.
The real product pattern

The orchestrator coordinates expertise. Code checks the handoffs.

Ready: scoped request enters the governed pipeline.
Orchestratororder · state · routing
→
Intake skillextract typed scope
→
Human gateapprove scope
↳
Ingestion + costcollection design
DetectionKQL + MITRE coverage
→
G3 · codedo the designs agree?
Solution architectsynthesize architecture IR
→
ArchStudiovalidate + draw.io
→
Document builderHLD · LLD · plan
→
Expert approvalrelease pack
5 conflicts caughtG3 found real collection↔detection disagreements in the golden run.
Resume, do not restartCompleted model work is skipped after an interruption.
Deterministic outputsDocuments, diagrams and manifests are rendered by code.
OWASP Top 10 for LLM Applications · 2025

Useful AI applications create new attack paths. Use a shared checklist.

The governed agentic stack

Security and traceability cross every layer.

Security · may it do this?
Auditability · what happened?
The model can change. The governed control system is the durable asset.
Current controls versus next layer

Show what is enforced now—and what is not yet.

Running today VERIFIED

✓Accounts, signed sessions, quota, rate limits and owner-scoped runs
✓LangGraph/YAML orchestration, fail-closed gates and human pauses
✓Structured outputs, per-node token budgets and revision limits
✓Read-only Learn/AWS MCP and Sentinel catalogue connectors with allowlists
✓Deterministic schema gates, journal and artifact lineage manifest
✓Optional Langfuse tracer integration; prompt and response capture is opt-in
✓Locked dependencies, Dependabot, tests, lint and container smoke
✓No live cloud write actions in Phase 1
Evidence boundary · local main d7ecd96 · private GitHub main verified · Railway health returned 10 skills · authenticated Railway configuration was not inspected
Langfuse and SIEM

Langfuse explains the AI run. The SIEM explains the security event.

Langfuse · AI workflow

What happened inside the run?

session · user · environment
orchestrator span
model generation
tool + retrieval span
validation + score
compact security event
+ trace ID
SIEM · security operations

What does it mean across the environment?

  • Identity and authentication
  • Application, cloud and network events
  • Cross-source correlation
  • Alert, incident and response
  • Retention and investigation
Ready.
Langfuse is not a SIEM replacement.
MINIMIZE · MASK · CONTROL ACCESS TO PROMPTS, RESPONSES, RETRIEVED DATA AND TOOL PAYLOADS
Live demonstration

Follow one request from intent to reviewed evidence.

Design a Microsoft Sentinel platform for a hybrid customer. Produce the HLD, LLD, project plan and editable architecture. Show assumptions, unresolved decisions and validation evidence.
Rehearsal simulation ready. No backend request has been sent.
Intent → evidence
Normalize
scope
→
Approval
gate
→
Route
specialists
→
Gather
evidence
→
Run
checks
→
Generate
architecture
→
Assemble
pack
Watch

Where does it pause?

Verify

What is checked mechanically?

Evidence

Can we explain the run?

What to remember

Package expertise. Bound action. Preserve evidence.

Example: the security-architect skill versions the review method, references, schemas and deterministic helpers.
What belongs in the skill, what belongs in the orchestrator, and what must remain a human decision?
Next session · Build and Secure Your First Cybersecurity Agent: tenant isolation, policy-as-code, adversarial testing, evaluation datasets, SIEM correlation and controlled write actions.
Explore the Learning Hub →
Space / →Reveal, then move forward←Reverse reveal, then move backM · N · RMenu · notes · referencesF · ?Fullscreen · help
LLM
Language model that generates or transforms text and code from context.
Skill
Reusable operating playbook with instructions, references and deterministic assets.
Agent
Model inside a controlled loop that can choose the next allowed step.
Tool
Approved function or service the application may execute.
MCP
Standard connection method through which a server exposes tools or context.
Runtime
Software managing model turns, tool requests, limits and structured results.
Orchestrator
Workflow controller for order, state, routing, retries and approval pauses.
Trace
Time-ordered record of the application run.