Cybersecurity × Agentic AI
Opening
⌂00:00
Session 1 · Foundations

From chat.
To repeatable capability.

How LLMs, projects, skills, tools and agents change the way cybersecurity work gets delivered.

60 minutesNo engineering prerequisiteInteractive
Repeated
security task
Prompt
Skill
Agent
Evidence
Deliverable
space / click to reveal
Audience pulse

Where does AI help you today?

Select the closest answer. There is no wrong starting point.
01Start with a real task
02Package what repeats
03Add agency only when earned
Mental model

An LLM predicts a useful continuation.

It transforms input into tokens, uses learned patterns plus current context, and generates the next token repeatedly.

OpenAI API concepts
Choose the next token
The suspicious sign-in should be _____

Illustrative probabilities. Context changes the distribution.

Prompt anatomy

The model sees tokens.
The application must preserve trust.

System
Operating boundaries

Role, output contract, prohibited actions and escalation rules.

User
Requested outcome

What the user wants accomplished in this interaction.

Context
Evidence and history

Documents, retrieved passages, prior decisions and run state.

Tools
Observed results

Structured results returned by approved capabilities.

Untrusted data
Never implicit authority

Email, web pages and uploads may contain hostile instructions.

Official OpenAI prompting guidance
The first upgrade

Chat solves a moment.
A project preserves working context.

Individual chat

Cold start

  • Re-explain scope and standards
  • Manually attach the right files
  • Copy results into the next task
  • Quality depends on the operator
Project / workspace

Persistent context

  • Shared instructions and reference files
  • Conversation continuity
  • Less repeated setup
  • Still human-directed, step by step
A workspace remembers the work. It does not yet govern the work.
The maturity ladder

Add structure before autonomy.

01

Chat

One conversation

02

Project

Persistent workspace

03

Skill

Packaged expertise

04

Tool

Callable action

05

Agent

Bounded loop

06

Workflow

Governed execution

07

Multi-agent

Specialist system

Best for exploration, explanation and one-off analysis. Human provides context and owns every transition.
Candidate scorecard

A good first candidate is bounded, repeated and reviewable.

Candidate
Repeats
Clear output
Easy to review
Safe scope
Architecture review draft
High
Findings schema
Architect validates
Read-only
KQL query drafting
High
Query + metadata
Needs dry-run
Dev workspace
Deploy controls to production
Variable
Many side effects
Hard to reverse
High impact
Start: architecture reviewStart: document packStart: schema validationLater: production writes
The reusable unit

A skill packages the task—not merely the prompt.

A versioned bundle can carry instructions, specialist references, output schemas, deterministic scripts and reusable assets.

OpenAI Skills documentation
security-architect/
├── SKILL.md # trigger + workflow
├── references/
│  ├── threat-modeling.md
│  └── cloud-security.md
├── schemas/
│  └── findings.schema.json
├── scripts/
│  ├── iac_scan.py
│  └── report_html.py
└── assets/
   └── report-template/
Interactive X-ray

Inside a production cybersecurity skill.

Selected component

SKILL.md · the operating contract

Defines when the skill applies, its workflow, boundaries, evidence rules and definition of done.

It routes work. It does not contain every reference inline.

Hybrid execution

Expert judgment and deterministic code do different jobs.

Agent judgment

Interpret context

  • Identify trust boundaries
  • Reason about abuse paths
  • Explain business impact
  • Prioritize remediation
Deterministic scripts

Guarantee the floor

  • Validate JSON Schema
  • Scan known IaC patterns
  • Calculate and transform
  • Render stable documents
Assessmodel
Validatecode
Reviewhuman / policy
Rendercode
Capability map

Tools, MCP and retrieval solve different problems.

Model

Requests a named action

Arguments follow a schema.

→
Application

validate_findings()

Code authenticates, authorizes, executes and returns a structured result.

→
Result

Pass / fail

The model receives data—not implicit authority.

OpenAI MCP and connectors guide
PLANnext step
ACTrequest tool
OBSERVEread result
CHECKdone / retry
Bounded
agent
The loop

Reason. Act. Observe.
Stop safely.

Bounds

Tool allowlist · iteration cap · token budget · timeout

Contracts

Structured input · structured tool arguments · validated output

Escalation

Needs input · needs review · blocked · failed safely

OpenAI Agents SDK overview
Do not anthropomorphize

The difference is controlled action—not personality.

Property
Chat assistant
Agent
Governed workflow
Primary behavior
Respond
Choose and act
Execute controlled stages
State
Conversation
Working memory
Durable run state
Tools
Optional
Core capability
Scoped per node
Control
Human drives turns
Runtime limits
Gates, policy, evidence
Agency starts when the system may choose an action. Governance starts when that choice is controlled and recorded.
Multi-agent patterns

Choose who owns the final answer.

Managerowns response
Architectagent as tool
Complianceagent as tool
Detectionagent as tool

Best when one component must combine specialist outputs and enforce shared controls.

OpenAI orchestration and handoffs
Complexity gate

More agents do not automatically create more value.

Stay single-agent when
  • One prompt and a few tools solve the task
  • All steps use the same context
  • There is one definition of done
  • Specialists would only repeat each other
Consider multi-agent when
  • Expert roles have different evidence and instructions
  • Work can proceed independently
  • Hand-offs are typed artifacts
  • An independent reviewer adds real assurance
+Specialization
+Parallel work
−Coordination cost
−Larger attack surface
Decision practice

Match the architecture to the task.

Task
Start with
Why
Upgrade trigger
Explain one Sentinel alert
Chat / project
Human owns context
Repeated triage procedure
Architecture review report
Skill + tools
Stable method and output
Multi-step evidence collection
Greenfield delivery pack
Governed workflow
Parallel specialists + artifacts
Independent assurance
Deploy controls automatically
Later-stage agent
High side effects
Policy, simulation, approval, rollback
A practical 90-day path

Adopt capability in measured increments.

Week 1

Select one repeated, bounded task.

Week 2

Baseline quality, time and rework.

Week 3–4

Package a reviewed skill and schema.

Month 2

Add read-only tools and deterministic checks.

Month 3

Add gates, traces and adversarial tests.

Decision

Scale only when evidence supports it.

Measure qualityMeasure review effortMeasure defects caughtMeasure safe failure
Session 1 synthesis

Package expertise.
Bound action. Measure value.

01 · Start

Repeated task

Not a fascination with agents.

02 · Build

Skill before autonomy

Make expertise versioned and reviewable.

03 · Scale

Governed workflow

Only when controls and evidence grow with capability.

Next: what changes when AI can read untrusted data and take action?

Presenter controls

Space / →Reveal or advance←Previous reveal or sceneMSession mapNSpeaker notesRReferencesFFull screen