From chat.
To repeatable capability.
How LLMs, projects, skills, tools and agents change the way cybersecurity work gets delivered.
security task
An LLM predicts a useful continuation.
It transforms input into tokens, uses learned patterns plus current context, and generates the next token repeatedly.
OpenAI API conceptsIllustrative probabilities. Context changes the distribution.
The model sees tokens.
The application must preserve trust.
Role, output contract, prohibited actions and escalation rules.
What the user wants accomplished in this interaction.
Documents, retrieved passages, prior decisions and run state.
Structured results returned by approved capabilities.
Email, web pages and uploads may contain hostile instructions.
Chat solves a moment.
A project preserves working context.
Cold start
- Re-explain scope and standards
- Manually attach the right files
- Copy results into the next task
- Quality depends on the operator
Persistent context
- Shared instructions and reference files
- Conversation continuity
- Less repeated setup
- Still human-directed, step by step
Add structure before autonomy.
Chat
One conversation
Project
Persistent workspace
Skill
Packaged expertise
Tool
Callable action
Agent
Bounded loop
Workflow
Governed execution
Multi-agent
Specialist system
A good first candidate is bounded, repeated and reviewable.
A skill packages the task—not merely the prompt.
A versioned bundle can carry instructions, specialist references, output schemas, deterministic scripts and reusable assets.
OpenAI Skills documentation├── SKILL.md # trigger + workflow
├── references/
│ ├── threat-modeling.md
│ └── cloud-security.md
├── schemas/
│ └── findings.schema.json
├── scripts/
│ ├── iac_scan.py
│ └── report_html.py
└── assets/
└── report-template/
Inside a production cybersecurity skill.
SKILL.md · the operating contract
Defines when the skill applies, its workflow, boundaries, evidence rules and definition of done.
It routes work. It does not contain every reference inline.
Expert judgment and deterministic code do different jobs.
Interpret context
- Identify trust boundaries
- Reason about abuse paths
- Explain business impact
- Prioritize remediation
Guarantee the floor
- Validate JSON Schema
- Scan known IaC patterns
- Calculate and transform
- Render stable documents
Tools, MCP and retrieval solve different problems.
Requests a named action
Arguments follow a schema.
validate_findings()
Code authenticates, authorizes, executes and returns a structured result.
Pass / fail
The model receives data—not implicit authority.
Agent host
Discovers approved capabilities.
Model Context Protocol
Standard messages for tools, resources and prompts.
Approved servers
MS Learn · ArchStudio · internal knowledge
Question
What evidence is relevant?
Retriever
Searches an indexed corpus.
Grounded context
Relevant passages + citations reach the model.
agent
Reason. Act. Observe.
Stop safely.
Tool allowlist · iteration cap · token budget · timeout
Structured input · structured tool arguments · validated output
Needs input · needs review · blocked · failed safely
The difference is controlled action—not personality.
Choose who owns the final answer.
Best when one component must combine specialist outputs and enforce shared controls.
Best when a specialist should directly own the interaction after routing.
Best when sequence, approvals and failure behavior must be explicit and testable.
More agents do not automatically create more value.
- One prompt and a few tools solve the task
- All steps use the same context
- There is one definition of done
- Specialists would only repeat each other
- Expert roles have different evidence and instructions
- Work can proceed independently
- Hand-offs are typed artifacts
- An independent reviewer adds real assurance
Match the architecture to the task.
Adopt capability in measured increments.
Select one repeated, bounded task.
Baseline quality, time and rework.
Package a reviewed skill and schema.
Add read-only tools and deterministic checks.
Add gates, traces and adversarial tests.
Scale only when evidence supports it.
Package expertise.
Bound action. Measure value.
Repeated task
Not a fascination with agents.
Skill before autonomy
Make expertise versioned and reviewable.
Governed workflow
Only when controls and evidence grow with capability.