LLM observability · evaluation · SOC operations
See the agent.Defend the system.
Langfuse gives engineers and security teams execution-level evidence for LLM applications. A SIEM turns security signals into enterprise correlation, incidents and response. Production SOC operations need both roles to be explicit.
The decision in one sentence
Langfuse answers
What did the agent see, call, generate, score and cost—and how did that change across a session or release?
The SIEM answers
Is this behavior part of a wider attack, which assets and identities are affected, and what incident or response should follow?
What Langfuse is
Langfuse is an open-source LLM engineering platform. Its observability model captures traces containing model generations and non-LLM observations such as retrievals, tool calls and application steps. Sessions group related traces; users support aggregate analysis; scores attach quality or security judgments; evaluations compare behavior across datasets and releases; prompt management versions runtime prompts.
How it is built
In the documented architecture, SDK or OpenTelemetry events reach the Langfuse API, payloads are persisted to object storage and queued, workers process them asynchronously, ClickHouse serves high-volume trace analytics, and PostgreSQL stores transactional platform data. Cloud and self-hosted editions share the core product model but have different operational responsibilities.
The operating model
| Langfuse object | What it records | SOC interpretation |
|---|---|---|
| Trace | One end-to-end request or workflow | Investigation timeline and correlation anchor |
| Observation / span | Retrieval, tool, guardrail or application step | Which control ran, which tool acted, and what failed |
| Generation | Model input/output, model, usage, latency and cost | Model behavior, sensitive-data exposure and anomaly context |
| Session | Related traces across a conversation or workflow | Multi-turn attack path, persistence and escalation story |
| User | Pseudonymous user-level aggregation | Usage patterns and abuse investigation—when identity mapping is governed |
| Score | Numeric, categorical, boolean or text judgment | Prompt-injection risk, policy result, quality or human verdict |
How this helps a SOC team
Replay an investigation
Follow a session across prompts, retrieval, tools and model calls to understand why the agent took an action.
Expose unsafe behavior
Score prompt-injection indicators, authorization failures, sensitive-data detections and policy denials at trace or observation level.
Find operational anomalies
Investigate latency, cost, token spikes, error rates and tool-call patterns by user, session, release or prompt version.
Prove release quality
Run code checks, human annotation and LLM-as-judge evaluations against datasets before and after a change.
Langfuse is especially useful during AI-specific triage: reconstructing a multi-turn prompt-injection attempt, checking which retrieved content influenced the model, confirming whether a privileged tool actually ran, comparing guardrail results, or finding the prompt/model version behind a regression. Those are details a general SIEM usually does not model natively.
Langfuse vs SIEM
| Capability | Langfuse | SIEM |
|---|---|---|
| LLM prompts, generations and tool spans | Primary strength | Usually custom, flattened telemetry |
| Evaluation, annotation and prompt versions | Built for this workflow | Not the primary job |
| Identity, endpoint, network and cloud correlation | Limited to supplied context | Primary strength |
| Detection rules, incidents and case management | Scores and analysis, not full SOC case handling | Primary strength |
| Hunting and automated response | AI execution investigation | Enterprise hunting, automation and response |
| Recommended role | System of insight for agent behavior | System of record for security operations |
Recommended integration blueprint
Keep the rich, sanitized LLM execution record in Langfuse. Send a smaller, normalized security event to the SIEM when a decision matters: policy denial, suspicious prompt score, unauthorized tool request, sensitive-data detector result, unusual cost threshold, or production evaluation regression. Include correlation IDs and a controlled investigation link—not the full prompt or output by default.
Security safeguards before production
- Mask at source: redact secrets, personal data and sensitive security evidence before telemetry leaves the application.
- Pseudonymize identity: use stable internal identifiers and keep the identity-resolution boundary governed.
- Separate environments: isolate development, test and production projects, keys, access and retention.
- Limit access: apply role-based access, avoid unsafe public trace sharing, rotate keys and audit administrative activity.
- Control data movement: decide cloud versus self-hosting from classification, residency, isolation and operating-capability requirements.
- Treat scores as evidence, not truth: calibrate thresholds, retain evaluator versions, and use human review for consequential decisions.
Current Cybersecurity Orchestrator implementation
A practical adoption roadmap
01 · Observe
Instrument traces, sessions and tool spans. Establish masking and environment boundaries.
02 · Evaluate
Add code scores, human annotation and release datasets. Baseline quality and risk.
03 · Operationalize
Emit high-signal security decisions to the SIEM, correlate incidents, and test response playbooks.
Fact-checked sources
Product capabilities and architecture were checked against current Langfuse documentation; SIEM responsibilities were checked against Microsoft Sentinel documentation. The telemetry split is a recommended architecture inferred from those documented capabilities, not a claim of a built-in integration.
Next in the path
Turn observability into a governed security operating model.