KSKS Security Research
Home / Agentic AI & security / Operate

LLM observability · evaluation · SOC operations

See the agent.Defend the system.

Langfuse gives engineers and security teams execution-level evidence for LLM applications. A SIEM turns security signals into enterprise correlation, incidents and response. Production SOC operations need both roles to be explicit.

LangfuseLLM observabilityEvaluationSOC + SIEM

The decision in one sentence

Langfuse is not a SIEM
Use Langfuse to explain and evaluate agent behavior. Use a SIEM to correlate that behavior with identity, endpoint, network and cloud telemetry, create incidents, support hunting, and trigger response workflows.

Langfuse answers

What did the agent see, call, generate, score and cost—and how did that change across a session or release?

The SIEM answers

Is this behavior part of a wider attack, which assets and identities are affected, and what incident or response should follow?

What Langfuse is

Langfuse is an open-source LLM engineering platform. Its observability model captures traces containing model generations and non-LLM observations such as retrievals, tool calls and application steps. Sessions group related traces; users support aggregate analysis; scores attach quality or security judgments; evaluations compare behavior across datasets and releases; prompt management versions runtime prompts.

Security interpretation
This is application telemetry for probabilistic systems. It becomes security-relevant when you attach trusted identity context, policy decisions, tool authorization outcomes, guardrail results and release metadata.

How it is built

In the documented architecture, SDK or OpenTelemetry events reach the Langfuse API, payloads are persisted to object storage and queued, workers process them asynchronously, ClickHouse serves high-volume trace analytics, and PostgreSQL stores transactional platform data. Cloud and self-hosted editions share the core product model but have different operational responsibilities.

Langfuse event ingestion and storage architectureThe agent sends telemetry to the Langfuse API. Payloads are persisted and queued, a worker processes them, and data is stored in ClickHouse and PostgreSQL.Agent applicationLangfuse SDK / OTelLangfuse APIingestion + webObject storagedurable event payloadsRedis / Valkeywork queueWorkerasync processingClickHousetrace analyticsPostgreSQLtransactional dataSimplified logical view · Cloud and self-hosted deployments differ operationally

The operating model

Langfuse objectWhat it recordsSOC interpretation
TraceOne end-to-end request or workflowInvestigation timeline and correlation anchor
Observation / spanRetrieval, tool, guardrail or application stepWhich control ran, which tool acted, and what failed
GenerationModel input/output, model, usage, latency and costModel behavior, sensitive-data exposure and anomaly context
SessionRelated traces across a conversation or workflowMulti-turn attack path, persistence and escalation story
UserPseudonymous user-level aggregationUsage patterns and abuse investigation—when identity mapping is governed
ScoreNumeric, categorical, boolean or text judgmentPrompt-injection risk, policy result, quality or human verdict

How this helps a SOC team

Replay an investigation

Follow a session across prompts, retrieval, tools and model calls to understand why the agent took an action.

Expose unsafe behavior

Score prompt-injection indicators, authorization failures, sensitive-data detections and policy denials at trace or observation level.

Find operational anomalies

Investigate latency, cost, token spikes, error rates and tool-call patterns by user, session, release or prompt version.

Prove release quality

Run code checks, human annotation and LLM-as-judge evaluations against datasets before and after a change.

Langfuse is especially useful during AI-specific triage: reconstructing a multi-turn prompt-injection attempt, checking which retrieved content influenced the model, confirming whether a privileged tool actually ran, comparing guardrail results, or finding the prompt/model version behind a regression. Those are details a general SIEM usually does not model natively.

Langfuse vs SIEM

CapabilityLangfuseSIEM
LLM prompts, generations and tool spansPrimary strengthUsually custom, flattened telemetry
Evaluation, annotation and prompt versionsBuilt for this workflowNot the primary job
Identity, endpoint, network and cloud correlationLimited to supplied contextPrimary strength
Detection rules, incidents and case managementScores and analysis, not full SOC case handlingPrimary strength
Hunting and automated responseAI execution investigationEnterprise hunting, automation and response
Recommended roleSystem of insight for agent behaviorSystem of record for security operations

Recommended integration blueprint

Keep the rich, sanitized LLM execution record in Langfuse. Send a smaller, normalized security event to the SIEM when a decision matters: policy denial, suspicious prompt score, unauthorized tool request, sensitive-data detector result, unusual cost threshold, or production evaluation regression. Include correlation IDs and a controlled investigation link—not the full prompt or output by default.

Recommended Langfuse and SIEM telemetry splitThe agent sends sanitized detailed AI telemetry to Langfuse and minimal security events to a SIEM. The SOC correlates SIEM alerts with Langfuse investigation links.Cybersecurity agentidentity · tools · policyLangfusesanitized prompts · tool spanssessions · scores · evaluationsSIEMsecurity decisions · identitiescloud · endpoint · incidentsSOC analystcorrelate + respondDeep AI execution evidenceMinimal normalized security events
Implementation note
Do not assume a native Langfuse-to-SIEM connector. A practical design can dual-emit security decisions from the application, or use Langfuse's public and metrics APIs through a governed custom integration. Validate the connector, schema and retention model for your environment.

Security safeguards before production

  • Mask at source: redact secrets, personal data and sensitive security evidence before telemetry leaves the application.
  • Pseudonymize identity: use stable internal identifiers and keep the identity-resolution boundary governed.
  • Separate environments: isolate development, test and production projects, keys, access and retention.
  • Limit access: apply role-based access, avoid unsafe public trace sharing, rotate keys and audit administrative activity.
  • Control data movement: decide cloud versus self-hosting from classification, residency, isolation and operating-capability requirements.
  • Treat scores as evidence, not truth: calibrate thresholds, retain evaluator versions, and use human review for consequential decisions.

Current Cybersecurity Orchestrator implementation

Version-aware snapshot
The current project is pinned to Langfuse Python 3.15.0. The platform documentation evolves quickly, so examples should be checked against the installed SDK before implementation.
In useTracing, users, sessions and deterministic code scores for workflow and policy evidence.
Deliberate choiceRuntime prompts remain in version-controlled skills rather than Langfuse Prompt Management.
NextHuman annotation, LLM-as-judge evaluation and datasets/experiments for repeatable regression testing.
Audit boundaryLangfuse is observability and evaluation evidence; the project journal remains the authoritative action/audit record.

A practical adoption roadmap

01 · Observe

Instrument traces, sessions and tool spans. Establish masking and environment boundaries.

02 · Evaluate

Add code scores, human annotation and release datasets. Baseline quality and risk.

03 · Operationalize

Emit high-signal security decisions to the SIEM, correlate incidents, and test response playbooks.

Fact-checked sources

Product capabilities and architecture were checked against current Langfuse documentation; SIEM responsibilities were checked against Microsoft Sentinel documentation. The telemetry split is a recommended architecture inferred from those documented capabilities, not a claim of a built-in integration.

Next in the path

Turn observability into a governed security operating model.

Explore Langfuse docs