KSKS Security Research
Home / DevOps & automation

Validating Sentinel detections in CI/CD

Detection-as-Code for Sentinel· part 2 of 2
SentinelKQLCI/CDSAST / DAST

A rule that deploys cleanly can still be wrong — bad KQL, missing MITRE tags, or a query that never fires. Validation is how the pipeline catches that before the SOC does. The trick: borrow the tiers you already know from application security.

TL;DR
KQL validation is three tiers, not one. Static lint (like SAST) and syntax/schema check (like compile) run in CI on the pull request — the gate that blocks a bad merge. Functional testing (like DAST) runs after merge, against a dev workspace with sample data. Cheap checks gate early; the expensive one runs once, later.

Why validate at all

In the portal-only world, a broken rule is found by a human noticing the SOC queue went quiet — or flooded. Detection-as-Code moves that discovery left: the pipeline rejects a bad rule on the pull request, before it can reach a workspace. But "bad" has three different meanings, and each needs a different kind of check.

The three tiers, mapped to app-sec

If you've seen SAST, DAST, and a compiler, you already have the mental model. Each maps cleanly onto a tier of detection validation:

the checkapp-sec analogpipeline stageStatic lintstructure & metadata≈ SASTCI · on the PRKQL syntax + schemaresolve vs tables≈ compileCI · read-onlyFunctional testdoes it fire? noisy?≈ DASTCD · after merge
Static lint and a syntax/schema check are cheap and run on the PR; the functional test needs a live workspace and runs after merge.
TierCatchesNeeds a workspace?
Static lint (≈ SAST)missing severity/MITRE/entities, bad naming, empty query, hardcoded workspace IDsno
Syntax + schema (≈ compile)KQL that doesn't parse, or references a table/column that doesn't existread-only
Functional test (≈ DAST)a rule that deploys fine but never fires, or fires on everythingyes (dev)

Where each check sits in the pipeline

The static tiers are the CI gate: they run on the pull request and must pass before the merge button unlocks. The functional test is a CD stage: it runs after merge, against the dev workspace, because it needs a live Sentinel to execute against.

Open pull requestfeature branch → PRCI — static checkssyntax · metadata · validate≈ SAST + lintPeer reviewlogic & false positives≈ code reviewMerge → deploy to devauto-deploy to dev workspaceCD — dynamic testdoes it fire? noisy?≈ DASTPromote to prodPR to main → prod deployCI gateCD stage
Static checks gate the merge; a human judges the logic; the does-it-fire test runs after merge against dev.
One-line answer to 'where in CI?'
A validate job triggered on pull_request, with steps that run before any deploy job — and branch protection requires that job to pass before merge. The deploy-and-fire test is a separate job on push to develop.

How to run each check

Tier 1 — static lint (no Azure). A script parses the exported ARM/YAML and asserts the metadata is complete. This is the fast gate on every PR:

metadata lint — runs on the PR, no workspace
python3 scripts/lint_rule.py rules/*.json
# asserts: displayName, valid severity, non-empty query, MITRE
# tactics + techniques, entityMappings, naming convention, and
# no hardcoded workspace GUIDs. Non-zero exit fails the PR.

Tier 2 — KQL syntax + schema. Run the query against a workspace over a tiny window with | take 0 appended: no rows return, but the service still parses the KQL and resolves every table and column against the real schema.

kql check — read-only, validates against real schema
az monitor log-analytics query \
  --workspace <dev-workspace-guid> \
  --analytics-query "SigninLogs | where ResultType == 50126 | take 0"
# fails if the KQL is malformed or a column doesn't exist.
# Fully offline alternative: parse with the Kusto.Language library.

Tier 3 — functional test (needs dev). Seed the workspace with sample logs containing a known-malicious pattern, run the detection, and assert it returns the planted true positive and nothing benign:

functional assertion — runs in CD, after merge
Detection
| summarize hits = count(),
    true_positive  = countif(UserPrincipalName == "attacker.target@contoso.com"),
    false_positive = countif(UserPrincipalName != "attacker.target@contoso.com")
// pass when true_positive == 1 and false_positive == 0
Runnable companion example

Clone the repo and open examples/detection-as-code-cicd — the analytic rule (ARM + Terraform), the validation scripts, and the GitHub Actions pipeline from this article, ready to run in your own lab. Try it now: python3 scripts/lint_rule.py rules/*.json

Why this order — shift-left

The ordering isn't arbitrary; it's the same economics as unit tests vs integration tests. Cheap, fast checks that need nothing run on every commit and block the merge. The expensive check — deploying to a live workspace and querying it — runs once, after merge, on the agreed-upon result.

  • Static lint: milliseconds, no dependencies → every PR.
  • Syntax/schema: seconds, read-only auth → every PR.
  • Functional test: minutes, write access, live data → once, post-merge.

Push the cheap checks as far left as you can (even a pre-commit hook), and reserve the workspace-dependent test for the stage that already has a workspace. A human still owns the one thing no check can judge: whether the detection logic is actually good.

Try it

The companion example above is a complete, runnable version of everything here — one brute-force rule in ARM and Terraform, the lint and KQL-check scripts, the functional-test recipe, and the GitHub Actions workflow that wires the tiers together. Break a field in the rule JSON and watch the linter reject it: that's the CI gate doing its job.