Validating Sentinel detections in CI/CD
A rule that deploys cleanly can still be wrong — bad KQL, missing MITRE tags, or a query that never fires. Validation is how the pipeline catches that before the SOC does. The trick: borrow the tiers you already know from application security.
Why validate at all
In the portal-only world, a broken rule is found by a human noticing the SOC queue went quiet — or flooded. Detection-as-Code moves that discovery left: the pipeline rejects a bad rule on the pull request, before it can reach a workspace. But "bad" has three different meanings, and each needs a different kind of check.
The three tiers, mapped to app-sec
If you've seen SAST, DAST, and a compiler, you already have the mental model. Each maps cleanly onto a tier of detection validation:
| Tier | Catches | Needs a workspace? |
|---|---|---|
| Static lint (≈ SAST) | missing severity/MITRE/entities, bad naming, empty query, hardcoded workspace IDs | no |
| Syntax + schema (≈ compile) | KQL that doesn't parse, or references a table/column that doesn't exist | read-only |
| Functional test (≈ DAST) | a rule that deploys fine but never fires, or fires on everything | yes (dev) |
Where each check sits in the pipeline
The static tiers are the CI gate: they run on the pull request and must pass before the merge button unlocks. The functional test is a CD stage: it runs after merge, against the dev workspace, because it needs a live Sentinel to execute against.
validate job triggered on pull_request, with steps that run before any deploy job — and branch protection requires that job to pass before merge. The deploy-and-fire test is a separate job on push to develop.How to run each check
Tier 1 — static lint (no Azure). A script parses the exported ARM/YAML and asserts the metadata is complete. This is the fast gate on every PR:
python3 scripts/lint_rule.py rules/*.json
# asserts: displayName, valid severity, non-empty query, MITRE
# tactics + techniques, entityMappings, naming convention, and
# no hardcoded workspace GUIDs. Non-zero exit fails the PR.Tier 2 — KQL syntax + schema. Run the query against a workspace over a tiny window with | take 0 appended: no rows return, but the service still parses the KQL and resolves every table and column against the real schema.
az monitor log-analytics query \
--workspace <dev-workspace-guid> \
--analytics-query "SigninLogs | where ResultType == 50126 | take 0"
# fails if the KQL is malformed or a column doesn't exist.
# Fully offline alternative: parse with the Kusto.Language library.Tier 3 — functional test (needs dev). Seed the workspace with sample logs containing a known-malicious pattern, run the detection, and assert it returns the planted true positive and nothing benign:
Detection
| summarize hits = count(),
true_positive = countif(UserPrincipalName == "attacker.target@contoso.com"),
false_positive = countif(UserPrincipalName != "attacker.target@contoso.com")
// pass when true_positive == 1 and false_positive == 0Clone the repo and open examples/detection-as-code-cicd — the analytic rule (ARM + Terraform), the validation scripts, and the GitHub Actions pipeline from this article, ready to run in your own lab. Try it now: python3 scripts/lint_rule.py rules/*.json
Why this order — shift-left
The ordering isn't arbitrary; it's the same economics as unit tests vs integration tests. Cheap, fast checks that need nothing run on every commit and block the merge. The expensive check — deploying to a live workspace and querying it — runs once, after merge, on the agreed-upon result.
- Static lint: milliseconds, no dependencies → every PR.
- Syntax/schema: seconds, read-only auth → every PR.
- Functional test: minutes, write access, live data → once, post-merge.
Push the cheap checks as far left as you can (even a pre-commit hook), and reserve the workspace-dependent test for the stage that already has a workspace. A human still owns the one thing no check can judge: whether the detection logic is actually good.
Try it
The companion example above is a complete, runnable version of everything here — one brute-force rule in ARM and Terraform, the lint and KQL-check scripts, the functional-test recipe, and the GitHub Actions workflow that wires the tiers together. Break a field in the rule JSON and watch the linter reject it: that's the CI gate doing its job.