KSKS Security Research
Learning map / Session 4 / Scenes 14–20

WORKSHOP 03 · EVALUATE

A demo proves possibility.Evaluation proves repeatability.

Combine deterministic tests, human review, model-based evaluation, adversarial cases, operational thresholds, and rollback evidence into a ship decision.

evaluationsLLM-as-judgered teamingrelease readiness
Learning guide
Level
Applied
Reading time
14 min
Presentation
Session 4
Progress
3 of 3

01 · Mental model

Evaluate the behavior, controls, and operating envelope

Deterministic tests validate schemas, permissions, required fields, and known invariants. Human annotation captures expert judgment and ambiguous quality. Model-based evaluators scale rubric application but must be calibrated. Adversarial datasets test injection, overreach, leakage, and unsafe persistence. Operational gates cover latency, cost, error rate, observability, ownership, and rollback.

The release decision belongs to accountable humans and policy, supported by versioned datasets, evaluator definitions, thresholds, exceptions, and evidence from the exact candidate being promoted.

02 · Visual explanation

01Schemamechanical validity
02Code evalknown invariants
03Human reviewexpert judgment
04LLM judgescaled rubric
05Red teamadversarial behavior
The evaluation ladderConfidence grows from multiple evidence types; no single score proves safety.

03 · Compare and decide

Use complementary evidence

Decision lensEvaluation methodBest use
DeterministicSchema, exact match, policy, tool and calculation checksFast gates with reproducible pass/fail
HumanNuance, impact, usefulness, and expert acceptanceGold labels, calibration, consequential review
LLM-as-judgeRubric-based semantic comparison at scaleRegression detection after calibration
AdversarialInjection, leakage, overreach, denial, and unsafe persistenceSecurity behavior and safe failure

04 · Cybersecurity example

Release gate for the review agent

A new model and skill version claim better architecture findings.

01

Run the frozen representative dataset.

02

Compare quality, security, cost, and latency to baseline.

03

Review regressions and adversarial failures.

04

Promote only the exact candidate with rollback to the previous version.

Outcome: The upgrade decision is based on repeatable evidence instead of a stronger-looking demonstration.

05 · What to remember

The 60-second recall

01

Freeze representative datasets and version every evaluator.

02

Calibrate model-based judges against expert labels.

03

Release readiness includes operations, security, ownership, and rollback—not only answer quality.

Teach-back prompt: Explain this concept to a teammate using the diagram, then name one failure mode and the control that stops it.

06 · Questions people ask

FAQ

No. It can scale a calibrated rubric, but it inherits model limitations and should be checked against expert judgments, especially for consequential decisions.

07 · Primary sources

Continue with authoritative guidance