Measure detection quality as an evidence chain: telemetry prerequisites and source health, validation results by detection version, known blind spots, alert context and duplication, investigation outcomes, tuning changes and retest age. A false-positive rate or ATT&CK mapping alone is not sufficient evidence of detection effectiveness.
External reference
NIST SP 800-55 Vol. 1 supplies the measurement principles. The detection-specific evidence chain is Cybatar-authored and links to the existing Cybatar Detection Engineering and Detection Evidence libraries.
Primary source →Measures to retain
Prerequisite coverage
Track whether required telemetry is present, parsable, timely and sufficiently complete for each material detection objective.
Validation by version
Record test date, test case, expected result, actual result and known limitations for the specific deployed detection version.
Triage outcome quality
Use dispositions, incident escalation, duplicate clusters, missing context and reopened investigations as feedback on usefulness.
Blind-spot inventory
Keep known missing sources, unsupported conditions and untested scenarios visible rather than reporting only successful coverage.
Retest and tuning age
Measure how long material detections have gone without validation after logic, telemetry or environmental change.
Operating method
Define the detection objective
State the behaviour or condition and required telemetry before measuring outcomes.
Measure prerequisites
Coverage and source quality are part of detection quality.
Measure tests and operations
Combine controlled validation evidence with actual triage and incident outcomes.
Measure known limits
A useful report makes blind spots and stale tests visible.
Measurement anti-patterns
Relevant Cybatar sources
Claim boundary
Detection metrics demonstrate only the stated scope and evidence. They do not prove complete ATT&CK coverage, universal attack detection, absence of false negatives or effectiveness under conditions that were not observed or tested.
Cybatar publishes measurement and reporting guidance as a first-party operating model. Examples are not universal benchmarks, regulatory thresholds, promises of security outcomes or evidence that a particular deployment is effective.