Skip to content
Paul Marinos
Menu

Program Metrics

Why MTTD and MTTR mislead, the difference between coverage and capability, and a small set of measures that survive being reported quarterly.

Detection programs are measured because they cost money, and the measures chosen tend to be the ones that are easy to compute rather than the ones that describe whether the program works. That has a specific cost: teams optimize what is reported.

Mean time to detect and mean time to respond are the default pair, and both are weaker than they look.

They are means over skewed distributions. Detection times are heavily long-tailed — most incidents found quickly, a few after months. The mean sits in a region where nothing actually happens, and one long incident moves it more than a quarter of good work. Report medians and percentiles, as with any skewed security distribution.

They only include incidents you detected. The undetected intrusion contributes nothing to MTTD. A program that misses everything subtle has an excellent MTTD, computed over the loud incidents it caught. This is survivorship bias determining a headline number.

The clock start is ambiguous. Detection time from first malicious action, from first observable evidence, or from the alert firing? These differ by weeks, and the choice is rarely documented — which makes cross-period comparison unreliable.

They are improvable without improving. Narrow the definition of “incident” and both numbers get better.

None of that makes them useless. Tracked as distributions, with a documented clock and alongside a measure of what you missed, they are informative. Reported as two means on a slide, they are theater.

The distinction most programs elide.

Coverage is a claim about existence: rules exist for these techniques. It is cheap to improve and easy to overstate — a rule that fires only where an agent is deployed, or has never been tested, still colors a Navigator cell green.

Capability is a claim about outcome: when this technique is executed here, we detect it, in time, and an analyst can act. It is expensive to establish, because it requires emulation, and it is the only one that predicts incident outcomes.

Report them separately, and never let coverage stand in for capability. The honest artifact has three states — untested, tested, validated by emulation — and the interesting number is the ratio between them.

A small set, chosen because they resist gaming and describe the program rather than its paperwork:

Measure What it tells you
Alert true-positive rate, per rule and overall Whether the queue is workable at all
Validated coverage — techniques confirmed by emulation Capability, not existence
Detection health — rules silent, drifting, or without owners The decay you’re carrying
Time to answer “are we affected by X?” Data accessibility, tested the way leadership tests it
Backlog age and intake source mix Whether high-quality intake is reaching the backlog
Coverage gaps by log tier What you structurally cannot see

Two properties make this set work. Each maps to an action — a bad number tells you what to do. And several get worse when the program is neglected, which is what makes them honest: a metric that only improves is a metric being managed rather than measured.

The audience problem applies with force here.

Engineering wants per-rule health and backlog. Leadership wants trend and consequence: are we more capable than last quarter, where did the investment go, what remains uncovered and what would it cost. Boards want risk framing, not technique names.

Keep the measures stable across quarters. A dashboard whose metrics change every reporting cycle is measuring the reporting cycle. If a measure must be replaced, show both for a period rather than silently switching.

This is also where detection metrics meet GRC: continuous control monitoring and detection health are the same telemetry with different consumers, and producing the evidence once is cheaper than assembling it twice.

Two shapes dominate. Embedded — detection engineers inside the SOC — gives tight feedback from analysts and tends toward reactive, ticket-driven work. Separate — a detection engineering function with its own backlog — protects engineering time and risks drifting from operational reality.

The functional difference is usually whether detection engineers are on the analyst rotation, not the org chart. Engineers who work the queue they create build different rules.

The distinguishing feature of a workable metrics program is that nobody assembles it by hand. Each measure in the table above is a query against a system of record — rule metadata in the detection repo, dispositions in the case system, liveness from pipeline health — generated on a schedule into a report whose layout doesn’t change between quarters. Hand-built slides drift toward what looks good; regenerated ones can only be improved by improving the program. The same queries export their results as evidence for GRC’s monitoring controls, which is the produce-once principle made concrete.

The measurement discipline is statistics, and the presentation discipline is visualization — where traffic-light dashboards do their damage. Capability claims rest on testing and emulation. And the evidence overlaps with GRC continuous control monitoring closely enough that it should be generated once.

Graph View

Spotted an error on this page? Report it.