Skip to main content
Version: 0.1 (next)

SLOs, incidents and DORA

Infrared watches a Product through what it already records: its SLOs in VictoriaMetrics (installed by the gitops template), its Releases, its Changes and its incidents. Observability shows it all for one Product at a time.

SLOs​

Add objectives to the Product. Each one counts good and bad events with one indicator, over a window, in the zones you name (every delivery zone by default).

spec:
slos:
- name: availability
description: Requests that don't fail with a 5xx.
objective: "99.5" # percent of good events
window: 168h # must fit in the metrics retention (14 days by default)
indicator:
availability:
metric: http_requests_total # a request counter
errors: code=~"5.." # the default
filter: route!="/metrics" # optional label matchers
- name: latency
description: Requests answered within 250 ms.
objective: "99"
indicator:
latency:
histogram: http_request_duration_seconds # _bucket and _count are read
threshold: "0.25" # must be one of the histogram's le buckets

For each SLO and zone, Infrared keeps a VMRule, infrared-slo-<product>, in the org namespace with:

  • recording rules for the error ratio over 5m, 30m, 1h, 6h and the window (infrared:slo_errors:ratio_*);
  • a fast burn alert (critical): the 1-hour and 5-minute error ratios above 14.4 times the budget, which would spend the whole budget in about a fourteenth of the window;
  • a slow burn alert (warning): the 6-hour and 30-minute ratios above 6 times the budget;
  • a zone down alert (critical) per zone: no available pods for 3 minutes.

The Product's status says whether the rules are in place (status.observability). The SLO table shows each objective's attainment, how much error budget is left, the burn rates over 1 and 6 hours, and a state: meeting the objective, burning budget, budget spent, or no traffic yet.

Alerts open incidents​

Infrared routes the org's alerts from Alertmanager to its API (a VMAlertmanagerConfig named infrared-incidents, authenticated with the token in the Secret infrared-alert-webhook). A firing alert opens an incident: one per alert, with its severity, zone and a link to the rule. When the alert clears, Infrared records it and resolves the incident.

On an incident's page you Acknowledge it (you're on it), add notes as you work, and Resolve it with what fixed it. Everything lands on the incident's timeline. Open an incident records something no alert caught.

DORA metrics​

Infrared computes the four DORA metrics for the last 30 days from what it runs, and rates each Elite, High, Medium or Low:

MetricHow Infrared measures it
Deployment frequencyReleases that reached a release zone Healthy.
Lead time for changesMedian time from a Change's merge (or a Release's start, when it was started by hand) to running Healthy in a release zone.
Change failure rateDeployments to a release zone that failed to promote, or were followed by a warning or critical incident there within 24 hours.
Time to restore serviceMedian time from opening to resolving the Product's warning and critical incidents.

Status page​

Every Product has a status page at /status/<org>/<product>: the state of each zone (operational, degraded or down, from open incidents and burning SLOs), its objectives, and its incidents with their updates, without who did what. Signed-in users can always see it. To show it to anyone, publish it:

spec:
statusPage:
public: true
zones: [checkout-prod] # default: the release zones

API​

EndpointWhat it returns
GET /v1/orgs/<org>/products/<product>/slosEach SLO in each zone.
GET /v1/orgs/<org>/products/<product>/dora?days=30The four DORA metrics.
GET /v1/orgs/<org>/incidents, POST .../incidents/<name>/acknowledge, .../resolve, .../notesIncidents and their lifecycle.
GET /v1/status/<org>/<product>The public status page, when published.