SLOs, incidents and DORA
Infrared watches a Product through what it already records: its SLOs in VictoriaMetrics (installed by the gitops template), its Releases, its Changes and its incidents. Observability shows it all for one Product at a time.
SLOs
Add objectives to the Product. Each one counts good and bad events with one indicator, over a window, in the zones you name (every delivery zone by default).
spec:
slos:
- name: availability
description: Requests that don't fail with a 5xx.
objective: "99.5" # percent of good events
window: 168h # must fit in the metrics retention (14 days by default)
indicator:
availability:
metric: http_requests_total # a request counter
errors: code=~"5.." # the default
filter: route!="/metrics" # optional label matchers
- name: latency
description: Requests answered within 250 ms.
objective: "99"
indicator:
latency:
histogram: http_request_duration_seconds # _bucket and _count are read
threshold: "0.25" # must be one of the histogram's le buckets
For each SLO and zone, Infrared keeps a VMRule, infrared-slo-<product>, in the org namespace with:
- recording rules for the error ratio over 5m, 30m, 1h, 6h and the window (
infrared:slo_errors:ratio_*); - a fast burn alert (critical): the 1-hour and 5-minute error ratios above 14.4 times the budget, which would spend the whole budget in about a fourteenth of the window;
- a slow burn alert (warning): the 6-hour and 30-minute ratios above 6 times the budget;
- a zone down alert (critical) per zone: no available pods for 3 minutes.
The Product's status says whether the rules are in place (status.observability). The SLO table shows each objective's attainment, how much error budget is left, the burn rates over 1 and 6 hours, and a state: meeting the objective, burning budget, budget spent, or no traffic yet.
Alerts open incidents
Infrared routes the org's alerts from Alertmanager to its API (a VMAlertmanagerConfig named infrared-incidents, authenticated with the token in the Secret infrared-alert-webhook). A firing alert opens an incident: one per alert, with its severity, zone and a link to the rule. When the alert clears, Infrared records it and resolves the incident.
On an incident's page you Acknowledge it (you're on it), add notes as you work, and Resolve it with what fixed it. Everything lands on the incident's timeline. Open an incident records something no alert caught.
DORA metrics
Infrared computes the four DORA metrics for the last 30 days from what it runs, and rates each Elite, High, Medium or Low:
| Metric | How Infrared measures it |
|---|---|
| Deployment frequency | Releases that reached a release zone Healthy. |
| Lead time for changes | Median time from a Change's merge (or a Release's start, when it was started by hand) to running Healthy in a release zone. |
| Change failure rate | Deployments to a release zone that failed to promote, or were followed by a warning or critical incident there within 24 hours. |
| Time to restore service | Median time from opening to resolving the Product's warning and critical incidents. |
Status page
Every Product has a status page at /status/<org>/<product>: the state of each zone (operational, degraded or down, from open incidents and burning SLOs), its objectives, and its incidents with their updates, without who did what. Signed-in users can always see it. To show it to anyone, publish it:
spec:
statusPage:
public: true
zones: [checkout-prod] # default: the release zones
API
| Endpoint | What it returns |
|---|---|
GET /v1/orgs/<org>/products/<product>/slos | Each SLO in each zone. |
GET /v1/orgs/<org>/products/<product>/dora?days=30 | The four DORA metrics. |
GET /v1/orgs/<org>/incidents, POST .../incidents/<name>/acknowledge, .../resolve, .../notes | Incidents and their lifecycle. |
GET /v1/status/<org>/<product> | The public status page, when published. |