Skip to main content
Version: 0.1 (next)

SRE daily audit

Audits every Product's zones each morning, opens issues for what is wrong, and writes a dated audit document to the internal docs.

Namesre-daily-audit
CategoryScheduled
Enabled by defaultYes
Budgetup to $4.00 per run, 60 turns
Catalogv0.2.0
Used inDaily audit

What it does​

The team starts each day knowing the true state of every zone. The SRE daily audit is accountable for finding Degraded or OutOfSync Applications, error budget burn, stale pins and expiring certificates before they become outages, opening one actionable issue per problem, and keeping a written daily record in the internal docs.

Step by step:

  1. List every Application in the gitops repo's registry and record its sync status (Synced, OutOfSync) and health (Healthy, Progressing, Degraded, Suspended, Missing). Anything not Synced and Healthy for more than an hour is a finding.
  2. Compute error budget burn per SLO for the last 24 hours and the last 7 days. Flag any SLO burning faster than its budget allows.
  3. Compare each zone's pin (tag@sha256) against the latest built image for its Product. Flag pins more than 14 days behind the Release environment, and Pre-release zones that have not been promoted in 7 days.
  4. List TLS certificates, cert-manager Certificates and ExternalSecrets. Flag certificates expiring within 21 days and ExternalSecrets that failed their last refresh.
  5. Check node pressure, PersistentVolume usage above 80 percent, Pods in CrashLoopBackOff, and restarts above baseline.
  6. Search open issues before filing. Open one issue per new problem with the zone, the Application, the evidence and a suggested fix; comment on existing issues when the problem persists.
  7. Write the daily audit document to docs/audits/.md in the internal docs repo, with a summary table and links to each issue, and open a pull request for it.
  8. Finish with exactly one verdict: HEALTHY (no findings), ATTENTION (findings filed as issues), or DEGRADED (a Release environment is Degraded or an SLO budget is exhausted, and it was escalated).

When it runs​

  • On a schedule, 0 6 * * *. Every day at 06:00 UTC, as part of the daily-audit AgentWorkflow.

What it reads​

SourceWhat it uses it for
clusterArgo CD Applications, Pods, nodes, PersistentVolumes, Certificates and ExternalSecrets in every zone.
metricsSLO error ratios, latency and saturation metrics from the metrics stack.
repoThe gitops repo registry, including every zone's current pin.
issuesOpen issues labeled sre, so findings are not filed twice.
previous-reportYesterday's audit document, to report what changed.

What it produces​

  • A dated audit document in the internal docs, submitted as a pull request.
  • One issue per new problem, labeled sre, with evidence and a suggested fix.
  • Comments on existing issues whose problem persists.

How it proves it​

Every run attaches this evidence to its AgentWorkflowRun step.

EvidenceWhat it showsRequired
DocumentThe daily audit document with a summary table, per-zone status and links to issues.Yes
ReportMachine-readable table of every Application with sync status, health and pin.Yes
MetricThe SLO burn queries and values behind any error budget finding.No

Success criteria​

A run succeeds only when every statement holds.

  • An audit document for the date exists in the internal docs repo as a pull request.
  • Every Application not Synced and Healthy appears in the audit with a linked issue.
  • No problem is filed as a duplicate of an open issue.
  • Every issue names the zone, the Application and the evidence, and suggests a fix.
  • Certificates expiring within 21 days are all listed.
  • The audit states what changed since the previous day's audit.

Guardrails​

  • Never sync, refresh with prune, roll back or delete an Application. Report only.
  • Never edit a pin or open a promotion; the promotion flow owns that.
  • Never change SLO targets or alert rules to clear a finding.
  • Never read or print Secret values.
  • Never merge the audit pull request; a human merges it.
  • Do not file issues for Progressing Applications younger than one hour.

Permissions​

Deny wins over allow.

Tools allowedRead, Grep, Glob, Write, Edit, Bash(kubectl get:*), Bash(kubectl describe:*), Bash(kubectl logs:*), Bash(argocd app list:*), Bash(argocd app get:*), Bash(promtool query:*), Bash(git checkout -b:*), Bash(git add:*), Bash(git commit:*), Bash(git push:*), Bash(gh pr create:*), Bash(gh issue:*)
Tools deniedWebSearch, Bash(argocd app sync:*), Bash(argocd app delete:*), Bash(argocd app rollback:*), Bash(kubectl delete:*), Bash(kubectl apply:*), Bash(kubectl get secret:*), Bash(git push --force:*), Bash(gh pr merge:*)
Git scopescontents:read, contents:write, pull_requests:write, issues:write
Cluster verbsget, list, logs
Networkallowlist
Egress allowlistapi.github.com, github.com, victoria-metrics.infrared.svc, argocd-server.argocd.svc
May merge its own pull requestsNo

When it hands off to a human​

It dead-letters the work to @platform/sre if it has not finished after 30m, or as soon as any of these is true:

  • An Application in a Release environment is Degraded or Missing.
  • An SLO has exhausted its monthly error budget.
  • A certificate expires within 7 days.
  • Metrics or Argo CD were unreachable for the whole run.

Verdicts​

Every run ends with exactly one of these verdicts:

  • HEALTHY
  • ATTENTION
  • DEGRADED

Opinions​

Opinions are the org’s editable guidance for this role. Each one can be edited or switched off in the Infrared UI; an edited opinion is marked as the org’s own.

How to judge error budget burn​

slo-burn-alerting · origin catalog

Use multi-window burn rates. A 1-hour burn rate above 14.4 or a 6-hour burn rate above 6 is urgent. A 3-day burn rate above 1 is worth an issue. Always state the SLO target and the window.

When a pin is stale​

stale-pin-threshold · origin catalog

A Release environment pin more than 14 days behind the latest image that passed Pre-release is stale. Stale pins accumulate risk; the issue should propose the promotion, not perform it.

Audit document format​

audit-format · origin catalog

Start with a three-line summary, then a table of zones with sync, health, pin age and open issues. Keep the document under two screens. Link, do not paste, long logs.

One problem, one issue​

one-problem-one-issue · origin catalog

File one issue per distinct problem, titled with the zone and the symptom. Update the existing issue while it stays open instead of filing a new one each day.