SRE daily audit
Audits every Product's zones each morning, opens issues for what is wrong, and writes a dated audit document to the internal docs.
| Name | sre-daily-audit |
| Category | Scheduled |
| Enabled by default | Yes |
| Budget | up to $4.00 per run, 60 turns |
| Catalog | v0.2.0 |
| Used in | Daily audit |
What it does
The team starts each day knowing the true state of every zone. The SRE daily audit is accountable for finding Degraded or OutOfSync Applications, error budget burn, stale pins and expiring certificates before they become outages, opening one actionable issue per problem, and keeping a written daily record in the internal docs.
Step by step:
- List every Application in the gitops repo's registry and record its sync status (Synced, OutOfSync) and health (Healthy, Progressing, Degraded, Suspended, Missing). Anything not Synced and Healthy for more than an hour is a finding.
- Compute error budget burn per SLO for the last 24 hours and the last 7 days. Flag any SLO burning faster than its budget allows.
- Compare each zone's pin (tag@sha256) against the latest built image for its Product. Flag pins more than 14 days behind the Release environment, and Pre-release zones that have not been promoted in 7 days.
- List TLS certificates, cert-manager Certificates and ExternalSecrets. Flag certificates expiring within 21 days and ExternalSecrets that failed their last refresh.
- Check node pressure, PersistentVolume usage above 80 percent, Pods in CrashLoopBackOff, and restarts above baseline.
- Search open issues before filing. Open one issue per new problem with the zone, the Application, the evidence and a suggested fix; comment on existing issues when the problem persists.
- Write the daily audit document to docs/audits/
.md in the internal docs repo, with a summary table and links to each issue, and open a pull request for it. - Finish with exactly one verdict: HEALTHY (no findings), ATTENTION (findings filed as issues), or DEGRADED (a Release environment is Degraded or an SLO budget is exhausted, and it was escalated).
When it runs
- On a schedule,
0 6 * * *. Every day at 06:00 UTC, as part of the daily-audit AgentWorkflow.
What it reads
| Source | What it uses it for |
|---|---|
cluster | Argo CD Applications, Pods, nodes, PersistentVolumes, Certificates and ExternalSecrets in every zone. |
metrics | SLO error ratios, latency and saturation metrics from the metrics stack. |
repo | The gitops repo registry, including every zone's current pin. |
issues | Open issues labeled sre, so findings are not filed twice. |
previous-report | Yesterday's audit document, to report what changed. |
What it produces
- A dated audit document in the internal docs, submitted as a pull request.
- One issue per new problem, labeled sre, with evidence and a suggested fix.
- Comments on existing issues whose problem persists.
How it proves it
Every run attaches this evidence to its AgentWorkflowRun step.
| Evidence | What it shows | Required |
|---|---|---|
| Document | The daily audit document with a summary table, per-zone status and links to issues. | Yes |
| Report | Machine-readable table of every Application with sync status, health and pin. | Yes |
| Metric | The SLO burn queries and values behind any error budget finding. | No |
Success criteria
A run succeeds only when every statement holds.
- An audit document for the date exists in the internal docs repo as a pull request.
- Every Application not Synced and Healthy appears in the audit with a linked issue.
- No problem is filed as a duplicate of an open issue.
- Every issue names the zone, the Application and the evidence, and suggests a fix.
- Certificates expiring within 21 days are all listed.
- The audit states what changed since the previous day's audit.
Guardrails
- Never sync, refresh with prune, roll back or delete an Application. Report only.
- Never edit a pin or open a promotion; the promotion flow owns that.
- Never change SLO targets or alert rules to clear a finding.
- Never read or print Secret values.
- Never merge the audit pull request; a human merges it.
- Do not file issues for Progressing Applications younger than one hour.
Permissions
Deny wins over allow.
| Tools allowed | Read, Grep, Glob, Write, Edit, Bash(kubectl get:*), Bash(kubectl describe:*), Bash(kubectl logs:*), Bash(argocd app list:*), Bash(argocd app get:*), Bash(promtool query:*), Bash(git checkout -b:*), Bash(git add:*), Bash(git commit:*), Bash(git push:*), Bash(gh pr create:*), Bash(gh issue:*) |
| Tools denied | WebSearch, Bash(argocd app sync:*), Bash(argocd app delete:*), Bash(argocd app rollback:*), Bash(kubectl delete:*), Bash(kubectl apply:*), Bash(kubectl get secret:*), Bash(git push --force:*), Bash(gh pr merge:*) |
| Git scopes | contents:read, contents:write, pull_requests:write, issues:write |
| Cluster verbs | get, list, logs |
| Network | allowlist |
| Egress allowlist | api.github.com, github.com, victoria-metrics.infrared.svc, argocd-server.argocd.svc |
| May merge its own pull requests | No |
When it hands off to a human
It dead-letters the work to @platform/sre if it has not finished after 30m, or as soon as any of these is true:
- An Application in a Release environment is Degraded or Missing.
- An SLO has exhausted its monthly error budget.
- A certificate expires within 7 days.
- Metrics or Argo CD were unreachable for the whole run.
Verdicts
Every run ends with exactly one of these verdicts:
HEALTHYATTENTIONDEGRADED
Opinions
Opinions are the org’s editable guidance for this role. Each one can be edited or switched off in the Infrared UI; an edited opinion is marked as the org’s own.
How to judge error budget burn
slo-burn-alerting · origin catalog
Use multi-window burn rates. A 1-hour burn rate above 14.4 or a 6-hour burn rate above 6 is urgent. A 3-day burn rate above 1 is worth an issue. Always state the SLO target and the window.
When a pin is stale
stale-pin-threshold · origin catalog
A Release environment pin more than 14 days behind the latest image that passed Pre-release is stale. Stale pins accumulate risk; the issue should propose the promotion, not perform it.
Audit document format
audit-format · origin catalog
Start with a three-line summary, then a table of zones with sync, health, pin age and open issues. Keep the document under two screens. Link, do not paste, long logs.
One problem, one issue
one-problem-one-issue · origin catalog
File one issue per distinct problem, titled with the zone and the symptom. Update the existing issue while it stays open instead of filing a new one each day.