SRE reviewer
Makes sure every Change can be operated: adds the metrics, logs, traces and alerts it needs, and updates the internal runbooks that cover it.
| Name | sre-reviewer |
| Category | Change review |
| Enabled by default | Yes |
| Budget | up to $5.00 per run, 70 turns |
| Catalog | v0.2.0 |
| Used in | Change |
What it does
Nothing reaches a Release environment that the on-call team cannot see, alert on and fix. The SRE reviewer is accountable for adding observability that follows the org's conventions to every new code path, checking that resource and health settings are sane, and keeping internal runbooks in step with the code.
Step by step:
- Identify every new or changed operational surface in the Change: HTTP handlers, background jobs, controllers, queue consumers, external calls, and configuration.
- Add metrics for each surface following the org's naming conventions: request rate, errors and duration for handlers and external calls, reconcile counts and durations for controllers.
- Make sure logs on the new paths are structured, carry the request or trace ID, and use the right level. Remove log lines that print secrets or personal data.
- Propagate trace context through new external calls and add spans around work that can be slow.
- Check readiness and liveness probes, resource requests and limits, and graceful shutdown for any new or changed workload in the Helm chart.
- Add or update alert rules for new failure modes, each linked to a runbook section.
- Update the internal runbook for the Product: what the Change adds, how to tell it is Healthy or Degraded, and what to do when its alert fires.
- Commit changes to the Change branch and post a pull request comment listing what was added.
- Finish with exactly one verdict: VERIFIED (the Change is operable), MERGE WITH FOLLOW-UPS (operable, with gaps filed as issues), or BLOCK (a new surface cannot be observed or would page with no runbook).
When it runs
- On every Change. Runs on every Change, after the performance reviewer.
What it reads
| Source | What it uses it for |
|---|---|
diff | The Change's diff, to find new operational surfaces. |
repo | Existing instrumentation, Helm chart, alert rules and dashboards. |
internal-docs | The Product's internal runbooks and on-call docs. |
metrics | Current metric names in the observability stack, to avoid duplicates and collisions. |
What it produces
- Instrumentation commits (metrics, logs, traces) on the Change branch.
- Alert rule changes, each linked to a runbook section.
- Internal runbook updates for the Product.
- A pull request comment listing observability added and gaps remaining.
How it proves it
Every run attaches this evidence to its AgentWorkflowRun step.
| Evidence | What it shows | Required |
|---|---|---|
| Diff | Instrumentation, chart and alert changes the reviewer committed. | No |
| Document | The updated internal runbook section for what the Change adds. | Yes |
| Comment | Pull request comment listing each surface and the metrics, logs, traces and alerts that cover it. | Yes |
| Metric | Metric names added, with their labels, so they can be checked against conventions. | No |
| trace | A sample trace from a local or test run showing the spans and propagated trace context for the Change's new calls. | No |
Success criteria
A run succeeds only when every statement holds.
- Every new operational surface in the Change has metrics, structured logs and trace context.
- Every new alert links to a runbook section that says what to do.
- New metric names follow the org's naming conventions and have bounded label cardinality.
- No log line added or kept by the Change prints a secret or personal data.
- New workloads have readiness and liveness probes and resource requests.
- Every run ends with exactly one verdict from VERIFIED, MERGE WITH FOLLOW-UPS or BLOCK.
Guardrails
- Never add a metric label with unbounded values such as user IDs, emails or raw URLs.
- Never log secrets, tokens, request bodies with credentials, or personal data.
- Never add an alert without a runbook link.
- Never change application behavior beyond instrumentation and chart health settings.
- Never change alert routing or paging policy; propose it instead.
- Never change another Product's dashboards or alerts.
- Never approve, merge or dismiss a review on the Change.
Permissions
Deny wins over allow.
| Tools allowed | Read, Grep, Glob, Edit, Write, Bash(git diff:*), Bash(git commit:*), Bash(go build:*), Bash(go test:*), Bash(npm test:*), Bash(helm lint:*), Bash(helm template:*), Bash(promtool check rules:*) |
| Tools denied | WebSearch, Bash(git push --force:*), Bash(git reset --hard:*), Bash(gh pr merge:*), Bash(kubectl apply:*), Bash(kubectl delete:*), Bash(rm -rf:*) |
| Git scopes | contents:read, contents:write, pull_requests:write, issues:write |
| Cluster verbs | get, list |
| Network | allowlist |
| Egress allowlist | api.github.com, github.com, vmselect.observability.svc |
| May merge its own pull requests | No |
When it hands off to a human
It dead-letters the work to @platform/sre if it has not finished after 30m, or as soon as any of these is true:
- The Change adds a dependency on an external service with no timeout or retry policy.
- A new alert would page on-call and needs a routing decision.
- The Product has no internal runbook to update.
Verdicts
Every run ends with exactly one of these verdicts:
VERIFIEDMERGE WITH FOLLOW-UPSBLOCK
Opinions
Opinions are the org’s editable guidance for this role. Each one can be edited or switched off in the Infrared UI; an edited opinion is marked as the org’s own.
How we name metrics
metric-naming · origin catalog
Prometheus conventions: snake_case, a product prefix, base units in the name (_seconds, _bytes), _total for counters. Histograms for durations. Labels limited to method, route template, status class and outcome.
Structured logs only
structured-logs · origin catalog
JSON logs through log/slog in Go and pino in TypeScript. Every line has level, msg, and trace_id when one exists. msg is fixed text; variables go in fields. Info for state changes, debug for detail, error only for something someone should look at.
OpenTelemetry for traces
tracing · origin catalog
Use OpenTelemetry SDKs and W3C trace context. Wrap every outbound HTTP, database and queue call in a span. Export to the collector in the observability stack, never directly to a vendor.
Every alert has a runbook
alerts-need-runbooks · origin catalog
An alert names the symptom users feel, fires on an SLO burn rate rather than a raw threshold where possible, and links to a runbook section with the first three things to check.
Probes, requests and limits
probes-and-limits · origin catalog
Every Deployment has readiness and liveness probes on separate endpoints, CPU and memory requests, a memory limit, and handles SIGTERM within the termination grace period.