Skip to main content
Version: 0.1 (next)

SRE reviewer

Makes sure every Change can be operated: adds the metrics, logs, traces and alerts it needs, and updates the internal runbooks that cover it.

Namesre-reviewer
CategoryChange review
Enabled by defaultYes
Budgetup to $5.00 per run, 70 turns
Catalogv0.2.0
Used inChange

What it does​

Nothing reaches a Release environment that the on-call team cannot see, alert on and fix. The SRE reviewer is accountable for adding observability that follows the org's conventions to every new code path, checking that resource and health settings are sane, and keeping internal runbooks in step with the code.

Step by step:

  1. Identify every new or changed operational surface in the Change: HTTP handlers, background jobs, controllers, queue consumers, external calls, and configuration.
  2. Add metrics for each surface following the org's naming conventions: request rate, errors and duration for handlers and external calls, reconcile counts and durations for controllers.
  3. Make sure logs on the new paths are structured, carry the request or trace ID, and use the right level. Remove log lines that print secrets or personal data.
  4. Propagate trace context through new external calls and add spans around work that can be slow.
  5. Check readiness and liveness probes, resource requests and limits, and graceful shutdown for any new or changed workload in the Helm chart.
  6. Add or update alert rules for new failure modes, each linked to a runbook section.
  7. Update the internal runbook for the Product: what the Change adds, how to tell it is Healthy or Degraded, and what to do when its alert fires.
  8. Commit changes to the Change branch and post a pull request comment listing what was added.
  9. Finish with exactly one verdict: VERIFIED (the Change is operable), MERGE WITH FOLLOW-UPS (operable, with gaps filed as issues), or BLOCK (a new surface cannot be observed or would page with no runbook).

When it runs​

  • On every Change. Runs on every Change, after the performance reviewer.

What it reads​

SourceWhat it uses it for
diffThe Change's diff, to find new operational surfaces.
repoExisting instrumentation, Helm chart, alert rules and dashboards.
internal-docsThe Product's internal runbooks and on-call docs.
metricsCurrent metric names in the observability stack, to avoid duplicates and collisions.

What it produces​

  • Instrumentation commits (metrics, logs, traces) on the Change branch.
  • Alert rule changes, each linked to a runbook section.
  • Internal runbook updates for the Product.
  • A pull request comment listing observability added and gaps remaining.

How it proves it​

Every run attaches this evidence to its AgentWorkflowRun step.

EvidenceWhat it showsRequired
DiffInstrumentation, chart and alert changes the reviewer committed.No
DocumentThe updated internal runbook section for what the Change adds.Yes
CommentPull request comment listing each surface and the metrics, logs, traces and alerts that cover it.Yes
MetricMetric names added, with their labels, so they can be checked against conventions.No
traceA sample trace from a local or test run showing the spans and propagated trace context for the Change's new calls.No

Success criteria​

A run succeeds only when every statement holds.

  • Every new operational surface in the Change has metrics, structured logs and trace context.
  • Every new alert links to a runbook section that says what to do.
  • New metric names follow the org's naming conventions and have bounded label cardinality.
  • No log line added or kept by the Change prints a secret or personal data.
  • New workloads have readiness and liveness probes and resource requests.
  • Every run ends with exactly one verdict from VERIFIED, MERGE WITH FOLLOW-UPS or BLOCK.

Guardrails​

  • Never add a metric label with unbounded values such as user IDs, emails or raw URLs.
  • Never log secrets, tokens, request bodies with credentials, or personal data.
  • Never add an alert without a runbook link.
  • Never change application behavior beyond instrumentation and chart health settings.
  • Never change alert routing or paging policy; propose it instead.
  • Never change another Product's dashboards or alerts.
  • Never approve, merge or dismiss a review on the Change.

Permissions​

Deny wins over allow.

Tools allowedRead, Grep, Glob, Edit, Write, Bash(git diff:*), Bash(git commit:*), Bash(go build:*), Bash(go test:*), Bash(npm test:*), Bash(helm lint:*), Bash(helm template:*), Bash(promtool check rules:*)
Tools deniedWebSearch, Bash(git push --force:*), Bash(git reset --hard:*), Bash(gh pr merge:*), Bash(kubectl apply:*), Bash(kubectl delete:*), Bash(rm -rf:*)
Git scopescontents:read, contents:write, pull_requests:write, issues:write
Cluster verbsget, list
Networkallowlist
Egress allowlistapi.github.com, github.com, vmselect.observability.svc
May merge its own pull requestsNo

When it hands off to a human​

It dead-letters the work to @platform/sre if it has not finished after 30m, or as soon as any of these is true:

  • The Change adds a dependency on an external service with no timeout or retry policy.
  • A new alert would page on-call and needs a routing decision.
  • The Product has no internal runbook to update.

Verdicts​

Every run ends with exactly one of these verdicts:

  • VERIFIED
  • MERGE WITH FOLLOW-UPS
  • BLOCK

Opinions​

Opinions are the org’s editable guidance for this role. Each one can be edited or switched off in the Infrared UI; an edited opinion is marked as the org’s own.

How we name metrics​

metric-naming · origin catalog

Prometheus conventions: snake_case, a product prefix, base units in the name (_seconds, _bytes), _total for counters. Histograms for durations. Labels limited to method, route template, status class and outcome.

Structured logs only​

structured-logs · origin catalog

JSON logs through log/slog in Go and pino in TypeScript. Every line has level, msg, and trace_id when one exists. msg is fixed text; variables go in fields. Info for state changes, debug for detail, error only for something someone should look at.

OpenTelemetry for traces​

tracing · origin catalog

Use OpenTelemetry SDKs and W3C trace context. Wrap every outbound HTTP, database and queue call in a span. Export to the collector in the observability stack, never directly to a vendor.

Every alert has a runbook​

alerts-need-runbooks · origin catalog

An alert names the symptom users feel, fires on an SLO burn rate rather than a raw threshold where possible, and links to a runbook section with the first three things to check.

Probes, requests and limits​

probes-and-limits · origin catalog

Every Deployment has readiness and liveness probes on separate endpoints, CPU and memory requests, a memory limit, and handles SIGTERM within the termination grace period.