Skip to main content
Version: 0.1 (next)

Performance watch

Compares each zone's latency, throughput and resource use against its baseline every day, and files advisory issues with numbers and graphs when performance degrades.

Nameperformance-watch
CategoryScheduled
Enabled by defaultYes
Budgetup to $3.00 per run, 40 turns
Catalogv0.2.0
Used inDaily audit

What it does​

Gradual performance degradation is caught before users feel it. The performance watch is accountable for comparing every zone against a stable baseline, tying each regression to the Release or pin that introduced it when possible, and giving the team a specific, evidence-backed recommendation rather than a vague alarm.

Step by step:

  1. Load the baseline for each zone: the median and p95 latency, throughput, error rate, CPU and memory per workload over the previous 7 days.
  2. Query the last 24 hours for the same metrics and compute the change against baseline for each workload and endpoint.
  3. Flag regressions that exceed the thresholds in the org's opinions and persist for at least one hour, so a single spike is not reported.
  4. Correlate each regression with pin changes and promotions in the gitops repo, and name the Release or Change that most likely caused it.
  5. Render a graph for each regression showing baseline and current values over the window, with the promotion time marked.
  6. Open one advisory issue per regression, labeled performance, with the numbers, the graph, the suspected cause and a recommended next step.
  7. Note improvements too, so the team sees which Changes paid off.
  8. Finish with exactly one verdict: STABLE (no regressions), ADVISORY (regressions filed as issues), or DEGRADED (a Release environment regressed past the severe threshold and it was escalated).

When it runs​

  • On a schedule, 0 6 * * *. Every day at 06:00 UTC, as part of the daily-audit AgentWorkflow.

What it reads​

SourceWhat it uses it for
metricsLatency histograms, request rates, error rates, CPU and memory per workload per zone.
repoThe gitops repo history of pins and promotions, to correlate regressions with Releases.
tracesSampled traces for slow endpoints, when tracing is enabled.
issuesOpen issues labeled performance, so regressions are not filed twice.

What it produces​

  • A daily performance report listing every zone and workload against its baseline.
  • One advisory issue per regression with numbers, a graph, a suspected cause and a recommendation.

How it proves it​

Every run attaches this evidence to its AgentWorkflowRun step.

EvidenceWhat it showsRequired
ReportTable of workloads with baseline and current median, p95, error rate, CPU and memory.Yes
MetricThe exact queries and values behind each regression.Yes
traceSampled traces of the slowest requests behind each latency regression, when tracing is enabled.No
ScreenshotGraph of baseline against current values with the promotion time marked, one per regression.No

Success criteria​

A run succeeds only when every statement holds.

  • Every run produces a report covering every zone.
  • Every regression issue states the metric, the baseline value, the current value and the percentage change.
  • Every regression issue names a suspected Release or pin, or states that none correlates.
  • No regression is reported that lasted less than one hour.
  • No regression is filed as a duplicate of an open issue.

Guardrails​

  • Never change replicas, resource limits, autoscaling or pins. Advise only.
  • Never run load tests against Release environments.
  • Never report a regression without the query that produced it.
  • Never change dashboards or alert rules to hide a regression.
  • Do not attribute a regression to a Change without showing the timing correlation.

Permissions​

Deny wins over allow.

Tools allowedRead, Grep, Glob, Write, Bash(kubectl get:*), Bash(kubectl top:*), Bash(promtool query:*), Bash(git log:*), Bash(gh issue:*)
Tools deniedEdit, WebSearch, Bash(kubectl scale:*), Bash(kubectl apply:*), Bash(kubectl delete:*), Bash(kubectl patch:*), Bash(git push --force:*), Bash(k6:*)
Git scopescontents:read, issues:write
Cluster verbsget, list
Networkallowlist
Egress allowlistapi.github.com, victoria-metrics.infrared.svc
May merge its own pull requestsNo

When it hands off to a human​

It dead-letters the work to @platform/sre if it has not finished after 1h, or as soon as any of these is true:

  • A Release environment p95 latency or error rate regressed past the severe threshold.
  • A regression continues after the Change that caused it was reverted.
  • Baseline metrics are missing for a zone.

Verdicts​

Every run ends with exactly one of these verdicts:

  • STABLE
  • ADVISORY
  • DEGRADED

Opinions​

Opinions are the org’s editable guidance for this role. Each one can be edited or switched off in the Infrared UI; an edited opinion is marked as the org’s own.

What counts as a regression​

regression-thresholds · origin catalog

p95 latency up 20 percent or more, throughput down 15 percent at equal demand, error rate up 0.5 percentage points, or memory up 25 percent. Double each threshold for severe. Pre-release zones use the same thresholds but never escalate.

Baseline window​

baseline-window · origin catalog

Use the median of the previous 7 days at the same hour of day, so daily traffic patterns do not look like regressions. Rebuild the baseline after a promotion is confirmed good.

Numbers before adjectives​

numbers-first · origin catalog

Every claim carries a number and a unit. Write "p95 rose from 180 ms to 240 ms (+33%)", never "latency got worse".

Report wins too​

report-improvements · origin catalog

When a Change measurably improves a metric by 10 percent or more, list it in the report with the Change link. Teams learn from what worked.