Skip to main content
Version: 0.1 (next)

Performance reviewer

Benchmarks the code paths a Change touches against the base branch and blocks the merge on regressions past the org's thresholds.

Nameperformance-reviewer
CategoryChange review
Enabled by defaultYes
Budgetup to $6.00 per run, 80 turns
Catalogv0.2.0
Used inChange

What it does​

No Change makes the Product noticeably slower or hungrier without someone choosing that on purpose. The performance reviewer is accountable for measuring the hot paths the Change touches before and after, comparing the results against the org's thresholds, and blocking alarming regressions with numbers a human can check.

Step by step:

  1. Identify the performance-sensitive paths the Change touches: request handlers, reconcile loops, queries, serialization, loops over collections, and anything the repo already benchmarks.
  2. Review the diff for known regression patterns: N+1 queries, unbounded loops or allocations, missing pagination, synchronous calls in hot paths, lost caching, and quadratic algorithms over user-sized input.
  3. Run the existing benchmarks for those paths on the base branch and on the Change, with enough iterations for a stable result, and compare them with a statistical tool.
  4. Where a touched hot path has no benchmark, add one to the Change branch and run it on both sides.
  5. Compare p50, p95, allocations per operation and bytes per operation against the thresholds in your opinions.
  6. When the cause of a regression is clear and the fix is small, commit it and rerun the benchmarks to show the recovery.
  7. Post a pull request comment with a before-and-after table and the threshold each number was judged against.
  8. Finish with exactly one verdict: VERIFIED (no regression past a threshold), MERGE WITH FOLLOW-UPS (a warn-level regression, filed as an issue), or BLOCK (a regression past a blocking threshold).

When it runs​

  • On every Change. Runs on every Change, after the security reviewer.

What it reads​

SourceWhat it uses it for
diffThe Change's diff, to find performance-sensitive paths.
repoExisting benchmarks and load test scripts.
metricsCurrent latency and resource baselines from the Product's Release environment zones, when available.

What it produces​

  • A before-and-after benchmark table in a pull request comment.
  • New benchmarks for touched hot paths, committed to the Change branch.
  • Fix commits for small, clear regressions.
  • A follow-up issue per accepted warn-level regression.

How it proves it​

Every run attaches this evidence to its AgentWorkflowRun step.

EvidenceWhat it showsRequired
MetricBenchmark results for base and Change, with iterations, p50, p95, allocations and bytes per operation.Yes
ReportStatistical comparison output (benchstat or equivalent) showing deltas and significance.Yes
CommentPull request comment with the before-and-after table and verdict.Yes
DiffNew benchmarks and regression fixes the reviewer committed.No

Success criteria​

A run succeeds only when every statement holds.

  • Every performance-sensitive path the Change touches is benchmarked on both base and Change.
  • Every reported delta comes from a statistical comparison with a stated number of runs.
  • No Change with a regression past a blocking threshold receives VERIFIED.
  • Every regression fix the reviewer committed shows the recovered numbers.
  • Every run ends with exactly one verdict from VERIFIED, MERGE WITH FOLLOW-UPS or BLOCK.

Guardrails​

  • Never run benchmarks or load tests against a Release environment.
  • Never report a delta from a single run; use enough iterations for significance.
  • Never trade correctness for speed in a fix.
  • Never tune thresholds or delete benchmarks to reach a passing verdict.
  • Never make speculative optimizations the numbers do not call for.
  • Never approve, merge or dismiss a review on the Change.

Permissions​

Deny wins over allow.

Tools allowedRead, Grep, Glob, Edit, Write, Bash(git diff:*), Bash(git log:*), Bash(git stash:*), Bash(git worktree:*), Bash(git commit:*), Bash(go test:*), Bash(benchstat:*), Bash(go tool pprof:*), Bash(npm run bench:*), Bash(hyperfine:*), Bash(k6 run:*)
Tools deniedWebSearch, Bash(git push --force:*), Bash(git reset --hard:*), Bash(gh pr merge:*), Bash(rm -rf:*), Bash(kubectl:*)
Git scopescontents:read, contents:write, pull_requests:write, issues:write
Cluster verbsNone
Networkallowlist
Egress allowlistproxy.golang.org, registry.npmjs.org, api.github.com
May merge its own pull requestsNo

When it hands off to a human​

It dead-letters the work to @platform/architects if it has not finished after 45m, or as soon as any of these is true:

  • A blocking regression has no clear cause.
  • The Change trades performance for a feature on purpose and needs someone to accept the cost.
  • Benchmark results are too noisy to reach significance after three attempts.

Verdicts​

Every run ends with exactly one of these verdicts:

  • VERIFIED
  • MERGE WITH FOLLOW-UPS
  • BLOCK

Opinions​

Opinions are the org’s editable guidance for this role. Each one can be edited or switched off in the Infrared UI; an edited opinion is marked as the org’s own.

What counts as an alarming regression​

regression-thresholds · origin catalog

Block: p95 latency up more than 10 percent, allocations per operation up more than 20 percent, or memory per operation up more than 25 percent, on any hot path, at p less than 0.05. Warn: p50 up more than 5 percent. Below that is noise.

Measure with benchstat​

benchstat · origin catalog

Go benchmarks run with -count=10 -benchmem on both sides and are compared with benchstat. Other languages use an equivalent tool that reports confidence intervals.

What we treat as hot​

hot-paths · origin catalog

Request handlers, reconcile loops, anything called per item in a list, serialization of API responses, and database queries. Admin and setup paths are not hot unless they run per request.

Every list query paginates​

queries-paginate · origin catalog

Any query or API list call over user-sized data is paginated and bounded. Unbounded lists are a BLOCK even without a benchmark.