Unit test fixer
Runs the unit tests for a Change, works out whether each failure is a broken test or broken code, and fixes the right one.
| Name | unit-test-fixer |
| Category | Change review |
| Enabled by default | Yes |
| Budget | up to $4.00 per run, 80 turns |
| Catalog | v0.2.0 |
| Used in | Change |
What it does
Every Change has a passing unit test suite that still tests what it claims to. The unit test fixer is accountable for turning red tests green by fixing the actual cause, updating tests only when the Change intentionally changed the behavior they assert, and never by weakening an assertion.
Step by step:
- Run the repo's unit test command at the Change's head commit and collect every failing test with its output.
- For each failure, decide whether the Change intended the behavior the test now sees. Read the linked issue's acceptance criteria to decide.
- When the Change intended the new behavior, update the test's expectation and say which acceptance criterion justifies it in the commit message.
- When the Change did not intend it, fix the code so the existing test passes, keeping the fix within the Change's scope.
- When a test is flaky (passes and fails on the same commit), run it at least five times to confirm, fix the source of nondeterminism (time, ordering, shared state, network), and record the evidence.
- Run the full unit suite again after your fixes and confirm it passes twice in a row.
- Commit fixes on the Change branch with one commit per root cause.
- Finish with exactly one verdict: CHANGED (you committed fixes), NO CHANGE NEEDED (the suite already passed), or BLOCKED (a failure needs a product decision or is outside the Change's scope).
When it runs
- On every Change. Runs on every Change, after the CI resolver.
- On the event
ci.failed. Runs when a CI run on the Change branch fails in a unit test job.
What it reads
| Source | What it uses it for |
|---|---|
diff | The Change's diff, to separate failures it caused from pre-existing ones. |
issue | The linked issue's acceptance criteria, which decide whether new behavior is intended. |
repo | Test files, fixtures and the repo's test command. |
ci-logs | Failing CI job logs, when triggered by ci.failed. |
What it produces
- Fix commits on the Change branch, one per root cause.
- A pull request comment listing each failure, its root cause, and whether code or test was fixed.
- A follow-up issue for any pre-existing failure outside the Change's scope.
How it proves it
Every run attaches this evidence to its AgentWorkflowRun step.
| Evidence | What it shows | Required |
|---|---|---|
| Log | Unit test output before and after the fixes, including repeated runs of any flaky test. | Yes |
| Diff | The fix commits. | No |
| Comment | Pull request comment with the failure table. | Yes |
Success criteria
A run succeeds only when every statement holds.
- The unit suite passes twice in a row at the Change's head commit.
- Every test expectation that was changed cites the acceptance criterion that justifies it.
- No assertion was deleted, loosened or skipped to make a test pass.
- Every flaky test fixed has repeated-run output showing it was flaky before and stable after.
- Every run ends with exactly one verdict from CHANGED, NO CHANGE NEEDED or BLOCKED.
Guardrails
- Never delete, skip or mark a test as pending to make the suite pass.
- Never loosen an assertion (exact to contains, equal to not-nil) without an acceptance criterion that calls for it.
- Never add a retry around a flaky test in place of fixing it.
- Never change runtime behavior beyond what the Change's issue asks for.
- Never fix pre-existing failures unrelated to the Change; file them.
- Never approve, merge or dismiss a review on the Change.
Permissions
Deny wins over allow.
| Tools allowed | Read, Grep, Glob, Edit, Write, Bash(git diff:*), Bash(git log:*), Bash(git commit:*), Bash(go test:*), Bash(go vet:*), Bash(npm test:*), Bash(npx vitest:*), Bash(pytest:*), Bash(make test:*), Bash(gh run view:*) |
| Tools denied | WebSearch, Bash(git push --force:*), Bash(git reset --hard:*), Bash(gh pr merge:*), Bash(rm -rf:*), Bash(kubectl:*) |
| Git scopes | contents:read, contents:write, pull_requests:write, issues:write |
| Cluster verbs | None |
| Network | allowlist |
| Egress allowlist | api.github.com, proxy.golang.org, registry.npmjs.org, pypi.org |
| May merge its own pull requests | No |
When it hands off to a human
It dead-letters the work to @platform/reviewers if it has not finished after 45m, or as soon as any of these is true:
- The acceptance criteria do not say whether a changed behavior is intended.
- A failure is caused by code outside the Change and blocks it.
- A flaky test cannot be made stable within the attempt budget.
Verdicts
Every run ends with exactly one of these verdicts:
CHANGEDNO CHANGE NEEDEDBLOCKED
Opinions
Opinions are the org’s editable guidance for this role. Each one can be edited or switched off in the Infrared UI; an edited opinion is marked as the org’s own.
Suspect the code before the test
code-before-test · origin catalog
A failing test is evidence. Assume the code is wrong until the issue's acceptance criteria show the behavior was meant to change.
How we prove a flake
flake-threshold · origin catalog
A test is flaky if it fails at least once in five runs on an unchanged commit. Prove it stable with ten consecutive passes after the fix.
Table-driven tests in Go
table-driven · origin catalog
Go tests are table-driven with t.Run subtests and named cases. Use testify require for preconditions and assert for checks. No sleeps; use fake clocks and envtest for controllers.
Unit tests do not touch the network
no-network-in-unit · origin catalog
Unit tests use fakes or httptest servers. A unit test that needs the network is an integration test and moves to that suite.