Agent Skill / language agnostic / MIT
pr-review
A line-by-line diff answers a question nobody asked. This reviews a pull request by running it — and proves what changed instead of asserting it.
tests before
tests before
tests after
Schematic. The scenarios are the author's own; only the build they run against changes.
The load-bearing run
mix
base tells you what is already red. head tells you whether the PR is green. Neither answers the question people actually ask.
mix does: the PR's production code with the test code and test data from before it. A scenario that passed before and fails now changed behavior — stated by the author's own suite, not by a reviewer's reading.
Then one question decides how it reads. Did the PR edit that test or its fixtures? Edited means the author knew. Untouched means nobody wrote it down.
What execution cannot see
probes
A test run only speaks about scenarios somebody already wrote down. Changed code that no scenario executed is not safe, it is unknown.
A probe is a throwaway test case written for measurement and run on both sides, aimed at whatever the diff narrowed: a value added to a blacklist, a condition replaced by a lookup, a newly required input.
One measurement is not an observation — the probe is the pair. This is where the regressions nobody tested for turn up.
A real report
charmbracelet/crush #3456 produced by this skill open → 2541scenarios, run on each of three builds 0changed result 3 / 4probes diverged 1behavior change the PR does not mentionThe file this PR changes had no test file before the branch, so the full suite moved nothing. A review that stopped at the test run would have concluded “no behavior change” — the opposite of the truth. Every finding came from probes.
Install
A skill is a directory with a SKILL.md. Put it where your agent looks —
for Claude Code that is ~/.claude/skills/:
Cursor, Codex, OpenCode, Gemini CLI, Goose, Copilot, Amp and others read the same format from their own directory. Then ask for a PR review, a diff summary, or “what does this branch actually change”.
It needs git and the ability to run the project's test suite. The bundled
scripts are bash and dependency-free Python 3.8+.
What is in the box
| SKILL.md | the method: five steps, and the rules that keep the report honest |
| references/execution.md | worktrees, priming a build that will not compile, runner formats |
| references/probes.md | where to place a probe, how to keep the pair honest, pitfalls |
| references/report.md | report structure and the numbers the verdict must carry |
| scripts/setup_worktrees.sh | builds base, head and mix, including the revert that is easy to get wrong |
| scripts/collect_results.py | normalizes go test -json, JUnit XML and Jest JSON into one shape |
| scripts/compare_runs.py | three runs to a behavior-change list, flakes filtered out |
| assets/report-template.html | working skeleton of the report: levels, filtering, themes |
What it is not
- Not a code-quality review. No opinions about style, naming or structure.
- Not a list of suggested refactors, and not a per-file metrics dashboard.
- Not a summary of the PR description — a gap between what the author claimed and what the run showed is the finding.