Review agents stop at findings. Pit-stop finishes the job.
Other agent tools stop short: PR reviewers only read pull requests; single-file review skills diagnose without acting; workflow skills teach phases but never run them end to end; architecture skills hand over a report and leave. Pit-stop stays in the car through the whole lap.
Pit-stop is a cross-runtime agent skill: one instruction runs a full improvement loop — load → find → propose → fix → ideas → report — with no mid-run questions.
Use pit-stop on <project path>
You say one line. The agent loads context, finds weaknesses (every one with
path:line evidence), fixes them in a review→fix loop, verifies with real tool
output, and returns once with a report split into verified / unverified /
remaining. Push and publish never happen without your explicit word.
Built to fix three failure modes every agent owner has met:
- #1: The agent reports, never fixes. Reviews end with "you should…", and the diff never happens. Fix: the loop doesn't end at findings — fix, re-review (max 3 rounds, cross-round ledger, escalation on stagnation), then report.
- #2: "Done" with no proof. "Fixed!", "tests pass!" — with no command output behind the words. Fix: evidence before claims. No verification run in the turn = no success claim. Reports carry tool outputs, not adjectives.
- #3: Improvement means bloat. Suggestions pile on abstractions nobody asked for.
Fix: every finding is tagged (
delete/stdlib/native/yagni/shrink/perf/security/obs), Speculative items are reported, never built. Nothing found:Lean already. Ship.
A real run: findings with
path:lineevidence 8/8 (hand-rolled first draft: 5/8), two verbal-only "success" claims intercepted into tool receipts. Read the full before/after → — honest baseline: n=1 repo, independent re-runs welcome.
Six phases, one pass, zero mid-run questions — guardrails on top, escalation exit below. ▶ Interactive version
- Load — read the project's own instructions,
git status, recent-commit hot spots; write one MODE line: what counts as a finding here. - Find — scope before scanning; every finding carries
path:lineevidence, a tag, and a strength. - Propose — Strong items only: symptom, evidence, impact, minimal fix, cost — plus a NOT-doing list.
- Fix — review→fix loop with an independent reviewer, max 3 rounds; stagnation escalates to you.
- Ideas — read-only brainstorm once the loop closes: up to 3 anchored
[idea]candidates, each with a kill-probe, never built this run. - Report — changed / verified (tool output) / unverified / remaining (the ideas ride behind needs-human items); no claim without a fresh verification run.
The L2 hook (hooks/block-destructive.sh) runs in Claude-family harnesses and fails closed — a command it can't parse is blocked. Feed it a command it recognizes as destructive and it exits 2 with the reason:
$ printf '{"tool_input":{"command":"rm -rf ~/projects"}}' | bash hooks/block-destructive.sh
BLOCKED by pit-stop: 'rm -rf ~/projects' matches file deletion (rm). Destructive ops need the human's explicit word.
$ echo $?
2
Safe commands pass through untouched (grep, man rm, git commit -m "... rm ...").
It covers rm in command position (incl. sudo/do/command/env/nohup/time/xargs/\rm
variants, flag values like sudo -u root rm, a path-qualified /bin/rm, and a destructive op on a
later line of a multi-line command), find -exec rm/-delete, non-git VCS force push, and
destructive git subcommands (push, reset --hard, clean -f, branch -D, checkout ., restore ., git rm).
Pattern matching is a tripwire, not a sandbox — sh -c 'rm …'-style vectors stay
with the L1 ban and the L3 sweep. The 90-case matrix lives in
hooks/test-block-destructive.sh.
Install (Claude Code — one file)
npx skills add Finn763/pit-stopTwo ways in, one skill: the plugin/registry paths subscribe (updates arrive on
their own); copying skills/ forks (you own the files, you edit them).
Pick your agent when asked; update later with npx skills update. Per harness:
| Harness | Install |
|---|---|
| Claude Code | /plugin marketplace add Finn763/pit-stop, then /plugin install pit-stop@pit-stop |
| Codex | plugin from .codex-plugin (see repo), or npx skills above |
| Cursor | rule from .cursor/rules/ (auto-loaded); .cursor-plugin/ for Cursor Plugins |
| Gemini CLI | gemini extensions install https://github.com/Finn763/pit-stop |
| Pi | pi install npm:@finn763/pit-stop (or copy skills/) |
| OpenCode | command from .opencode/command/, or copy skills/ |
| Hermes | plugin from .hermes-plugin, or copy skills/ |
| Devin | plugin manifest in .devin-plugin/ (see repo); universal fallback below |
| Kimi | plugin manifest in .kimi-plugin/ (see repo); universal fallback below |
| Windsurf | rule from .windsurf/rules/ |
| Anything else | cp -r skills/pit-stop ~/.agents/skills/ |
No per-repo setup for the skill itself. The hook is the one optional extra (repo-local, per instructions above); nothing else to configure.
| Area | What's pinned down |
|---|---|
| Run | Six phases, one pass: load → find → propose → fix → ideas → report (verification is a hard gate, not a phase). Zero mid-run questions |
| Findings | Every one with path:line evidence, a tag, and a strength — Speculative items are reported, never built |
| Fix loop | Review→fix with an independent reviewer, max 3 rounds, cross-round ledger, stagnation escalates to human |
| Ideas | Phase 5, read-only: ≤3 anchored [idea] candidates from this run's ledger/diff, each with a kill-probe — never built in the run that raises them |
| Verification | No verification run in the turn = no success claim. Reports carry tool output, not adjectives |
| Cost | Phases report spend; one phase past $20 stops itself |
| Push / publish | Never automatic. Everything waits in the workdir for one explicit word |
Repo layout
skills/pit-stop/SKILL.md # the skill (<500-word core)
skills/pit-stop/references/ # per-phase rules (audit/fix/review/report/ideas) + guardrails, verification
skills/pit-stop/templates/ # report template
hooks/block-destructive.sh # L2 guardrail (Claude-family hooks)
hooks/test-block-destructive.sh # 90-case guardrail matrix (CI on 3 OS)
hooks/check-consistency.sh # six-phase / [idea] / manifest-version / adapter-registration / sudo gate (CI)
commands/ .opencode/ # slash-command entries
.claude-plugin/ .codex-plugin/ .cursor-plugin/ .devin-plugin/
.kimi-plugin/ .hermes-plugin/ .pi/ .cursor/ .windsurf/ # per-harness adapters
.claude/settings.json .github/workflows/ci.yml # dogfood hook + CI matrix
gemini-extension.json GEMINI.md package.json
examples/before-after.md # real run, real diff
docs/SPEC.md # full specification (v4)
Repo work rules live in AGENTS.md; fixes welcome.
Lean already. Ship.