feat(pr-review): the reviewer fixes its own review when the format check flags it - #542
Conversation
There was a problem hiding this comment.
Mogplex PR Review
Status: Attention needed
Request changes: acceptsRevision in lib/workflows/pr-review-self-revision.ts never compares hasIssues, so a flagged review's revision can flip the verdict that AGENTS.md and docs/decisions.md explicitly bind it to keep — flipping the published check conclusion and removing the auto-merge block. Also flagged: reviewFormatPassed is stamped when the format check fails open rather than runs, and one docs/decisions.md row needs rewording.
2 findings were added inline.
Suggestions
-
Reword the garbled
review_formatlocation cell in docs/decisions.md (docs/decisions.md)
The updated "Where it runs" cell forreview_formatindocs/decisions.mdreads: "Every structured PR review: a native review's draft inside the run, and any review not already passed before the check run, native review, and timeline comment are published (pr_review)". The edit splices the new in-run location into the old publish-time wording and no longer parses as a sentence.Since
docs/decisions.mdis the binding reference for this layer, reword it while keeping the row's style, e.g.: "A native review's draft inside the run, and at publish time any structured PR review not already passed — before the check run, native review, and timeline comment are published (pr_review)".
| revised: PrReviewHarnessResult | ||
| ): boolean { | ||
| return ( | ||
| revised.source === "structured" && |
There was a problem hiding this comment.
Critical: Revision can flip the review verdict: acceptsRevision never checks hasIssues
acceptsRevision accepts a revision when its extraction is structured and its findings count is at least the draft's — it never compares hasIssues.
Escape path: the draft files reportReview with hasIssues: true and findings, the format check flags it, and the revision files hasIssues: false while keeping at least as many findings. The merged steps still contain the draft's call, so claimedIssues is true, but the last accepted report has findings, so clearsReviewWithoutFindings is false and the state stays filed/structured (see pr-review-report-state.ts); only the flip that also drops every finding is caught via dropped_findings.
Once accepted, the flipped verdict flows to publication: finalizePrReviewSuccess derives the check conclusion and reason from reviewOutcome.hasIssues, so the review publishes as success/noFindings where the draft said issues exist, and getPrReviewAutoMergeBlockReason, which keys on hasIssues, stops blocking flow auto-merge.
This contradicts the contract this PR itself ships: AGENTS.md's binding rule ("the revision must keep the verdict and every finding"), docs/decisions.md ("accepted only when it keeps the verdict and every finding", stated twice), and this file's own comment at lines 62-64 ("a revision that loses the verdict or a finding, leaves the draft as it was"). The dropped_findings guard exists because this exact model failure was already observed in production (pr-review-report-state.ts: "reviewers told that hasIssues=true needs findings sometimes flip hasIssues to false instead of listing them"), so the revision turn is a second, currently unguarded opportunity for it.
Fix: add revised.reviewOutcome.hasIssues === draft.reviewOutcome.hasIssues to acceptsRevision. Consider also enforcing "every finding" by identity rather than count — every draft (severity, title, path) triple present in the revised findings — since a count-preserving swap (drop one finding, add another) also violates the documented rule, and the publish-time rewrite contract already treats title, severity, path, and line as immutable.
Add a regression test for the flip-with-findings-kept case; the suite currently covers only the flip-with-zero-findings case.
| try { | ||
| const problems = await input.judge(draft); | ||
| if (problems.length === 0) { | ||
| return { ...input.result, reviewFormatPassed: true }; |
There was a problem hiding this comment.
Warning: reviewFormatPassed is stamped when the check never ran (fail-open is indistinguishable from a pass)
reviseFlaggedReview stamps reviewFormatPassed whenever the judge returns no problems, but an empty result also means "never checked".
findReviewFormatProblems returns [] when no gateway credential exists, and decide() fails open to act: false on a timeout, an outage, an open circuit breaker, a slow account-setting read, DECISIONS_DISABLED, or shadow mode — none of which throw, per the decision layer's "decide() never throws" rule. All of those paths stamp the marker, and finalizePrReviewSuccess then skips polishPrReviewForPublish for good.
The comment at lines 62-64 claims a failed check "leaves the draft as it was, and publishing still runs the platform rewrite on it" — that holds for a throwing judge (the catch at line 89), but the most common real check failures return [] and get recorded as a pass. The unit test "should not follow up when the check cannot run" simulates unavailability with a judge that throws, so the tested behavior and production fail-open behavior diverge.
Impact is bounded — the check only polishes formatting, and in the off/unavailable states the publish-time check would be off too — but a transient evaluator outage during the run permanently skips a publish-time check that could have run seconds later, and the persisted marker claims the report "passed the format check inside the run" when it was never judged.
Fix: distinguish "not evaluated" from "evaluated and clean" — for example, have judgeReviewFormat return null when the check did not run, and only stamp the marker on a genuine pass; on null, leave the draft unmarked so finalizePrReviewSuccess still judges it at publish time.
58540a7 to
8317cd6
Compare
|
Rebased on the updated #541 and addressed the review:
|
Mogplex PR ReviewStatus: No material issues found Approve-ready: the self-revision loop is fail-open at every failure point (an unjudged draft, a thrown revision, and an untrusted revision each leave the draft), the acceptance gate requires a trusted report that keeps the verdict and every finding, and publishing keeps its rewrite path for reviews the in-run check cannot judge, skipping only text the evaluator judged in its published shape. Test coverage matches every claimed behavior, including sad paths. Two non-blocking suggestions are filed: the finding key discards a whole revision when a finding gains a line number, and nothing pins the default judge's Affected files
Suggestions
|
…judged report is never a pass
…ewer what a retraction publishes
8317cd6 to
355e27b
Compare
|
Rebased onto main now that #541 is merged, and retargeted to main. The last review had no warnings; its three suggestions are addressed in 'chore(pr-review): record the job run on in-run checks and tell a reviewer what a retraction publishes':
It also carries the retraction wording promised on #541: the dropped-findings request now tells the reviewer that reporting no issues after all publishes the review as incomplete. |
|
Answers to the last review (no warnings, so merging as is):
|
Builds on #541, which is merged.
When the
review_formatcheck flags a review today, a platform model rewrites the text at publish time. A rewrite can only reword what is already there. The #535 review said "three minor suggestions, detailed in the comment" and listed none. The check'sdanglingReferencequestion catches exactly that, but the only fix is to state the suggestions, and only the reviewer knows them.This moves the check inside a native review run, while the reviewer's conversation still exists:
review_format.reportReview. It uses the same forced-then-unforced ask as the repair, now shared asaskReviewerForReport.structured, notdropped_findings) and keeps at least as many findings. Otherwise the draft stands.reviewFormatPassed, and publishing skips its own check for it.review_rewrite_faithfulexactly as before: harness (Codex / Claude Code) reviews, which have no conversation to continue, and revisions that did not fix it.Nothing here can fail a review. A check or revision that throws leaves the draft as it was.
Cost: the Jev judgments stay on the platform credential. The revision is a turn of the review on the review's own model and is billed like the rest of the review. It only happens when the check flags the draft.
docs/decisions.mdand the Decision layer paragraph in AGENTS.md now say so.findReviewFormatProblemsandprReviewDecisionScopeare extracted from the publish-time check so both stages record decisions the same way. The in-run judgment is taggedstage: reviewer_draftin its decision metadata.Tests cover a passing draft (no follow-up), a flagged draft fixed by the reviewer, a revision still flagged, a revision that drops a finding, clears the issues, or is itself rejected (draft kept in each case), failing model and judge calls, the repair-then-revise order, the runner wiring, and publishing skipping a report already passed. Each behavior was mutation-checked.