Skip to content

ci(lint, governed-surface-guard): split each full-history job wall into a checkout budget plus a body budget - #22033

Merged
objectstack-fleet[bot] merged 1 commit into
mainfrom
claude/issue-22020-checkout-wall-budget
Oct 6, 2026
Merged

objectstack-fleet[bot] merged 1 commit into
mainfrom
claude/issue-22020-checkout-wall-budget

Conversation

@objectstack-fleet

@objectstack-fleet objectstack-fleet Bot commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Fixes #22020
Clause-②: no

Which route

The card's done-when offers two routes. This PR takes the first one, in the split form triage asked for (6022336293). Every full-history checkout under a fixed job wall gets its own step-level timeout-minutes: 20. Each job wall becomes that checkout budget plus the job's re-measured body budget. No gate, command, fetch depth, job id or check name changes. The seven required contexts are untouched, and the merge-base anchor of the authorable-surface deletion gate keeps its full history (fetch-depth: 0 is unchanged on every job).

job file wall before checkout step body budget wall after
Type Check · source gates (the card) lint.yml 10 20 10 30
Type Check · workspace (adjacent) lint.yml 30 20 35 55
Type Check · debt ledger (adjacent) lint.yml 15 20 15 35
Type Check · consumer gates (adjacent) lint.yml 20 20 25 45
Governed Surface Queue Guard (adjacent) governed-surface-guard.yml 10 20 10 30

In-scope adjacent fixes. All four have the same defect class and get the same mechanical treatment:

  • The other three Type Check lanes, added by the PM after merge-group run 37509811445. In that run, job 112427734781 ("Type Check · consumer gates") spent 11m05s in the checkout and was cancelled at its 20-minute wall. The aggregate failed and PR fix(deps): take the fixes for sharp and shell-quote that turn main's OSV scan red #22016 was ejected from the queue.
  • The Governed Surface Queue Guard, which the claim named (run 37508992265, job 112424902156).

The lint job is not touched. Open PR #22002 edits its region.

Measured

Window. I read the 300 most recent completed runs of Lint & Type Check (2026-10-05T15:32Z to 2026-10-06T18:41Z) and of Governed Surface Guard (2026-10-05T14:57Z to 2026-10-06T18:59Z). Every attempt was included (GET /actions/runs/{id}/jobs?filter=all), using step timestamps. The controls come from the 300 most recent CI runs (2026-10-05T15:15Z to 2026-10-06T18:22Z).

Checkout per job. "Wall hits" counts the runs that the old wall cancelled behind a slow checkout:

job n p50 p90 p99 max wall hits
Type Check · source gates 293 31 s 122 s 599 s 600 s (cut by the wall) 8
Type Check · workspace 296 31 s 100 s 395 s 421 s 0
Type Check · debt ledger 293 31 s 123 s 599 s 774 s 2
Type Check · consumer gates 293 31 s 151 s 574 s 735 s 2 (both merge_group)
Governed Surface Queue Guard 292 31 s 156 s 595 s 599 s (cut by the wall) 4

Pooled. Adding the same fetch from ci.yml's Test Core shards gives n = 3,459: p50 31 s, p90 119 s, p99 553 s, max 895 s. Another 19 samples were cut off by a wall and are lower bounds only.

Control. The shallow (fetch-depth 1) checkouts in ci.yml over the same hours (n = 2,743) read p50 15 s, p99 63 s, max 107 s. None went over 180 s. So the slow tail belongs to the full-history fetch, not to the runner pool.

Body budgets. The body is everything after the checkout, on successful runs:

job max 2 × max budget
source gates 3.2 min 6.4 10 (the floor)
workspace 17.9 min 35.8 35
debt ledger 7.4 min 14.8 15
consumer gates 11.5 min 23 25
guard 1.1 min 2.2 10 (the floor)

The rule is the lint job's: 2× the measured max, on a 5-minute grain, never below 10.

Checkout budget. 20 minutes is 1.3× the pooled 895 s max. The rule's 2× would give 30, but the queue's window binds first: the workspace lane's 35 plus 30 is 65, past the 55 cap the lint job argues, while 35 + 20 = 55. It is one fetch, so every job gets the same budget.

What the required aggregate reads

Before. The lane's job wall stopped any checkout that left the body too little time. That covers slow checkouts that would have finished, such as the 555 s and 588 s cases, as well as stuck ones. The lane read cancelled and every gate was skipped. TypeScript Type Check printed "concluded cancelled -- expected success" and failed the PR.

After:

  • A checkout that ends inside 20 minutes: nothing happens. The body keeps its full budget and the lane goes green. In this window, that covers every checkout that completed.
  • A checkout still running at 20 minutes: the step fails with "The action 'Checkout repository' has timed out after 20 minutes." Later steps are skipped, the lane reads failure and the aggregate goes red. No layout can keep a truly hung checkout green, because the aggregate correctly refuses a lane that ran no gate. This layout names the step that hung, and it fires only beyond 1.3× anything measured.
  • A later step that hangs: the job wall cancels the job, the lane reads cancelled, and the aggregate goes red, as before but at the new wall.

Source check (premise 2). In actions/runner main at 67f01c27, StepsRunner.RunStepAsync evaluates every step's timeout-minutes, uses: steps included. When it expires, the runner sets TaskResult.Failed with that message. A job-level cancellation sets TaskResult.Canceled instead. No workflow in this repo had a step-level timeout before this PR.

Deviation from triage's direction

Triage wrote: "The gate steps keep the stall guard's intent through their own step-level timeouts." I did not add those. The PM thread has the reasoning:

  • Actions has no timeout over a group of steps. Covering the gates would take one step-level timeout per step, about 65 of them across the four lanes, each sized from a single 27-hour window.
  • Several of those steps are bimodal, for example a turbo cache hit in 4 s against a miss in 5 minutes. Timeouts sized that tightly would be a new source of random reds, which is this card's own defect class.

The cost: after a fast checkout, a hung gate now runs up to the new wall (30, 35, 45 or 55 minutes) instead of the old one. Every new wall stays at or under the 55 cap, inside the queue's 60-minute window.

This PR's own lane run

Read from the jobs API by the seat (step timestamps), run 37519814891 (Lint & Type Check, head f5060b251) and run 37519814823 (Governed Surface Guard):

job conclusion checkout job duration new wall
Type Check · source gates success 34 s 197 s (3.3 of 30 min) 30
Type Check · debt ledger success 147 s 217 s 35
Type Check · workspace success 242 s 280 s 55
Type Check · consumer gates success 33 s 298 s 45
TypeScript Type Check (aggregate) success — 3 s —
Governed Surface Queue Guard success 38 s 66 s 30

Section filled in by the domain:devx seat 2 PM from the run above, because the body is written before its own run exists.

Local gates (head f5060b251)

node scripts/pm/dispatch-gates.mjs --commands --repo objectstack-ai/objectstack derived 44 commands from the actual diff. Exit codes were written to disk before reading:

  • 42 exit 0. This includes check:required-contexts, check:stall-guard-budget, check:stall-guard-headroom, check:workflow-status-functions, check:workflow-step-name-quoting, check:aggregator-roster (+ --self-test) and check:step-collectors (+ --self-test). The full reconciliation is in the report.
  • I also ran the guard's own check-governed-queue-guard.mjs --self-test: 296 cases passed, including the fetch-depth: 0 pin.
  • check:type-check-debt ran its --self-test half (exit 0). Its --re-measure half needs the workspace build the debt lane runs first. It is declared to this PR's CI rather than measured here, because the diff touches no package and no ledger.

Acceptance notes

  • Mechanism, measured locally and not changed here. fetch-depth: 0 makes actions/checkout v7 fetch +refs/heads/* and +refs/tags/*: 1,236 branch heads and about 7,950 tags today. From this container:

    fetch time pack
    main's full history alone 29 s and 32 s 397 MB
    the all-refs refspec CI uses 473 s 598 MB
    all heads without tags 686 s 597 MB
    depth 1 plus unshallowing main 120 s and 169 s not recorded

    The last variant kept the anchor exact on PR chore(pm): delete report-only check-widening-tells.mjs and its wiring (ruling 208) #22002's merge ref and on PR refactor(drivers)!: memory / mongodb 的 aggregate / distinct 收进 DriverQuery (#6212 批 C) #6356's old head. The cost is the ref set, not main's depth. Fetching less could make every one of these jobs faster, but the CI tail of any narrower fetch is unmeasured (n = 0), and its semantics need a per-gate audit across all five jobs. It is left as a question for the PM in the report.

  • ci.yml Test Core shards, same fetch, 45-minute wall: one shard was cancelled at the wall behind a 508 s checkout in the window (run 37338337323, Test Core (2/6)). This is outside this PR's file surface, and those steps run under the stall guard, whose budget gate reads that wall. The report names it for the seat.

  • Lint & Repo Gates is not exposed in the window: body max 34.5 min plus checkout max 730 s is 46.7, under its 55.


Generated by Claude Code

…to a checkout budget plus a body budget

The four Type Check lanes and the Governed Surface Queue Guard check out with
`fetch-depth: 0`, whose time follows GitHub's git server rather than the change
under test. Re-measured over the 300 most recent runs (pooled n = 3,459 with
ci.yml's Test Core shards): p50 31 s, p99 553 s, max 895 s, while the shallow
checkouts of the same hours never passed 107 s. Walls sized when that fetch took
~30 s were cancelling lanes behind it (source gates 8, debt 2, consumers 2,
guard 4 in the window), and the required aggregates failed each with no gate run.

Each checkout step now carries its own 20-minute budget and each wall is that
budget plus the lane's re-measured body budget (2x its max, 5-minute grain,
floor 10): source gates 30, debt 35, consumers 45, workspace 55, guard 30. No
gate, command, check name or fetch depth changes.

Claude-Session: https://claude.ai/code/session_01VF48aw8RPG6wzDnMgp6rtw
Co-authored-by: Claude <noreply@anthropic.com>
@github-actions github-actions Bot added the size/m label Oct 6, 2026
@objectstack-fleet objectstack-fleet Bot added the skip-changeset PR has no user-facing published change; bypasses the changeset gate label Oct 6, 2026
@github-actions github-actions Bot added the ci/cd label Oct 6, 2026
@objectstack-fleet
objectstack-fleet Bot marked this pull request as ready for review October 6, 2026 20:04
@objectstack-fleet
objectstack-fleet Bot added this pull request to the merge queue Oct 6, 2026
Merged via the queue into main with commit 099a94d Oct 6, 2026
39 checks passed
@objectstack-fleet
objectstack-fleet Bot deleted the claude/issue-22020-checkout-wall-budget branch October 6, 2026 20:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci/cd size/m skip-changeset PR has no user-facing published change; bypasses the changeset gate

Projects

None yet

2 participants