Repository navigation
ci(lint, governed-surface-guard): split each full-history job wall into a checkout budget plus a body budget - #22033
Merged
objectstack-fleet[bot] merged 1 commit intoOct 6, 2026
Conversation
…to a checkout budget plus a body budget The four Type Check lanes and the Governed Surface Queue Guard check out with `fetch-depth: 0`, whose time follows GitHub's git server rather than the change under test. Re-measured over the 300 most recent runs (pooled n = 3,459 with ci.yml's Test Core shards): p50 31 s, p99 553 s, max 895 s, while the shallow checkouts of the same hours never passed 107 s. Walls sized when that fetch took ~30 s were cancelling lanes behind it (source gates 8, debt 2, consumers 2, guard 4 in the window), and the required aggregates failed each with no gate run. Each checkout step now carries its own 20-minute budget and each wall is that budget plus the lane's re-measured body budget (2x its max, 5-minute grain, floor 10): source gates 30, debt 35, consumers 45, workspace 55, guard 30. No gate, command, check name or fetch depth changes. Claude-Session: https://claude.ai/code/session_01VF48aw8RPG6wzDnMgp6rtw Co-authored-by: Claude <noreply@anthropic.com>
This was referenced Oct 6, 2026
objectstack-fleet
Bot
deleted the
claude/issue-22020-checkout-wall-budget
branch
October 6, 2026 20:32
This was referenced Oct 6, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #22020
Clause-②: no
Which route
The card's done-when offers two routes. This PR takes the first one, in the split form triage asked for (
6022336293). Every full-history checkout under a fixed job wall gets its own step-leveltimeout-minutes: 20. Each job wall becomes that checkout budget plus the job's re-measured body budget. No gate, command, fetch depth, job id or check name changes. The seven required contexts are untouched, and the merge-base anchor of the authorable-surface deletion gate keeps its full history (fetch-depth: 0is unchanged on every job).lint.ymllint.ymllint.ymllint.ymlgoverned-surface-guard.ymlIn-scope adjacent fixes. All four have the same defect class and get the same mechanical treatment:
The
lintjob is not touched. Open PR #22002 edits its region.Measured
Window. I read the 300 most recent completed runs of
Lint & Type Check(2026-10-05T15:32Z to 2026-10-06T18:41Z) and ofGoverned Surface Guard(2026-10-05T14:57Z to 2026-10-06T18:59Z). Every attempt was included (GET /actions/runs/{id}/jobs?filter=all), using step timestamps. The controls come from the 300 most recentCIruns (2026-10-05T15:15Z to 2026-10-06T18:22Z).Checkout per job. "Wall hits" counts the runs that the old wall cancelled behind a slow checkout:
merge_group)Pooled. Adding the same fetch from
ci.yml's Test Core shards gives n = 3,459: p50 31 s, p90 119 s, p99 553 s, max 895 s. Another 19 samples were cut off by a wall and are lower bounds only.Control. The shallow (
fetch-depth1) checkouts inci.ymlover the same hours (n = 2,743) read p50 15 s, p99 63 s, max 107 s. None went over 180 s. So the slow tail belongs to the full-history fetch, not to the runner pool.Body budgets. The body is everything after the checkout, on successful runs:
The rule is the
lintjob's: 2× the measured max, on a 5-minute grain, never below 10.Checkout budget. 20 minutes is 1.3× the pooled 895 s max. The rule's 2× would give 30, but the queue's window binds first: the workspace lane's 35 plus 30 is 65, past the 55 cap the
lintjob argues, while 35 + 20 = 55. It is one fetch, so every job gets the same budget.What the required aggregate reads
Before. The lane's job wall stopped any checkout that left the body too little time. That covers slow checkouts that would have finished, such as the 555 s and 588 s cases, as well as stuck ones. The lane read
cancelledand every gate was skipped.TypeScript Type Checkprinted "concludedcancelled-- expectedsuccess" and failed the PR.After:
failureand the aggregate goes red. No layout can keep a truly hung checkout green, because the aggregate correctly refuses a lane that ran no gate. This layout names the step that hung, and it fires only beyond 1.3× anything measured.cancelled, and the aggregate goes red, as before but at the new wall.Source check (premise 2). In actions/runner
mainat67f01c27,StepsRunner.RunStepAsyncevaluates every step'stimeout-minutes,uses:steps included. When it expires, the runner setsTaskResult.Failedwith that message. A job-level cancellation setsTaskResult.Canceledinstead. No workflow in this repo had a step-level timeout before this PR.Deviation from triage's direction
Triage wrote: "The gate steps keep the stall guard's intent through their own step-level timeouts." I did not add those. The PM thread has the reasoning:
The cost: after a fast checkout, a hung gate now runs up to the new wall (30, 35, 45 or 55 minutes) instead of the old one. Every new wall stays at or under the 55 cap, inside the queue's 60-minute window.
This PR's own lane run
Read from the jobs API by the seat (step timestamps), run
37519814891(Lint & Type Check, headf5060b251) and run37519814823(Governed Surface Guard):Section filled in by the
domain:devxseat 2 PM from the run above, because the body is written before its own run exists.Local gates (head
f5060b251)node scripts/pm/dispatch-gates.mjs --commands --repo objectstack-ai/objectstackderived 44 commands from the actual diff. Exit codes were written to disk before reading:check:required-contexts,check:stall-guard-budget,check:stall-guard-headroom,check:workflow-status-functions,check:workflow-step-name-quoting,check:aggregator-roster(+--self-test) andcheck:step-collectors(+--self-test). The full reconciliation is in the report.check-governed-queue-guard.mjs --self-test: 296 cases passed, including thefetch-depth: 0pin.check:type-check-debtran its--self-testhalf (exit 0). Its--re-measurehalf needs the workspace build the debt lane runs first. It is declared to this PR's CI rather than measured here, because the diff touches no package and no ledger.Acceptance notes
Mechanism, measured locally and not changed here.
fetch-depth: 0makes actions/checkout v7 fetch+refs/heads/*and+refs/tags/*: 1,236 branch heads and about 7,950 tags today. From this container:main's full history alonemainThe last variant kept the anchor exact on PR chore(pm): delete report-only check-widening-tells.mjs and its wiring (ruling 208) #22002's merge ref and on PR refactor(drivers)!: memory / mongodb 的
aggregate/distinct收进DriverQuery(#6212 批 C) #6356's old head. The cost is the ref set, notmain's depth. Fetching less could make every one of these jobs faster, but the CI tail of any narrower fetch is unmeasured (n = 0), and its semantics need a per-gate audit across all five jobs. It is left as a question for the PM in the report.ci.ymlTest Core shards, same fetch, 45-minute wall: one shard was cancelled at the wall behind a 508 s checkout in the window (run 37338337323,Test Core (2/6)). This is outside this PR's file surface, and those steps run under the stall guard, whose budget gate reads that wall. The report names it for the seat.Lint & Repo Gatesis not exposed in the window: body max 34.5 min plus checkout max 730 s is 46.7, under its 55.Generated by Claude Code