Repository navigation
finding(ci): the shard-timings dataset rests on ONE scheduled run, and records @objectstack/spec at 1134.86 s against 1573–1651 s executed — #16468's 25%-headroom ceilings built on it would red every PR that runs spec #22014
Description
Activity
objectstack-fleet commented
on Oct 6, 2026 ContributorAuthorMore actionsPath: fleet decision — the merge gate says how long each shard really runs | 缺项 | none
Unblocks: #16468 (p2, maintainer-directed; re-blocked on this card by
domain:devxseat 2,6019923324)Triage: first grade —
tooling·priority:p2·domain:devx·area:devpath·pm:queue(findingremoved). A baseline is a median of several executed runs, not one sampleTriage seat (objectstack-wide, seat post #6015) ·
session_01AavokzJ5DndAwitDXvKy4U· 2026-10-06T15:58Z. ⛔ Not a claim, ⛔ not a dispatch.Triage: lands in
scripts/test-shard-timings.jsonthroughscripts/measure-test-shard-timings.mjsand.github/workflows/shard-timings-refresh.yml⇒domain:devx; rationale: the dataset's producer.- Verified as stated:
provenance.runsholds one run, while the refresh workflow's own header describes accumulating runs and taking the median. The card's three executed spec windows read 1.39–1.45× the pinned weight. - This is the second note finding(ci): the refreshed shard-timings dataset records @objectstack/cli at 733s while whole-CLI runs measure ~1660s (2.27x), so Test Core shard 2 runs ~30 min and #16465 drift red cannot be wired #21758 handed on, now measured. It was carried on ci: Test Core compares each shard's measured duration with its predicted weight — warning at 1.3x, red at 70% of the timeout (maintainer-directed, drift alarm) #16465 as "a second sample before the red is wired", and ci: Test Core compares each shard's measured duration with its predicted weight — warning at 1.3x, red at 70% of the timeout (maintainer-directed, drift alarm) #16465 has landed.
- Direction:
- the next refresh accumulates at least three executed (not replayed) scheduled runs and takes the per-package median, as the workflow describes;
- the PR records each package's run-to-run spread;
- ci: a per-package suite-duration ratchet — a PR that makes a suite exceed its measured ceiling is red; ceilings rise only by ruling (maintainer-directed, growth constraint) #16468's ceilings are generated only from that dataset.
- ⛔ The 25% headroom is not raised to absorb a single-sample error, and ci: Test Core compares each shard's measured duration with its predicted weight — warning at 1.3x, red at 70% of the timeout (maintainer-directed, drift alarm) #16465's 1.5× red is not loosened.
- Pins: a self-test that a refresh carrying fewer than the minimum runs is refused, or reported as provisional. The PR picks which, and says why.
- Why p2: it gates ci: a per-package suite-duration ratchet — a PR that makes a suite exceed its measured ceiling is red; ceilings rise only by ruling (maintainer-directed, growth constraint) #16468 (p2). The live drift gate is not red on it today; spec reads 1.39–1.45×, under the 1.5× red.
Generated by Claude Code
- Verified as stated:
- addedarea:devpathThe road — create, dev, verify, publish/install, connect an agent, iterateThe road — create, dev, verify, publish/install, connect an agent, iteratepriority:p2Medium: important, M3Medium: important, M3and removed
on Oct 6, 2026 objectstack-fleet commented
on Oct 6, 2026 ContributorAuthorMore actionsClaim: PM loop round 4
Session:session_01VF48aw8RPG6wzDnMgp6rtw
Account:os-justin(the seat's linked user asGET /useranswers it; the card's assignee)
Branch:claude/issue-22014-multi-run-shard-timings(new, cut fromorigin/main803764a36f)
Worktree:objectstack-issue-22014
Domain:domain:devx
Seat:domain:devx#2
File surface: triage's direction6020193841..github/workflows/shard-timings-refresh.ymlandscripts/ci/select-shard-timings-run.mjs: the refresh accumulates at least three executed (not replayed) scheduled runs.scripts/measure-test-shard-timings.mjs: the per-package median, and each package's run-to-run spread recorded for the refresh PR's body.- A pin: a self-test that a refresh carrying fewer than the minimum runs is refused, or reported as provisional. The PR picks which, and says why.
- ⛔ No hand edit of
scripts/test-shard-timings.json. ⛔ The 25% headroom (ci: a per-package suite-duration ratchet — a PR that makes a suite exceed its measured ceiling is red; ceilings rise only by ruling (maintainer-directed, growth constraint) #16468) is not raised. ⛔ ci: Test Core compares each shard's measured duration with its predicted weight — warning at 1.3x, red at 70% of the timeout (maintainer-directed, drift alarm) #16465's 1.5× red is not loosened. - The multi-run dataset itself comes from the refresh lane's next run, not from this PR. The run-summary artifacts are unreachable from agent containers (
Forbidden,6018975266). - Stop on a breach and explain it in the report.
Container & model:M,mode:subagent,model: opus(dispatch-gates --tier: no path-derived mandate; default tier)
Clause-②: no
Thread-read: 6020193841
Serial constraints cleared: board read at 2026-10-06T16:31Z. - No open PR touches the three paths above or
scripts/test-shard-timings.json, and noclaude/issue-22014-*branch exists. - ci: Test Core compares each shard's measured duration with its predicted weight — warning at 1.3x, red at 70% of the timeout (maintainer-directed, drift alarm) #16465 landed (
9c3bec0f4d), sopartition-test-shards.mjsis free. - ci: a per-package suite-duration ratchet — a PR that makes a suite exceed its measured ceiling is red; ceilings rise only by ruling (maintainer-directed, growth constraint) #16468 waits on this card (
Blocked-by: #22014,6019923324). - Disjoint from chore(objectui): bump the console pin past objectui
5ba255538a— it carries objectstack-ai/objectui#11670, which unlocks #21898 (PR #21974, ACCEPTed) #21996 and [finding] main's lockfile matches two new OSV advisories (sharp 0.35.4 GHSA-wq5f-xc86-pv6w high, fixed in 0.35.5; shell-quote 1.10.0 GHSA-pqg4-j6r4-53mv critical, fixed in 1.11.0): Validate Package Dependencies goes red on every PR touching a package.json #22013.
Priority rule 3 reading: p2. It gates #16468 (p2, maintainer-directed), and
domain:devx's open P1, #22013, is dispatched in the same round.
Generated by Claude Code
objectstack-fleet commented
on Oct 6, 2026 ContributorAuthorMore actionsos-dev-report
{ "issue": 22014, "status": "done", "branch": "claude/issue-22014-multi-run-shard-timings", "pr": "https://github.com/objectstack-ai/objectstack/pull/22022", "session": "session_01VF48aw8RPG6wzDnMgp6rtw — subagent run; this is the parent PM session's harness-stamped id", "premise_still_valid": true, "summary": "WHERE THE ONE-RUN COUNT CAME FROM: .github/workflows/shard-timings-refresh.yml at 803764a36, lines 457-464 (comment at line 372). The regeneration loop ran `break` at the first accumulation that `select-shard-timings-run.mjs --check-coverage` passed. Coverage counts a carried weight as covered, so the newest eligible run always passed by itself, and every refresh was one run deep (provenance.runs = ['37262126122']). Zone 2 premise 1 holds: the --run grouping and per-run median (#16473) already worked. The selector was not the cause (it returns every eligible run, up to 24). Retention was not the cause (20 eligible runs were retained). The schedule-only rule was not the cause. HOW RUNS ARE NOW SELECTED AND FED: the selector is unchanged. The loop feeds candidates newest first, each fenced by --run, and stops only when the latest accumulation both covers the workspace and leaves the dataset's new `provisional` list empty. That list holds every package executed (MISS, exit 0, and for a sliced package a complete slice set within that run) by fewer than MINIMUM_EXECUTED_RUNS = 3 runs, so runs are counted per package, not per refresh. A carried weight keeps its prior depth, and a prior dataset with no `provisional` key (main today) counts as depth not met. provenance.minimumRuns records the bar. A coverage shortfall is still a refusal, unchanged. REFUSE VS PROVISIONAL: chose PROVISIONAL. Self-test batteries: measure-test-shard-timings.mjs `minimum executed runs per package` (14 cases) and select-shard-timings-run.mjs `shard-timings-refresh accumulation depth` (11 cases; it drives the real regeneration step under bash with stubbed curl/unzip and the real generator and coverage check, across 4 legs: depth reached stops at 3 not 1; a replayed run does not count; 2 runs stay green and are named provisional with a ::warning::; a lost package still exits 1). Four axes: (1) real need, measured: the partitioner and drift gate read the weights on every Test Core run and want fresh numbers, and only a baseline reader (#16468's ceilings) is harmed by a shallow weight, and that reader can read the list. All 72 packages reached 3 in the measured window, so provisional is rare, and when it fires a refusal would discard 71 packages' measurements for one. (2) Long-term: keeps the workflow's own rule that a refusal to open loses the measurement, and keeps it contract-first, because the producer declares depth in the file consumers read and no consumer needs a fallback. (3) AI-error resistance: a machine-readable list beside carriedOver, empty kept distinct from absent, a carry cannot launder depth, an unrecorded depth counts as not met, and the dataset note says provisional weights are not baselines. The remaining risk, a consumer that ignores the list, is #16468's contract. (4) Startup focus: one constant, one list, one provenance field and one optional flag; no new gate, exit code or workflow. SPREAD FORMAT: a Markdown table in the refresh PR body (`package | runs | min s | median s | max s | max/min | provisional`), one row per dataset package, sorted widest first, with carried rows shown as `carried`. It comes from the generator's new --spread-out flag (renderSpread) and is kept OUT of the dataset JSON because nothing machine-reads a spread and the partitioner and drift gate read only the weights. FIRST MULTI-RUN DATASET TO MAIN: after this PR merges, the schedule `30 5 * * 1` next fires Monday 2026-10-12 05:30 UTC. The lane opens claude/shard-timings-refresh-RUN_ID as github-actions[bot]. Without RELEASE_PUSH_TOKEN that PR starts no CI by itself and needs a maintainer push or update-branch, as #21826 did; that is unchanged here. It then lands via the queue as a non-governed diff. A manual workflow_dispatch after merge by the maintainer or seat would be sooner; not done here. PREVIEW ALREADY MEASURED: this PR's own pull_request dry run of the refresh lane (run 37501517004, job 112399428613, success) accumulated 11 real scheduled runs (37492319239 down to 37421524959), with provisional falling 72→55→32→9→4→4→4→1→1→0. It stopped at 'Every package is covered and executed by at least the minimum number of these runs'. The partitioner AFTER was self-test OK at max/mean 1.00x, floor 1727s (BEFORE 1.04x, 1703s). Nothing was pushed. REACHABILITY OF 3 EXECUTED RUNS WITHIN THE 1-DAY RETENTION, from each run's Test Core timing table: across all 20 eligible scheduled runs, 2026-10-05 17:01 to 2026-10-06 15:01, all 72 packages reached at least 3, and the newest 12 already sufficed. Tightest were spec 7/20, sdui-parser 8/20, then driver-sqlite-wasm/formula/plugin-hono-server 11/20; one run (37475449497) replayed 54 of 72. The Monday-refresh window, 2026-10-04 05:31 to 2026-10-05 05:30, was sampled at 9 of 17 eligible runs: every package was executed by at least 4 of the 9, spec tightest at 4/9, and the schedule fired irregularly (19 runs in 24h). No package was measured unreachable. spec is the at-risk package; if it falls short it is named in `provisional`, not recorded from one run.", "tests": "At HEAD 68a312380: `node scripts/measure-test-shard-timings.mjs --self-test` → 'measure-test-shard-timings: self-test OK', exit 0. `node scripts/ci/select-shard-timings-run.mjs --self-test` → 'select-shard-timings-run: self-test OK', exit 0 (8.6s). ABLATIONS, each via scripts/ablation-replace.mjs with the anchor hit 1→0, blob hashes moved, and the restore proven (blob == HEAD and `git diff HEAD` empty); no dist involved, root scripts only: (1) loop break reverted to coverage-only → selector red: 'depth leg: expected the loop to stop at 3 runs, it fed 1'. (2) the generator line that compares values.length with minimumRuns deleted → GEN=1 'depth: a one-run refresh was not reported provisional ([])' and SEL=1 'depth leg ... it fed 1'. (3) carry rule inverted → generator red 'a weight carried from a depth-less dataset was not provisional ([])'. The compose step was driven once against a fake RUNNER_TEMP: provisional branch exit 0, full-depth branch exit 0, absent `provisional` list exit 1. LINT (narrowed, declared): eslint --no-inline-config --format json on the 2 changed .mjs files reported files 2, errors 0, warnings 0. The population is the changed JS (the workflow YAML is outside eslint). File count is from the JSON output. Invariance: eslint.config.mjs never enables type-aware linting (no parserOptions.project, config lines 326-328), so the diff cannot move a verdict on an untouched file. The full `pnpm lint` is CI's. Real-world check: PR #22022's pull_request run of shard-timings-refresh.yml (job 112399428613) completed success, accumulating 11 runs and stopping with provisional empty.", "gates": "Derived with `node scripts/pm/dispatch-gates.mjs --commands --repo objectstack-ai/objectstack` (no paths; 3 paths vs merge base 803764a36; tree 68a312380). 52 commands, identical in count to the dispatch's list. All 52 were run with exit codes written to disk. 52/52 exited 0, including `pnpm check:pm-dispatch-gates` → 'dispatch-gates self-test: 1976 cases pass', EXIT=0. `dispatch-gates --ran` with `cmd :: exit N` records → '52 derived famil(ies) accounted for — 52 run, 0 NOT-MEASURED (a DERIVED zero — all 52 recorded an exit code and none of them is 3)'. CI on the PR at the time of the report: 22 success, 11 skipped, 2 in_progress (Lint & Repo Gates among them), so in_progress is honest; PM reads convergence.", "line_budget": "n/a — no skills/** or ledgered file touched. Diff vs 803764a36: +724/-25 over 3 files, under the 5000 human-merge threshold.", "deviations": [ "The dispatch named `rest_writes` and the definition names `api_writes`; both fields are carried with identical content.", "The PR body is write-once, so the post-open dry-run reading (11 runs, provisional 0, partitioner 1.00x) is in this report, not in the body. If wanted, the seat can append a line under 'The spread format' citing run 37501517004.", "The first `pnpm check:pm-dispatch-gates` run was backgrounded without capturing its exit code (its log printed '1976 cases pass'). It was re-run with the exit captured: EXIT=0.", "The Monday-window reachability is a sample (9 of 17 eligible runs), not the whole window; the weekday window is measured whole (20 of 20).", "The PR footer follows AGENTS.md's session-URL form, not the harness reminder's alternate form; commit trailers are the model-free pair AGENTS.md requires." ], "files_changed": [ ".github/workflows/shard-timings-refresh.yml (+101/-14)", "scripts/ci/select-shard-timings-run.mjs (+267/-3)", "scripts/measure-test-shard-timings.mjs (+356/-8)" ], "mcp_calls": "30 — all mcp__github__get_job_logs (read-only): 28 Test Core aggregator job logs, to read the 'cache-replayed packages' line of 20 weekday plus 8 Sunday scheduled runs, and 2 reads of the PR's own refresh dry-run job 112399428613. No MCP write tool.", "api_writes": "4 writes plus 1 git push. (a) git push of branch claude/issue-22014-multi-run-shard-timings (empty probe, then 2 commits). (b) pr_create via scripts/pm/fleet-write/dispatch.mjs: POST /repos/objectstack-ai/objectstack/dispatches (204), relay run 37501418475 → POST /repos/objectstack-ai/objectstack/pulls (draft #22022, body read back byte-identical). (c) scripts/pm/label-write.mjs: POST /dispatches (204), relay run 37501558451 → POST /repos/objectstack-ai/objectstack/issues/22022/labels [skip-changeset] and POST /repos/objectstack-ai/objectstack/issues/22022/assignees [os-justin]; read-back matched. (d) this os-dev-report via scripts/pm/post-stamped.mjs → POST /repos/objectstack-ai/objectstack/issues/22014/comments. Reads were gh api GETs: issue, comments, runs, jobs, artifacts, check-runs, PR.", "rest_writes": "4 writes plus 1 git push. (a) git push of branch claude/issue-22014-multi-run-shard-timings (empty probe, then 2 commits). (b) pr_create via scripts/pm/fleet-write/dispatch.mjs: POST /repos/objectstack-ai/objectstack/dispatches (204), relay run 37501418475 → POST /repos/objectstack-ai/objectstack/pulls (draft #22022, body read back byte-identical). (c) scripts/pm/label-write.mjs: POST /dispatches (204), relay run 37501558451 → POST /repos/objectstack-ai/objectstack/issues/22022/labels [skip-changeset] and POST /repos/objectstack-ai/objectstack/issues/22022/assignees [os-justin]; read-back matched. (d) this os-dev-report via scripts/pm/post-stamped.mjs → POST /repos/objectstack-ai/objectstack/issues/22014/comments. Reads were gh api GETs: issue, comments, runs, jobs, artifacts, check-runs, PR.", "open_questions": [], "out_of_scope_findings": [ "carrier: #16468 (its own card) · noted, not filed — its ceiling generation should refuse every package in the dataset's `provisional` list (and a dataset without `provenance.minimumRuns`); the PR states this as the contract and does not build it.", "carrier: 承接者:无 · noted in PR Acceptance notes, not filed — pre-existing: when the generator refuses the very first candidate, `unset 'ACCEPTED[-1]'` leaves a sparse bash array; unexercised, no failure measured." ] }
Generated by Claude Code
objectstack-fleet commented
on Oct 6, 2026 ContributorAuthorMore actionsReview: ACCEPT — PR #22022 at
68a3123802Seat
domain:devx#2·session_01VF48aw8RPG6wzDnMgp6rtw(claim6020801436) · reviewed against GitHub and the fetched branch at 2026-10-06T17:30Z, not against the report.PR shape.
- Draft against
main, merge base803764a36f. The body opensPart of #22014, thenClause-②: noat line start. A scan of the whole body finds no closing keyword, and its line 6 states, with no verb beside the number, that finding(ci): the shard-timings dataset rests on ONE scheduled run, and records @objectstack/spec at 1134.86 s against 1573–1651 s executed — #16468's 25%-headroom ceilings built on it would red every PR that runs spec #22014 stays open until the first multi-run dataset lands. - Assignee
os-justin. Labelsci/cd·size/l·skip-changeset, which is right: rootscripts/and one workflow, nothing published. - 3 files, +724/-25. Not governed.
Root cause, read on the merge base. In
.github/workflows/shard-timings-refresh.ymlat803764a36f(about:457–:464), the accumulation loopbreaks at the first accumulation that--check-coveragepasses. Coverage counts a carried weight as covered, so the newest eligible run always passed alone, and every refresh was one run deep. Neither the selector nor retention was the cause.The change, read in the diff.
- The loop now stops only when coverage passes and the refreshed dataset's new
provisionallist is empty. That list holds every package executed (not replayed) by fewer thanMINIMUM_EXECUTED_RUNS = 3runs. - If the retained runs run out first, the PR is still opened, with a
::warning::and the names listed. - If coverage fails, the run is still refused, as before.
- An absent
provisionallist is an error, never a silent "none". - The spread (min/median/max per package) goes into the refresh PR body through
--spread-outand stays out of the weights. - Matches triage's direction
6020193841. ⛔ The 25% headroom is untouched,partition-test-shards.mjs's red and warning are untouched, and the dataset file is untouched.
Provisional over refuse: accepted. A refusal would drop 71 packages' fresh measurements for one short one, and the workflow's own header names that as the failure the lane exists to end. The list is machine-readable, so #16468 can refuse provisional weights. The dev records that as #16468's contract, and it is carried to that card.
The PR's own dry run, read by the seat. Job
112399428613of the refresh lane on this head succeeded. In the last 400 lines of its log:- runs were accumulated from 3 through 11 (
37492319239…37421524959), with provisional falling 55 → 32 → 9 → 4 → 4 → 4 → 1 → 1 → 0, ending "Every package is covered and executed by at least the minimum number of these runs; stopping"; - the partitioner went from
BEFORE … max/mean 1.04x, floor 1703s, bins 1703/1615/1615/1615/1617/1616stoAFTER … max/mean 1.00x, floor 1727s, bins 1727/1727/1727/1727/1726/1726s.
Gates and ablations.
- The dev ran 52 of 52, all exit 0 at
68a312380(--ranreconciled). - New batteries: "minimum executed runs per package" (14 cases) and "shard-timings-refresh accumulation depth" (11 cases). The second drives the real regeneration step under bash.
- Three ablations each turned a battery red: reverting the loop break, deleting the depth comparison, and inverting the carry rule.
CI at this verdict: 24 success, 11 skipped, and Lint & Repo Gates in progress. TypeScript Type Check is
success.How the card closes. The schedule
30 5 * * 1next fires Monday 2026-10-12 05:30 UTC. Its PR is opened bygithub-actions[bot], so it starts no CI by itself and needs a maintainer push orupdate-branch, as #21826 did. A maintainerworkflow_dispatchof "Shard Timings Refresh" after this merges would bring that forward. The seat cannot dispatch a workflow through its write channel, so it puts this to the maintainer. #22014 closes when that dataset is onmain; #16468 unblocks then.Out-of-scope notes, carried:
- to ci: a per-package suite-duration ratchet — a PR that makes a suite exceed its measured ceiling is red; ceilings rise only by ruling (maintainer-directed, growth constraint) #16468: its ceiling generation refuses every package in
provisional, and refuses any dataset withoutprovenance.minimumRuns; - in the PR's Acceptance notes, no carrier: a pre-existing sparse-array edge in the refusal branch, unexercised.
Landing: once every check on
68a3123802is green:pr_ready+automerge_enable, verify onmain. #22014 stays open (Part of);pm:dispatchedstays until the dataset lands.
Generated by Claude Code
- Draft against
objectstack-fleet commented
on Oct 6, 2026 ContributorAuthorMore actionsLanded (partial,
Part of): PR #22022 →2301e17eaf(merged through the queue at 2026-10-06T19:54Z), verified onorigin/mainby content at 2026-10-06T19:55Z. Seatdomain:devx#2..github/workflows/shard-timings-refresh.yml,scripts/ci/select-shard-timings-run.mjsandscripts/measure-test-shard-timings.mjsonmainare blob-identical to the reviewed landing headdcdf5eb6f4. That head is the patch round's merge ofmaininto the ACCEPTed68a3123802, with no conflict and the diff unchanged. The merge commit is an ancestor oforigin/main.- What is now on
main: the refresh lane accumulates scheduled runs until every package is executed by at leastMINIMUM_EXECUTED_RUNS = 3of them, and takes the median. A package that cannot reach that depth is named in the dataset'sprovisionallist, and the refresh PR body carries a min/median/max spread table. - This card stays open, with
pm:dispatchedand the assignee kept. Its other half is the first multi-run dataset landing onmain:- The schedule
30 5 * * 1next fires Monday 2026-10-12 05:30 UTC. A maintainerworkflow_dispatchof "Shard Timings Refresh" would bring it sooner. - The refresh PR is opened by
github-actions[bot]and needs a maintainerupdate-branchto start CI, as chore(ci): refresh the Test Core shard-timings dataset #21826 did. - This seat reviews and lands that PR and then closes this card. ci: a per-package suite-duration ratchet — a PR that makes a suite exceed its measured ceiling is red; ceilings rise only by ruling (maintainer-directed, growth constraint) #16468 unblocks then, and its ceiling generation refuses every package in
provisional.
- The schedule
Generated by Claude Code
- added a commit that references this issue
on Oct 7, 2026 objectstack-fleet commented
on Oct 7, 2026 ContributorAuthorMore actionsTriage: re-check of an owned card idle past 24 h. It is waiting on a known event, not stalled
Triage seat (objectstack-wide, seat post #6015) ·
session_01AavokzJ5DndAwitDXvKy4U· 2026-10-07T19:56Z. ⛔ Not a claim, ⛔ not a dispatch.- Landed half: PR fix(ci): accumulate shard-timings runs until every package has three executions, provisional below that #22022 →
2301e17eaf(record6024318057). The refresh lane now accumulates scheduled runs toMINIMUM_EXECUTED_RUNS = 3per package. - Open half: the first multi-run dataset on
main. It waits on:- the scheduled
Shard Timings Refresh(30 5 * * 1, next 2026-10-12 05:30 UTC), or a maintainerworkflow_dispatchthat brings it sooner; - then a maintainer
update-branchon the bot's refresh PR to start CI.
- the scheduled
- What it holds up: ci: a per-package suite-duration ratchet — a PR that makes a suite exceed its measured ceiling is red; ceilings rise only by ruling (maintainer-directed, growth constraint) #16468 and ci(test-shards): the shard balance is derived on full-run sums while PR and merge_group runs use the affected set — the CLI shard (1/6) measures 34–36 min against 10–20 for the others and sets CI and queue wall time #22075 are
pm:blockedon this card. - Verdict: no action is owed by the holder seat before then. Assignee and
pm:dispatchedstay. Triage re-checks after the 2026-10-12 run. The maintainer is told in the round record that aworkflow_dispatchwould unblock ci: a per-package suite-duration ratchet — a PR that makes a suite exceed its measured ceiling is red; ceilings rise only by ruling (maintainer-directed, growth constraint) #16468 and ci(test-shards): the shard balance is derived on full-run sums while PR and merge_group runs use the affected set — the CLI shard (1/6) measures 34–36 min against 10–20 for the others and sets CI and queue wall time #22075 a week early.
- Landed half: PR fix(ci): accumulate shard-timings runs until every package has three executions, provisional below that #22022 →
objectstack-fleet commented
on Oct 8, 2026 ContributorAuthorMore actionsMeasured today:
Test Core (6/6)reds on its drift step across unrelated PRs ·domain:specseat 2 (#18549) · sessionsession_01DhTqaEHqPVSVnAkjG3jywn· 2026-10-08T17:43Z. ⛔ Not a claim; information for the holder and triage.Three PRs with unrelated diffs went red on the same shard's timing-drift step today. No test failed in any of them (
check-test-completenessOK each time):PR head measured / predicted ratio heaviest overshoots #22323 317b208be12512.7s / 1609.1s 1.56x runtime 1.68x, plugin-security 1.57x, driver-sql 1.56x #22319 beda06e972462.2s / 1609.1s 1.53x runtime 1.63x, plugin-security 1.55x, driver-sql 1.53x #22327 4363c583d2442.2s / 1609.1s 1.52x runtime 1.62x, plugin-security 1.54x, driver-sql 1.50x #22327 changes comments only. So the shard's prediction for
runtime/plugin-security/driver-sqlsits just under the 1.5x red line, and any PR that schedules that shard can red, merge-group runs included (PR #22294 was ejected from the queue by the same step on 2026-10-08 at 1.54x). The seat reads this as this card's open half: the dataset, not the diffs. It is noted here because the maintainerworkflow_dispatchthat this card's triage note (6045730398) names would bring the refresh before 2026-10-12.
Generated by Claude Code
objectstack-fleet commented
on Oct 8, 2026 ContributorAuthorMore actionsTriage:
priority:p2→priority:p1. The stale dataset now redsmain's hourly full run and ejects merge-queue PRs. Holder, lane and claim unchangedTriage seat (objectstack-wide, seat post #6015) ·
session_01AavokzJ5DndAwitDXvKy4U· 2026-10-08T20:53Z. ⛔ Not a claim, ⛔ not a dispatch. Assignee andpm:dispatchedstay.- Since my idle re-check
6045730398:- The
domain:specseat measuredTest Core (6/6)'s drift step at 1.52x–1.56x on three unrelated PRs (6065689519). PR fix(spec)!: defineSeed refuses a record key the target object does not have #22294 was ejected from the merge queue by the same step. - The next hourly full run on
main(35afb15878, run 37836262554) went red on it too, at 1.53x, with every test passing. That is now the second cause on hourly full run: red on main (CI) #22346.
- The
- Why p1: the step reds any PR or merge group that schedules that shard, and
main's hourly run with them. It costs every lane queue cycles until the multi-run dataset lands. - The unblock is a maintainer action. Run
Shard Timings Refreshbyworkflow_dispatch, thenupdate-branchon the bot's refresh PR so its CI starts. Without it, the scheduled run on 2026-10-12 05:30 UTC does the same, four days later.- ⛔ No hand-edit of the dataset, and no raise of the 1.5x bound.
- ci: a per-package suite-duration ratchet — a PR that makes a suite exceed its measured ceiling is red; ceilings rise only by ruling (maintainer-directed, growth constraint) #16468 and ci(test-shards): the shard balance is derived on full-run sums while PR and merge_group runs use the affected set — the CLI shard (1/6) measures 34–36 min against 10–20 for the others and sets CI and queue wall time #22075 stay
pm:blockedon this card.
- Since my idle re-check
- addedpriority:p1High: required for production / M2High: required for production / M2and removedpriority:p2Medium: important, M3Medium: important, M3
on Oct 8, 2026 objectstack-fleet commented
on Oct 9, 2026 ContributorAuthorMore actionsClaim: PM loop round 4 (takeover)
Session:session_0115N1oNnQS5WqofZ2DzaT3q
Account:os-sales(the seat's linked user asGET /useranswers it; the card's assignee from this act)
Branch:claude/issue-22014-multi-run-shard-timings(continued: its one PR, #22022, landed as2301e17eaf; the branch is deleted on the remote)
Worktree:objectstack-issue-22014
Domain:domain:devx
Seat:domain:devx#1
File surface: no new code. The open half is the first multi-run dataset onmain(record6024318057); this claim verifies it and closes the card.
Container & model:n/a — no dev dispatch; verification and closure by the seat
Clause-②: no
Responsibility:n/a — not a defect card
Thread-read: 6068864362
Serial constraints cleared:no open PR touches scripts/test-shard-timings.json, .github/workflows/shard-timings-refresh.yml, scripts/ci/select-shard-timings-run.mjs or scripts/measure-test-shard-timings.mjs— read at 2026-10-09T03:10Z
Release: claim6020801436(domain:devxseat 2,session_01VF48aw8RPG6wzDnMgp6rtw,os-justin) is taken over by this seat on the maintainer's order. Provenance: who — the maintainer; verbatim — 「22014 你负责接手,优先处理」; where — this seat's PM session chat (session_0115N1oNnQS5WqofZ2DzaT3q), received on 2026-10-09 before 03:09Z.② Assignee. Swapped in one write,
--unassign os-justin --assign os-sales, read back asos-sales(relay run 37877941374).④ Handover record.
- The old branch's last pushed head was
dcdf5eb6f4, the patch-round merge ofmaininto the ACCEPTed68a3123802. It landed through the queue as2301e17eaf(seat 2's record6024318057), and the branch is deleted on the remote. - Status at takeover: the code half is on
main. The open half (the first multi-run dataset) landed before this act: the maintainer'sworkflow_dispatchof "Shard Timings Refresh" (run 37856323158, 2026-10-08T22:54Z) opened PR chore(ci): refresh the Test Core shard-timings dataset #22368, whichos-zhuangmerged at 2026-10-08T23:52Z as040184752c. - So no dev is dispatched. The verification and closure follow in this seat's next comment.
Generated by Claude Code
- The old branch's last pushed head was
objectstack-fleet commented
on Oct 9, 2026 ContributorAuthorMore actionsLanded — closing
completed. The first multi-run shard-timings dataset is onmain· 2026-10-09T03:12Zdomain:devxseat 1 ·session_0115N1oNnQS5WqofZ2DzaT3q(takeover claim6073464580). The seat verified this onorigin/main11d119ab18by content, not from the PR bodies.Both halves are on
main:- The producer: PR fix(ci): accumulate shard-timings runs until every package has three executions, provisional below that #22022 →
2301e17eaf(seat 2's landing record6024318057). The refresh accumulates scheduled runs until every package hasMINIMUM_EXECUTED_RUNS = 3executions, takes the median, and names any package below that depth asprovisional. - The dataset: PR chore(ci): refresh the Test Core shard-timings dataset #22368 →
040184752c. The maintainer'sworkflow_dispatch(run 37856323158) opened it, andos-zhuangmerged it 2026-10-08T23:52Z. It changesscripts/test-shard-timings.jsononly, +298/−84.
Read in
scripts/test-shard-timings.jsononmain:provenance.runsholds 21 runs, where it held 1 (['37262126122']) when this card was filed.provenance.minimumRunsis3, andprovenance.carriedOveris[].- There are 72 package weights.
provisionalis["@objectstack/sdui-parser"]: 0 executed runs in the window, carried forward on a turbo cache hit.
Triage's direction (
6020193841), item by item:- At least three executed scheduled runs, taking the median: yes. The figures are above.
- The PR records each package's run-to-run spread: yes. PR chore(ci): refresh the Test Core shard-timings dataset #22368's body has a "Sample depth and run-to-run spread" table (min, median and max per package) with a PROVISIONAL warning for the one short package.
- ci: a per-package suite-duration ratchet — a PR that makes a suite exceed its measured ceiling is red; ceilings rise only by ruling (maintainer-directed, growth constraint) #16468's ceilings are generated only from this dataset: this is now possible. ci: a per-package suite-duration ratchet — a PR that makes a suite exceed its measured ceiling is red; ceilings rise only by ruling (maintainer-directed, growth constraint) #16468 unblocks in this act, and its generation refuses every package in
provisional(the contract seat 2 carried there). - ⛔ The 25% headroom is not raised and ci: Test Core compares each shard's measured duration with its predicted weight — warning at 1.3x, red at 70% of the timeout (maintainer-directed, drift alarm) #16465's 1.5× red is not loosened: both hold.
scripts/partition-test-shards.mjsis unchanged since9c3bec0f4d(ci(test-core): wire the shard timing-drift check: red past 1.5x measured/predicted, warning past 1.3x #21998), and PR chore(ci): refresh the Test Core shard-timings dataset #22368 touches only the dataset.
Effect on the p1 symptom (
6068864362).main's hourly full run on35afb15878(run 37836262554) failed onTest Core (6/6)'s drift step at 1.53×. The scheduled full runs since the refresh passed: 37869710053 onc8c803c293and 37873442424 one030d436bf. So did the push run 37873990782 on11d119ab18.In this act:
- This card closes
completed, andpm:dispatchedcomes off in the same write. - ci: a per-package suite-duration ratchet — a PR that makes a suite exceed its measured ceiling is red; ceilings rise only by ruling (maintainer-directed, growth constraint) #16468 and ci(test-shards): the shard balance is derived on full-run sums while PR and merge_group runs use the affected set — the CLI shard (1/6) measures 34–36 min against 10–20 for the others and sets CI and queue wall time #22075, both
pm:blockedon this card, are re-read and returned topm:queue, each with its own comment.
Generated by Claude Code
- The producer: PR fix(ci): accumulate shard-timings runs until every package has three executions, provisional below that #22022 →
- added a commit that references this issue
on Oct 9, 2026
Filed by PM seat
domain:devx#2(session_01VF48aw8RPG6wzDnMgp6rtw) while re-deriving #16468's blocker after #16465 landed (PR #21998 →9c3bec0f4d, record6019846983). Its measurements come from #21998's dev report (6018891729), and the seat re-read them onorigin/mainbefore filing. ⛔ Filed bare: grading and routing are triage's. ⛔ Not a claim.Filing gate: ① a defect, class (a). A named producer's output contradicts measurement.
scripts/test-shard-timings.json, written byscripts/measure-test-shard-timings.mjsthrough.github/workflows/shard-timings-refresh.yml, and last refreshed by chore(ci): refresh the Test Core shard-timings dataset #21826 (f2aa0c9fad).9c3bec0f4d, judges against it.Measured
One run. On
origin/main, the dataset'sprovenance.runsis['37262126122'], andcarriedOveris[]. Every weight is a single observation. The refresh workflow's own header says the regeneration step "accumulates runs, each under its own--run <id>group … and medians the per-run sums across runs". This dataset carries one.@objectstack/spec(test+test:reposummed on both sides, per #16550):ea7ff394b6, run36380128221)37262126122), the current weight374333817953745359838837467882762(70 executed, 0 replayed)The seat read
37453598388's Test Core timing table itself (aggregator job112248922952): spec 1573.20 against pinned 1134.86, and objectql 671.60 against 521.33 (1.29×).@objectstack/objectql: pinned at 521.33 s, it reads 1.29–1.38× on six of nine post-refresh runs (#21998's report).Run-to-run spread on one package:
@objectstack/cliread 1033.74 s and 1723.05 s an hour apart (#21758's acceptance reading6012200987). A single run is therefore a sample, not a baseline.Why it matters now
5924551316again, on a different package.Reader
RUN_COUNT: unbound variable—— 数据集已 22 天未动,而 pull_request 演练结构上跑不到这条腿 #18341, finding(ci): the refreshed shard-timings dataset records @objectstack/cli at 733s while whole-CLI runs measure ~1660s (2.27x), so Test Core shard 2 runs ~30 min and #16465 drift red cannot be wired #21758) sat indomain:devx.Blocked-by:written on it in the same act).Dedupe
MCP
search_issues, repo-scoped, sorted by update: 「shard timings dataset refresh rests on a single run, spec under-recorded, median across several runs」 → 13 hits. One is open, #21933 (a merge-queue runner fault, unrelated). The closed ones include #21758 (the cli instance of this shape, fixed by the #21826 refresh), #16473 (the median merge rule for sliced packages), #16550 (thetest:repofold), #18341 (the refresh's write-back crash) and #16464 (the scheduled refresh itself). None records a dataset resting on one run, or spec's under-weight.Dedupe words:
test-shard-timings one run provenance·spec 1134.86 1573 1651·ratchet ceiling 25% headroom standing redGenerated by Claude Code