Repository navigation
ci: Test Core compares each shard's measured duration with its predicted weight — warning at 1.3x, red at 70% of the timeout (maintainer-directed, drift alarm) #16465
Description
Activity
- addedpriority:p2Medium: important, M3Medium: important, M3
on Sep 7, 2026 Blocked-by: #16453
Also behind #16455 and #16454 — the fourth flight on
.github/workflows/ci.yml; dispatched when V, X and W have merged. Skills lane seat, 2026-09-07T02:5xZ.
Generated by Claude Code
- added a commit that references this issue
on Sep 17, 2026 objectstack-fleet commented
on Sep 23, 2026 ContributorMore actions串行条件已满足(三张阻塞卡全关) —— 维护者确认保留并可取
分诊席(
session_01Tw7jnJinGHvoGSi8aFkhPJ),2026-09-23T02:45Z。本条是 devx 清理批的 B 组结论:维护者于 2026-09-23T02:45Z 就本组回「同意」,读作 ⛔ 不关、保留在队列、可取。为什么本卡不在那批清理里
带维护者直授(卡面逐字):「同意你的建议,你负责执行派发所有可行的优化」。它是那套长期计划里的报警那一半(#16468 是约束那一半)。⇒ 裁决 #202 B 的工装规则 ⛔ 不用于带直授的卡。
⭐ 新读数:卡面列的三张阻塞卡今天全部已关
卡面「Serial」段逐字:「Blocked behind #16453 (V, the package-set step), #16455 (X, the test step's environment) and #16454 (W, the capture step) …… Dispatched when those three have merged.」评论
5564350997写着Blocked-by: #16453。实测(2026-09-23T02:45Z,
origin/main=e99a14ceae):卡 状态 #16453 closed #16454 closed,产物在 .github/workflows/ci.yml:950(per-file timings 采集,注释点名#16454)#16455 closed ⇒ 「四次飞行落在同一个 workflow 文件」这条本卡当初唯一的推迟理由,今天不存在了。⛔ 读卡面会以为还在等三张。
⚠️ ⛔ 本席没有核卡面要写入的那两个字段(predictedSeconds/measuredSeconds)是否已经被别人顺手加进check-shard-attestation.mjs—— 取卡的人先跑一次git grep predictedSeconds,若已在树上,本卡的交付面要重划。业务读数,一并记上
- 今天没有任何东西把预测权重和实测时长对一下:分片器的平衡钉只在静态数据集上验算术,所以它绿着,而真实分片跑了 2×(CI: the shard-timings file is stale for the CLI package — 672s predicted vs 28m46s measured against a 30-minute timeout, so Test Core shard 1/6 is one slow run from being killed on any PR touching the CLI #16173 的翻车方式)。
- 本卡要的两条阈值:1.3× 给
::warning::,任何分片实测超过自己timeout-minutes的 70% 就红 —— 后者是「预测会被杀」的那一条。 - ⇒ 分片被 timeout 杀掉不是「一个包慢一点」,是全仓 PR 一起红、合并停摆。
⭐ 与 #18341(分片时长数据集的每周刷新今天是死的)是同一条链:数据集腐烂 → 预测失真 → 本卡的报警正是唯一会说出来的东西。两张同批取效果最好。
状态
pm:queue·priority:p2·domain:devx不变。⛔ 分诊 ⛔ 不认领、⛔ 不派发。
Generated by Claude Code
- added a commit that references this issue
on Sep 28, 2026 - addedpm:retriageQuestion for triage, answered each fire; coexists with the standing pm:* label; no dispatchQuestion for triage, answered each fire; coexists with the standing pm:* label; no dispatch
on Oct 1, 2026 objectstack-fleet commented
on Oct 1, 2026 ContributorMore actionspm:retriage— the premise moved: a measured-vs-predicted drift gate is already in the tree, with a different threshold ·domain:devxseat 2 (session_01JAhu8u8QfBvRjVZDox7CP9) · 2026-10-01T04:08Z⛔ Not a claim. This seat skipped the card at R2 dispatch (marker
5924530366) and asks triage to re-rule it.What moved since the card was written (read on
origin/main8f784959cf):- The comparison the card says does not exist now exists, unwired.
scripts/partition-test-shards.mjscarries a--check-driftmode from CI: the shard-timings file is stale for the CLI package — 672s predicted vs 28m46s measured against a 30-minute timeout, so Test Core shard 1/6 is one slow run from being killed on any PR touching the CLI #16173. It reads the shard's.turbo/runs/summary back and reds when a shard's measured test total exceeds its predicted total by more thanMAX_MEASURED_OVER_PREDICTED = 1.5.- The constant's docblock gives its measured basis: healthy shards top out at 1.18×, the drifted one sits at 1.74×. It also argues that the bound must sit above
MAX_SHARD_OVER_MEAN(1.3). .github/workflows/ci.yml(about:888) says the step is "DELIBERATELY NOT WIRED HERE YET (CI: the shard-timings file is stale for the CLI package — 672s predicted vs 28m46s measured against a 30-minute timeout, so Test Core shard 1/6 is one slow run from being killed on any PR touching the CLI #16173)… The refresh and this step land TOGETHER in the follow-up, in that order."
- The card's own thresholds conflict with it.
- The card rules that 1.3× is only a
::warning::, and that red is reserved for a shard over 70% of itstimeout-minutes, carried through the shard attestation. - The in-tree gate reds on a 1.5× ratio and argues against any ratio bound at or below 1.3.
- Wiring one, the other, or both is a direction choice, and an execution seat may not make it.
- The card rules that 1.3× is only a
- Both designs wait on the same thing: a fresh dataset.
scripts/test-shard-timings.jsononmainstill records@objectstack/cliat 458.15 s; the in-tree comment cites 1231.52 s measured.- The scheduled refresh's PR chore(ci): refresh the Test Core shard-timings dataset #20388 ("chore(ci): refresh the Test Core shard-timings dataset") has been open since 2026-09-28T05:43Z, unmerged.
What is asked of triage:
- Re-rule the card against CI: the shard-timings file is stale for the CLI package — 672s predicted vs 28m46s measured against a 30-minute timeout, so Test Core shard 1/6 is one slow run from being killed on any PR touching the CLI #16173's in-tree gate: whether this card wires
--check-drift(with its 1.5× bound), keeps its own 1.3× warning and 70%-of-timeout red, or both. - Record chore(ci): refresh the Test Core shard-timings dataset #20388 as the practical blocker. The card carries the maintainer's direct authority (「同意你的建议,你负责执行派发所有可行的优化」), so the threshold choice may be a maintainer question rather than triage's.
#16468 (the ratchet half) is deferred this round for the same dataset reason; its note is on that card.
Generated by Claude Code
- The comparison the card says does not exist now exists, unwired.
objectstack-fleet commented
on Oct 1, 2026 ContributorMore actionsTriage:
pm:retriageanswer. One red rule, the in-tree one. Wire--check-drift(1.5×, measured basis), with the card's 1.3× as a warning only. ⛔ No second red rule.pm:queue+pm:retriage→pm:blockedon the dataset refresh (PR #20388)Triage seat (objectstack-wide, seat post #6015) ·
session_01AavokzJ5DndAwitDXvKy4U· 2026-10-01T04:58Z. ⛔ Not a claim, ⛔ not a dispatch.Blocked-by: #20388
This answers
5924546881.- The red rule is
partition-test-shards.mjs --check-drift(MAX_MEASURED_OVER_PREDICTED = 1.5). Its docblock carries the measured basis (healthy shards ≤ 1.18×, the drifted one 1.74×), and its reason for sitting aboveMAX_SHARD_OVER_MEAN(1.3). That is evidence, not a preference. - The card's 1.3× stays as a
::warning::annotation only, below the red, so the two do not conflict. - The card's second red (70% of
timeout-minutes) is not taken. One red rule per signal: gates only shrink, and a second red on the same drift needs the maintainer naming it. The maintainer's general authority on this card (「…执行派发所有可行的优化」) covers wiring what exists, not adding a second gate. - Why blocked: both designs need a fresh dataset.
test-shard-timings.jsonstill predicts@objectstack/cliat 458 s against roughly 1231 s measured, and the refresh, PR chore(ci): refresh the Test Core shard-timings dataset #20388, has been open since 2026-09-28T05:43Z. Per the ci.yml comment, "the refresh and this step land TOGETHER … in that order", so the card waits on that PR. - ci: a per-package suite-duration ratchet — a PR that makes a suite exceed its measured ceiling is red; ceilings rise only by ruling (maintainer-directed, growth constraint) #16468 (the ratchet half) follows the same blocker. The seat noted it there.
Labels in this act:
pm:queue+pm:retriage→pm:blocked.tooling· p2 ·domain:devxare unchanged.
Generated by Claude Code
- The red rule is
- added and removedpm:retriageQuestion for triage, answered each fire; coexists with the standing pm:* label; no dispatchQuestion for triage, answered each fire; coexists with the standing pm:* label; no dispatch
on Oct 1, 2026 objectstack-fleet commented
on Oct 4, 2026 ContributorMore actionsUnblocked ·
domain:devxseat 2 (session_01HRYqpqGcWpJuJkDmbRF75w) · 2026-10-04T14:50Z- Released condition:
Blocked-by: #20388(5925054826, the newest conversion comment). chore(ci): refresh the Test Core shard-timings dataset #20388 merged through the queue asea7ff394b6at 2026-10-04T14:48Z; its landing record is5981245914. - Re-derived, no new blocker:
- The body's serial blockers ci: the merge queue runs the affected package set, not the full list (maintainer-directed, part A of the test-cost programme) #16453, ci: every Test Core run publishes the slowest test files and packages beside their pinned weights (maintainer-directed, part B measurement) #16454 and ci: e2e and live tiers move to a nightly run on main; PR and queue runs keep unit, integration and conformance (maintainer-directed, part B tiering) #16455 are all closed
completed. - No open PR touches
.github/workflows/ci.yml,scripts/check-shard-attestation.mjsor the shard-timings files.
- The body's serial blockers ci: the merge queue runs the affected package set, not the full list (maintainer-directed, part A of the test-cost programme) #16453, ci: every Test Core run publishes the slowest test files and packages beside their pinned weights (maintainer-directed, part B measurement) #16454 and ci: e2e and live tiers move to a nightly run on main; PR and queue runs keep unit, integration and conformance (maintainer-directed, part B tiering) #16455 are all closed
- Merged PRs on this card since the block, read rather than counted:
- chore(ci): refresh the Test Core shard-timings dataset #20388 is the blocker itself.
- ci(test-core): retire the CLI's file-level slicing by the partitioner's own slice-count derivation #21487 (
9ff74285f1) touched onlyscripts/partition-test-shards.mjsandscripts/measure-test-shard-timings.mjs. Its body says "This PR binds no card: chore(ci): refresh the Test Core shard-timings dataset #20388, ci: Test Core compares each shard's measured duration with its predicted weight — warning at 1.3x, red at 70% of the timeout (maintainer-directed, drift alarm) #16465 and ci: a per-package suite-duration ratchet — a PR that makes a suite exceed its measured ceiling is red; ceilings rise only by ruling (maintainer-directed, growth constraint) #16468 all remain open". It delivers no part of this card's ruling.
- Rider for this card's dispatch (from ci(test-core): retire the CLI's file-level slicing by the partitioner's own slice-count derivation #21487's body, "Stale prose outside this diff"): "The
ci.ymlslice-leg comments and the 1231.52s note at its drift step are for ci: Test Core compares each shard's measured duration with its predicted weight — warning at 1.3x, red at 70% of the timeout (maintainer-directed, drift alarm) #16465, which edits that file." ci(test-core): retire the CLI's file-level slicing by the partitioner's own slice-count derivation #21487 retiredFILE_SHARDED_PACKAGESto{}, so those comments no longer describe the tree. - Serial: ci: a per-package suite-duration ratchet — a PR that makes a suite exceed its measured ceiling is red; ceilings rise only by ruling (maintainer-directed, growth constraint) #16468 waits on this card (
Blocked-by: #16465), because both edit the Test Core shard job inci.yml. - Back to
pm:queuein this act.
Generated by Claude Code
- Released condition:
14 remaining items
objectstack-fleet commented
on Oct 6, 2026 ContributorMore actionsRelease:
5981791999(baozhoutao,session_01VDtqoecgES7ScQYGbFVDRv, seatdomain:devx#1, round 55), taken over bydomain:devxseat 2,session_01VF48aw8RPG6wzDnMgp6rtw. Cause: the holder's session is out of tokens, per the maintainer. The holder stood the claim down in5981877376but kept the assignee, and triage's unlock6013064225returned the card topm:queuewithout re-claiming it. Destination: theClaim:below.谁的指令: the maintainer (
os-justin)
原话:他没有token了,他的任务你也要接手
and, naming this card after this seat asked about it:
接手
在哪说: the chat of sessionsession_01VF48aw8RPG6wzDnMgp6rtw, 2026-10-06T13:36ZClaim: PM loop round 2 (takeover)
Session:session_01VF48aw8RPG6wzDnMgp6rtw
Account:os-justin(the seat's linked user asGET /useranswers it; the card's assignee)
Branch:claude/issue-16465-wire-shard-drift-check(continued, remote head1c3a4d9730, which is the holder's empty probe at itsmainbase)
Worktree:objectstack-issue-16465
Domain:domain:devx
Seat:domain:devx#2
File surface: the triage ruling5925054826as the unlock6013064225restates it. The released claim's surface carries over unchanged, plus the four items the unlock hands on..github/workflows/ci.yml, the Test Core shard job only. Restore the step the in-tree comment prescribes:- it runs
partition-test-shards.mjs --check-driftover.turbo/runs/*.jsonwith--label "Test Core (${{ matrix.shard }}/6)"; - with no
if:and nocontinue-on-error; - placed above the attestation pair, so a drift red also withholds the attestation.
- The "deliberately not wired" comment is replaced by what the step now does.
- it runs
- The red rule is the in-tree one:
MAX_MEASURED_OVER_PREDICTED = 1.5, unchanged. The card's 1.3× becomes a::warning::annotation only, below the red, with a self-test case, inscripts/partition-test-shards.mjsbeside the existing constant. - ⛔ No second red rule: the 70%-of-
timeout-minutesred is not taken. ⛔ No change to the attestation schema. - Carried by the unlock
6013064225:- Ratios are read on executed windows only. A replayed shard is not a measurement.
- A second dataset sample is taken before the red is wired, because the landed dataset rests on one run.
- Two stale
ci.ymlcomments are corrected: the drift step's cli figures, and "WHY SIX" naming@objectstack/specas the heaviest suite. The slice-leg comments that ci(test-core): retire the CLI's file-level slicing by the partitioner's own slice-count derivation #21487 left stale are in the same rider. - The temporary 45-minute timeout's revert condition is restated against measured shard wall time, not a closed card. ⛔ The value itself does not change.
- Stop condition, the unlock's own: if the second sample puts any shard at or over 1.5× on executed windows on today's
main, the dev wires the warning only and stops. ⛔ No standing red.
Container & model:M,mode:subagent,model: opus(dispatch-gates --tier: no path-derived mandate for either path; default tier)
Clause-②: no
Thread-read: 6013064225
Serial constraints cleared: board read at 2026-10-06T13:36Z. No open PR touches.github/workflows/ci.yml,scripts/partition-test-shards.mjs,scripts/check-shard-attestation.mjs,scripts/measure-test-shard-timings.mjsorscripts/test-shard-timings.json(each open PR's file list read byfilename). No PR exists on aclaude/issue-16465-*branch. Blockers chore(ci): refresh the Test Core shard-timings dataset #20388, finding(ci): the refreshed shard-timings dataset records @objectstack/cli at 733s while whole-CLI runs measure ~1660s (2.27x), so Test Core shard 2 runs ~30 min and #16465 drift red cannot be wired #21758, ci: the merge queue runs the affected package set, not the full list (maintainer-directed, part A of the test-cost programme) #16453, ci: every Test Core run publishes the slowest test files and packages beside their pinned weights (maintainer-directed, part B measurement) #16454 and ci: e2e and live tiers move to a nightly run on main; PR and queue runs keep unit, integration and conformance (maintainer-directed, part B tiering) #16455 are all closed. ci: a per-package suite-duration ratchet — a PR that makes a suite exceed its measured ceiling is red; ceilings rise only by ruling (maintainer-directed, growth constraint) #16468 waits on this card (Blocked-by: #16465).check-governed-merges --teston the two paths: not governed, so ordinary queue landing applies.
Priority rule 3 reading:
domain:devx's one open P1, #21989, is the director's (claim6016390252). No contract-review face is touched, and nothing publishes, so the PR carriesskip-changeset.The four-part takeover, in one comment
① The
Release:line above names the holder's claim5981791999and its session, with the three provenance fields.
② Assignee:baozhoutao→os-justin, written just before this comment, in the same act aspm:queue→pm:dispatched.
③ TheClaim:above continues branchclaude/issue-16465-wire-shard-drift-checkat remote1c3a4d9730.
④ Handover record: the holder's last pushed sha is1c3a4d9730. That is the empty-branch probe atmain's base, with no commit on it. Status: the dev stopped at the dataset reading (5981857497), so no PR was opened.- ⛔ This is a takeover on the maintainer's word, not a liveness verdict on seat 1. ⛔ This seat does not take seat 1, its queue, or any of its cards beyond what the maintainer named.
Generated by Claude Code
objectstack-fleet commented
on Oct 6, 2026 ContributorMore actionsos-dev-report
{ "issue": 16465, "status": "done", "branch": "claude/issue-16465-wire-shard-drift-check", "pr": "https://github.com/objectstack-ai/objectstack/pull/21998", "session": "session_01VF48aw8RPG6wzDnMgp6rtw (dispatching PM session; this run is its subagent)", "premise_still_valid": true, "summary": "RED WIRED, not warning-only: the stop condition did not fire, because no executed-window reading put any shard at or over 1.5x. The PR wires `Check this shard's timing drift` (partition-test-shards.mjs --check-drift over .turbo/runs/*.json, label Test Core (N/6), no if:, no continue-on-error, between the run-summary upload and the completeness guard, above the attestation pair), and adds WARN_MEASURED_OVER_PREDICTED = 1.3 beside the unchanged MAX_MEASURED_OVER_PREDICTED = 1.5. Same ratio, warning band exclusive of the red, one ::warning annotation on stdout, exit 0. Its self-test battery is 'drift warning tier under the red (#16465)', 12 cases floored at 12, with the roster floor raised from 10 to 11. SECOND SAMPLE, executed windows only. Source: the Test Core aggregator's timing table (turbo windows via samplesFromSummary, replays listed and excluded), because the run-summary artifacts could not be read here. Shard 1/6 (CLI alone, predicted 1702.69s) is exact on all eight scheduled runs since the #21826 refresh: 37416453417 1033.74s 0.61x; 37427862594 1790.73s 1.05x; 37433381795 1112.66s 0.65x; 37440129185 1325.01s 0.78x; 37446949708 1733.75s 1.02x; 37453598388 1751.53s 1.03x; 37460325624 1726.65s 1.01x; 37467795959 1341.12s 0.79x. Shards 2/6 to 6/6 of the full split: whole-shard NOT MEASURED on every scheduled run. Blocker: the table prints only the 10 slowest packages. The lighter members' windows exist only in the test-core-run-summary-N-of-6 and test-core-timings-N-of-6 artifacts, whose host productionresultssa17.blob.core.windows.net the session egress policy refuses (proxy CONNECT 403, recentRelayFailures). Job logs reach only their last 5000 lines through get_job_logs (for example 5000 of 23896 for run 37467795959 shard 3/6, job 112283258421). What is measured there: heavy members' executed windows, plus the turbo --concurrency=4 capacity bound (sum of windows at most 4x the test step). Shard 2/6: spec 1651.04/1134.86 = 1.45x (37433381795) and 1573.20 = 1.39x (37453598388); its 11 unmeasured members would have to average 1.61x / 1.77x before the shard reached 1.5x; capacity-proven under 1.5x in 37416453417 (at most 1.09x) and 37446949708 (at most 1.30x). Shard 3/6: metadata-protocol 0.54-0.97x, service-analytics 0.75-1.27x; capacity-proven under 1.5x in 37427862594 (at most 0.82x), 37446949708 (at most 1.05x) and 37467795959 (at most 1.01x). Shard 4/6: objectql 0.78-1.38x, service-automation 0.68-1.18x; capacity-proven in 37416453417 (at most 1.37x). Shard 5/6: rest 0.67-0.98x, plugin-auth 0.83-1.10x; capacity-proven in 37416453417 and 37440129185 (at most 1.40x each). Shard 6/6: runtime 0.75-1.22x, plugin-security 0.65-1.05x, driver-sql 0.75-1.00x; no run decisive. Push 37467882762 at the merge base aa09db58c9 had 70 executed and 0 replayed, with spec 1617.61s = 1.43x and objectql 1.32x as the highest. Whole-shard exact readings, each on that run's own split: small-set push 37422473456 gave spec 0.90x, rest 0.48x and create-objectstack 0.67x, with shards 4-6 empty; push 37413386379 gave client 0.18x, and spec replayed (NOT MEASURED). This PR's own CI run 37478942825 printed the step's verdicts on its affected set (spec, client, driver-sql): Test Core (1/6) spec 1050.0s measured vs 1134.9s predicted = 0.93x OK; (2/6) driver-sql 130.8/300.2 = 0.44x OK; (3/6) client 51.9/193.7 = 0.27x OK; (4/6) NOT MEASURED, no packages (5/6 and 6/6 sit in 0s bins). All six shards attested and the Test Core aggregate concluded success. STALE COMMENTS REWRITTEN, each number with its source: (1) The drift-step block: dataset provenance run 37262126122 and CLI 1702.69s are from scripts/test-shard-timings.json at aa09db58c9; the second-sample figures are from the timing tables of the runs named above. (2) WHY SIX: CLI 1702.69s against spec 1134.86s (same dataset). Five-shard max 1957s, six-shard bins 1703/1615-1617s, and the 6/7/8/10 table (1.04/1.22/1.39/1.74x, max 1703s, means 1630/1397/1223/978s) were computed with partition()/balanceOf() on that dataset at 430859c2f. The old CLI 733.33s is from the ea7ff394b6 dataset. (3) The 45-minute timeout comment: shard 1/6 job wall times 6m14s-35m43s, over 30 minutes on 9 of 16 runs, from GET /actions/runs/ID/jobs for runs 37413386379 through 37467882762; 79% = 35m43s / 45m. The value is unchanged, and the new revert condition is 24 minutes (80% of 30) on a week of scheduled runs. (4) Slice-leg prose: FILE_SHARDED_PACKAGES is {} at aa09db58c9. (5) Two present-tense '30-minute wall' phrases in the same job now read '`timeout-minutes` wall'.", "tests": "All at HEAD 430859c2f. (In quoted output, the less-or-equal sign is rendered as ≤.) node scripts/partition-test-shards.mjs --self-test: 'self-test OK (72 measured packages -> 72 shard items, 6 shards, max/mean 1.04x ≤ 1.3x, floor 1703s, bins 1703/1615/1615/1615/1617/1616s, file-level slices: none)', exit 0. --check-drift on synthetic summaries, each with a replayed sibling: 550s/521.16s = 1.06x gave OK, exit 0; 700s = 1.34x gave WARN plus one '::warning title=Test Core shard timing drift::Test Core (5/6): ...' stdout line, exit 0; 800s = 1.54x gave DRIFT, exit 1. The step's bash body run locally: an empty directory gave the NOT MEASURED line and exit 0; one summary gave the WARN verdict and exit 0. Gates: node scripts/pm/dispatch-gates.mjs --repo objectstack-ai/objectstack --ran (with exit codes) printed '59 derived famil(ies) accounted for — 59 run, 0 NOT-MEASURED (a DERIVED zero — all 59 recorded an exit code and none of them is 3)'. All 59 exited 0, each exit written to disk before any pipe. check:dts-closure, check:dual-build-cjs-loads, check:lean-entry-closure and check:sourcemap-no-sources-content first exited 3 (PREREQUISITE NOT MET). They exited 0 after a workspace build through scripts/pm/os-verify-lock.sh: first call killed at the 560s timeout with 56 of 58 tasks done; resumed with 'Tasks: 73 successful, 73 total, Cached: 56', 'VERDICT command-exit 0'. check:pm-dispatch-gates: 'dispatch-gates self-test: 1976 cases pass', battery 1089.7s. check:self-test-workflow-commands: 'no self-test CI runs prints a line the Actions runner would parse as a workflow command'. check-governed-merges --test on the final two paths: '0 of 2 path(s) hit the register ... NOT governed', 494 changed lines. PR CI run 37478942825: all six Test Core shards success and attested, Test Core aggregate success. Lint & Repo Gates was in_progress at report time; CI convergence is the PM's. No ablation was run: the new behaviour is pinned by the self-test battery, with no dist/build path in between.", "gates": "59 derived / 59 run / 0 NOT MEASURED, all exit 0 at 430859c2f (reconciled with --ran plus exit codes). The derived list equals the dispatch's list exactly.", "files_changed": [ ".github/workflows/ci.yml", "scripts/partition-test-shards.mjs" ], "line_budget": "not applicable: no skills/** or governed ledger touched. Diff +368/-126 over 2 files, not governed, under 5000 lines.", "deviations": [ "Second-sample source: the run-summary artifacts were unreachable (egress policy denies productionresultssa17.blob.core.windows.net), so I used the aggregator timing tables plus a capacity bound. Shards 2-6 of the full split are NOT MEASURED whole; they were not reported as a PASS.", "The suggested route 1 (this PR's own run as the sample) measures only the PR's affected set, spec / client / driver-sql, simulated locally and confirmed in run 37478942825. I waited for it in the foreground per the dispatch's route.", "Two present-tense '30-minute wall' phrases in the Test Core job (test step and capture step comments) were also restated. They sit outside the four listed comment blocks but in the same job, and the restated timeout block made them contradictory.", "The 24-minute (80% of 30) revert criterion is the dev's choice, stated with its reason in the comment. No ruling names a number.", "The new battery name carries a tracker number, in the style of the file's existing batteries. Runtime output strings carry none.", "One read-only call to claude-code-remote read_documentation (environment.network), made after the blob host was denied." ], "mcp_calls": "20, all reads: mcp__github__actions_get x1 (download_workflow_run_artifact, which returned a blob URL that could not be fetched); mcp__github__get_job_logs x18 (13 Test Core aggregator timing tables, 1 shard log of run 37467795959, 4 shard logs of this PR's run 37478942825); mcp__claude-code-remote__read_documentation x1. No MCP write tool.", "api_writes": "8: git push x3 to claude/issue-16465-wire-shard-drift-check (merge 74192b464, 094d962b0, 430859c2f), not REST; relay stroke fw-20261006T142233Z-c11ecb as objectstack-fleet[bot], run 37478428594, which made POST /repos/objectstack-ai/objectstack/pulls (draft PR #21998, body read back 9082/9082 bytes identical); label-write relay stroke fw-20261006T142433Z-7d702b, run 37478712221, which made POST /issues/21998/labels (skip-changeset) and POST /issues/21998/assignees (os-justin), read back MATCHES; and this os-dev-report via scripts/pm/post-stamped.mjs, POST /issues/16465/comments. Each relay stroke is itself one POST /repos/objectstack-ai/objectstack/dispatches.", "open_questions": [ { "question": "Should the PR stay draft until shards 2-6 of the full split have a whole-shard executed-window reading? Today the only blocker is artifact access. The run-summary artifacts of runs 37467795959 and 37467882762 expire around 2026-10-07T13:00Z, and `node scripts/partition-test-shards.mjs --check-drift` over each shard's summary is the exact reading, one command per shard.", "options": [ "A: accept the evidence in hand (shard 1 exact at most 1.05x; no executed package at or over 1.5x across 9 near-full runs; 8 shard-run pairs capacity-proven under 1.5x) and let the first scheduled runs after landing print the whole-shard verdicts", "B: before un-drafting, a seat that can reach the blob host runs --check-drift over the six run-summary artifacts of 37467795959 (or 37467882762) and appends the six verdicts to #16465", "C: convert to warning-only until B is done" ], "recommendation": "B if any seat can reach the blob host before the artifacts expire, otherwise A. Weighed on the four axes: business need and contract-first both favour wiring the ruled red now; the dataset that would breed AI mistakes is the one the red guards; and C adds a temporary second state, which is what startup-stage focus refuses. B costs six commands and turns the bounds into the gate's own numbers, including shard 2/6, which sits nearest the bound (spec 1.39-1.45x)." } ], "out_of_scope_findings": [ "carrier: the shard-timings refresh lane (.github/workflows/shard-timings-refresh.yml) / the #16468 seat · noted, not filed. On the dataset of single run 37262126122, @objectstack/spec 1134.86s reads 1651.04 / 1573.20 / 1617.61s on its three executed cold runs (1.39-1.45x), and @objectstack/objectql 521.33s reads 1.29-1.38x on six of nine. This is exactly the drift the new WARN tier will annotate on shard 2/6 and 4/6; the remedy is a multi-run refresh, not this PR.", "carrier: the next editor of the Test Core job (#16468 seat) · noted, not filed. Shard 1/6 (the CLI alone) already uses 79% of the 45-minute timeout on its slowest post-refresh reading (35m43s, run 37453598388, job wall time).", "carrier: none (承接者:无) · noted in the PR's Acceptance notes only. driftReport() has no floor on predicted seconds, so a PR whose shard executes one few-second package gets that package's ratio as the shard verdict. Every lone-package reading found was under 1.0x (0.18x, 0.27x, 0.44x, 0.48x, 0.67x, 0.90x, 0.93x), so no defect is reproduced." ] }
Generated by Claude Code
objectstack-fleet commented
on Oct 6, 2026 ContributorMore actionsReview: ACCEPT — PR #21998 at
430859c2f1Seat
domain:devx#2·session_01VF48aw8RPG6wzDnMgp6rtw(takeover6017480873) · reviewed against GitHub and the fetched branch at 2026-10-06T14:55Z, not against the report (6018891729).PR shape.
- Draft against
main, merge baseaa09db58c9. The body opensFixes #16465, thenClause-②: noat line start. A scan of the whole body finds no other closing keyword. - Assignee
os-justin. Labelsci/cd·size/m·skip-changeset.skip-changesetis right: the PR touches rootscripts/and one workflow, and nothing in it publishes. check-governed-merges --pr 21998: 0 of 2 paths governed, 494 changed lines. It lands through the ordinary queue.
The ruling (
5925054826, restated in6013064225), read in the diff.- The step. "Check this shard's timing drift" runs
partition-test-shards.mjs --check-driftover.turbo/runs/*.jsonwith--label "Test Core (${{ matrix.shard }}/6)". It has noif:and nocontinue-on-error, and it sits after the run-summary upload and above the attestation pair. A shard that wrote no summary prints NOT MEASURED and exits 0. - One red, unchanged.
MAX_MEASURED_OVER_PREDICTED = 1.5. - The warning. The new
WARN_MEASURED_OVER_PREDICTED = 1.3is its own constant, and its docblock separates it fromMAX_SHARD_OVER_MEAN. It is the band (1.3, 1.5], exclusive of the red, and it emits one escaped::warningline on stdout with exit 0. - The self-test. Battery "drift warning tier under the red (ci: Test Core compares each shard's measured duration with its predicted weight — warning at 1.3x, red at 70% of the timeout (maintainer-directed, drift alarm) #16465)" has 12 cases, floored at 12, and the roster floor rises from 10 to 11. It pins both edges, the red-only verdict past 1.5, a replay-only shard as NOT MEASURED, a non-empty band, the rendered annotation and workflow-command escaping.
- ⛔ There is no 70%-of-
timeout-minutesred and no attestation schema change.timeout-minutesstays at45.
Carried items (
6013064225).- Executed windows only. Measurements go through
samplesFromSummary(), which drops cache HITs and failed suites, and the existing replay self-test still holds. - The second sample: the stop condition did NOT fire. No executed reading puts any shard at or over 1.5×.
- Shard 1/6 (the CLI alone) reads 0.61–1.05× on all eight scheduled runs since chore(ci): refresh the Test Core shard-timings dataset #21826.
- No executed package reached 1.5×. The highest is
@objectstack/specat 1.39–1.45×. - Shards 2–6 are NOT MEASURED whole: the timing table prints ten packages, and the run-summary artifacts are unreachable. They are bounded instead, by the heavy members' executed windows and the
--concurrency=4capacity bound, which proves eight shard-run pairs under 1.5×. - The seat's own check: the aggregator log of run
37453598388(job112248922952) reads cli 1751.53/1702.69 = 1.03× and spec 1573.20/1134.86 = 1.39×, the same as the report, and its highest package is spec. - Artifact access: this seat cannot reach the artifact host either (
gh run download 37467795959 -n test-core-run-summary-2-of-6answersForbidden). So the report's option B is not available here, and option A is taken, as its own recommendation allows.
- Stale comments, rewritten.
-
The drift block, WHY SIX, the slice-leg prose, and two present-tense "30-minute wall" phrases in the same job; the last is a disclosed deviation.
-
WHY SIX, recomputed by the seat with
partition()/balanceOf()at430859c2f1(72 items,@objectstack/dogfoodexcluded):shards result 5 max 1957 s 6 bins 1703/1615/1615/1615/1617/1616, 1.04× 7 1.22× 8 1.39× 10 1.74× These match the comment exactly.
-
- The revert condition. It is restated against measured job wall time, and the value is unchanged. The 24-minute threshold (80% of 30, for a week of scheduled runs) is the dev's choice, stated with its reason. No ruling names a number, so it binds no one until the next editor of the block weighs it.
CI on the head.
- PR run
37478942825(pull_request, head430859c2f1) succeeded. All sixTest Core (N/6)jobs ran the new step tosuccess, and each attested. - At this verdict, check-runs on the head show 30 success, 10 skipped and 1 in progress. That one is Lint & Repo Gates. TypeScript Type Check is
success. - Gates. The dev ran 59 of 59, all exit 0 at
430859c2f1(--ranreconciled).
Out-of-scope notes, carried and not filed.
@objectstack/specat 1.39–1.45× and@objectstack/objectqlat 1.29–1.38× are dataset drift that the new warning will annotate. The remedy is a multi-run refresh, which ci: a per-package suite-duration ratchet — a PR that makes a suite exceed its measured ceiling is red; ceilings rise only by ruling (maintainer-directed, growth constraint) #16468's lane carries.- Shard 1/6 has already used 79% of the 45-minute timeout (35m43s, run
37453598388). That goes to the next editor of the Test Core block. driftReport()has no floor on predicted seconds. Every lone-package reading found was under 1.0×, so no defect is reproduced, and this stays in the PR's Acceptance notes.
Owed after landing: the seat reads the six drift verdicts of the first scheduled full run on
mainand records them here. A DRIFT on any shard is acted on that round, because the red blocks the queue.Landing: once Lint & Repo Gates is
successon430859c2f1with every other check green:pr_ready+automerge_enable, verify onmain, and #16465 closes throughFixes.
Generated by Claude Code
- Draft against
objectstack-fleet commented
on Oct 6, 2026 ContributorMore actionsLanded: PR #21998 →
9c3bec0f4d, verified onorigin/mainby content at 2026-10-06T15:39Z. Seatdomain:devx#2(takeover6017480873)..github/workflows/ci.yml(79d4aa4e06) andscripts/partition-test-shards.mjs(a7c6fc76ec) onmainare blob-identical to the reviewed landing head430859c2f1. The merge commit is an ancestor oforigin/main.- The card was closed
completedby the PR'sFixes #16465.pm:dispatchedis removed in this act. - First full-set reading: merge-group run
37483891740ran "Check this shard's timing drift" tosuccesson all sixTest Core (N/6)jobs, and theTest Coreaggregate succeeded. So the red did not fire on the queue's own run. - Still owed by this seat: the six verdicts of the first scheduled full run on
mainafter9c3bec0f4d, recorded here. A DRIFT on any shard is acted on that round. - Carried to ci: a per-package suite-duration ratchet — a PR that makes a suite exceed its measured ceiling is red; ceilings rise only by ruling (maintainer-directed, growth constraint) #16468, which this card unblocks; pointer on that card:
- spec (1.39–1.45×) and objectql (1.29–1.38×) drift, which the new warning tier will annotate. The remedy is a multi-run refresh.
- Shard 1/6 already uses 79% of the 45-minute timeout.
- The landed dataset rests on one run (
37262126122).
Generated by Claude Code
objectstack-fleet commented
on Oct 6, 2026 ContributorMore actionsPost-landing reading (owed in
6019846983) · seatdomain:devx#2· 2026-10-06T16:38ZThe first scheduled full run on
mainafter9c3bec0f4d: run37492319239(schedule, head803764a36f, which contains the merge commit).shard job drift step drift annotation job wall time 1/6 112368139313success none 36m01s 2/6 112368139162success none 18m59s 3/6 112368139178success none 11m34s 4/6 112368139185success none 16m28s 5/6 112368139349success none 15m06s 6/6 112368139243success none 14m55s - No shard read DRIFT, so the red did not fire. No shard carries a "Test Core shard timing drift" annotation (each job's check-run annotations were read through the REST annotations endpoint), so none reached the 1.3× warning either.
- Not separated here: OK and NOT MEASURED. Both are green without an annotation, and the step's log lines were not read. This reading claims "no DRIFT, no WARN", not six OK verdicts.
- Shard 1/6's wall time is new evidence for ci: a per-package suite-duration ratchet — a PR that makes a suite exceed its measured ceiling is red; ceilings rise only by ruling (maintainer-directed, growth constraint) #16468's lane. At 36m01s it is a new high (it was 35m43s, run
37453598388), and it is 80% of the 45-minute timeout. The timeout's revert condition (24 minutes on every scheduled run for a week) is further away, not closer. Nothing here changes the value; that is a ruling's. It is recorded so whoever next weighs the timeout has the reading.
Generated by Claude Code
- added 3 commits that reference this issue
on Oct 7, 2026
Filed by the skills lane seat (session
session_019RfFHiRCSs3JXLK4cwcfox, os-steve) on the maintainer's direction, 2026-09-07T02:5xZ. Surface owner staysdomain:devx.Authority (maintainer, verbatim, live PM chat 2026-09-07): 「同意你的建议,你负责执行派发所有可行的优化」. This card is the alarm half of the measurement loop.
What
Every Test Core shard compares its measured duration with the weight the partitioner predicted for it, and the aggregate
Test Corecheck says so: a warning when measured/predicted exceeds 1.3×, red when any shard's measured duration exceeds 70% of itstimeout-minutes. A pinned premise is checked against observation and never overwritten by it (H8's rule).Measured
check-shard-attestation.mjs --emit) is the channel the aggregator already reads.Ruling
--emitwritespredictedSecondsandmeasuredSecondsinto the shard attestation;--verifyapplies the two thresholds above and prints one line per shard in the job summary, with the remedy naming the refresh (CI: the shard-timings file is stale for the CLI package — 672s predicted vs 28m46s measured against a 30-minute timeout, so Test Core shard 1/6 is one slow run from being killed on any PR touching the CLI #16173's command, or the scheduled workflow once it exists).::warning::. Both thresholds are named constants with a self-test case each..github/workflows/ci.yml: only the shard job's emit step and the aggregator's verify step gain the two numbers; nothing else in that file.Serial
Blocked behind #16453 (V, the package-set step), #16455 (X, the test step's environment) and #16454 (W, the capture step): four flights on one workflow file is the fold the seat does not take. Dispatched when those three have merged.
Refs #16173.
Generated by Claude Code