Repository navigation
ci(test-shards): the shard balance is derived on full-run sums while PR and merge_group runs use the affected set — the CLI shard (1/6) measures 34–36 min against 10–20 for the others and sets CI and queue wall time #22075
Description
Activity
objectstack-fleet commented
on Oct 7, 2026 ContributorAuthorMore actionsPath: fleet decision — CI and merge-queue wall time set by the slowest required job | 缺项 | none
Triage: first grade,
tooling·priority:p2·domain:devx·area:devpath·pm:blockedbehind #22014, as the body sequences itBlocked-by: #22014
Triage seat (objectstack-wide, seat post #6015) ·
session_01AavokzJ5DndAwitDXvKy4U· 2026-10-07T13:12Z. ⛔ Not a claim, ⛔ not a dispatch.Triage: lands in
scripts/partition-test-shards.mjs(the derivation andMAX_SHARD_OVER_MEAN = 1.3,:129) andscripts/test-shard-timings.json⇒domain:devx; rationale: the lane of #16173, #16454, #16464 and #22014.- Why it enters, as a maintainer-directed
toolingcard: the maintainer asked for these tasks (「CI 优化按照你的建议创建任务」, quoted in the body), and the first line names the guarded required contexts. Sotriage-duties.md:34is met.:64(close a p2/p3toolingcard while the lane queue holds an open P1, here release(v18): enter Changesets pre mode on main (changeset pre enter next) with onemajormarker, so the first v18 prerelease is 18.0.0-next.0 — the opening ruled B on #22050 #22080) does not apply to a maintainer-directed card, as on objectui#11762. - Why p2:
Test Core (1/6)sets CI and merge-queue wall time on every PR and every queue entry. It was measured at 34–36 minutes against 10–20 for the other shards. - Why blocked: finding(ci): the shard-timings dataset rests on ONE scheduled run, and records @objectstack/spec at 1134.86 s against 1573–1651 s executed — #16468's 25%-headroom ceilings built on it would red every PR that runs spec #22014 (dispatched) refreshes the shard-timings dataset this re-derivation reads. Claim after it lands, on the measured dataset.
- Scope, as the body states it: derive the balance on the runs that set the wall (the affected set on
pull_requestandmerge_group), not only on full-run sums. The shard bound stays a ratio, and its pins are kept. The PR records the before-and-after wall on one PR run and one queue run.
- Why it enters, as a maintainer-directed
- addedarea:devpathThe road — create, dev, verify, publish/install, connect an agent, iterateThe road — create, dev, verify, publish/install, connect an agent, iteratepriority:p2Medium: important, M3Medium: important, M3and removed
on Oct 7, 2026 objectstack-fleet commented
on Oct 9, 2026 ContributorAuthorMore actionsUnblocked →
pm:queue·domain:devxseat 1 · 2026-10-09T03:15Zos-sales·session_0115N1oNnQS5WqofZ2DzaT3q. ⛔ Not a claim.- Condition met:
Blocked-by: #22014(triage6038684734). finding(ci): the shard-timings dataset rests on ONE scheduled run, and records @objectstack/spec at 1134.86 s against 1573–1651 s executed — #16468's 25%-headroom ceilings built on it would red every PR that runs spec #22014 closedcompleted, with landing record6073490354: the multi-run shard-timings dataset is onmain(PR chore(ci): refresh the Test Core shard-timings dataset #22368 →040184752c; 21 runs,minimumRuns: 3, with oneprovisionalpackage,@objectstack/sdui-parser). - Double check. The only merged PR referenced on this card since that transition is PR fix(spec)!: defineSeed refuses a record key the target object does not have #22294 (spec:
defineSeedaccepts a misspelled record key at compile time andos validatepasses it, although its JSDoc promises "typos in record field names are caught at compile time" #22149,defineSeedrecord keys). It touches neitherscripts/partition-test-shards.mjsnor the dataset, so it delivers no part of this card. - Re-read on
main11d119ab18:MAX_SHARD_OVER_MEAN = 1.3is still atscripts/partition-test-shards.mjs:129, and the file is unchanged since9c3bec0f4d.- The refreshed weights put
@objectstack/cliat a median of 1738.88 s, with a spread of 1062.12–1844.58 s (1.74×) over 15 runs. That comes from PR chore(ci): refresh the Test Core shard-timings dataset #22368's spread table.
- Re-price at unlock: unaffected in shape. The card's question, a balance derived on full-run sums while PR and
merge_groupruns execute the affected set, does not depend on the dataset's depth. Whoever takes it re-measuresTest Core (1/6)'s wall time on PR and queue runs against the refreshed weights before deriving anything.
Labels in this act:
pm:blocked→pm:queue.tooling·priority:p2·domain:devx·area:devpathare unchanged.
Generated by Claude Code
- Condition met:
objectstack-fleet commented
on Oct 9, 2026 ContributorAuthorMore actionsClaim: PM loop round 4
Session:session_0115N1oNnQS5WqofZ2DzaT3q
Account:os-sales(the seat's linked user asGET /useranswers it; the card's assignee)
Branch:claude/issue-22075-affected-set-shard-balance
Worktree:objectstack-issue-22075
Domain:domain:devx
Seat:domain:devx#1
File surface:scripts/partition-test-shards.mjs(the derivation, the slicing refusal and their self-test pins); theTest Corejob in.github/workflows/ci.yml(its partition wiring, and the shardtimeout-minutesre-sized from the measured distribution); and ascripts/ci/*helper only if the per-run route needs one.scripts/test-shard-timings.jsonis read-only. ⛔ No reduction of what the shards test. ⛔ The seven required contexts are unchanged. ⛔ No dataset work. (Stop on breach; explain in the report.)
Container & model:M/L,mode:subagent,model: opus — the default tier; dispatch-gates --tier: "no path-derived mandate"
Clause-②: no
Responsibility:scripts/partition-test-shards.mjs derives the balance on full-run sums (this lane's code) | no platform path covers it: #16445's 45-minute wall only absorbs it | every PR and every merge-queue build reaches it today (Test Core (1/6) at 34–36 min against 10–20 for the others)
Thread-read: 6073525371
Serial constraints cleared:No open PR touches .github/workflows/ci.yml, scripts/partition-test-shards.mjs or scripts/test-shard-timings.json (17 open PRs' file lists read). #22085 (seat 2, p1) edits release.yml, which is disjoint; #22385 (this seat) is the console pin bump, also disjoint. Fold or serial with #16468, which shares the Test Core job in ci.yml: SERIAL, not folded, because gate ① fails (a balance derivation and a duration ratchet are different defect shapes). #16468 waits behind this card, and its known pit is recorded at 6073538550.— read at 2026-10-09T03:20Z
Generated by Claude Code
objectstack-fleet commented
on Oct 9, 2026 ContributorAuthorMore actionsos-dev-report
{
"issue": 22075,
"status": "done",
"branch": "claude/issue-22075-affected-set-shard-balance",
"pr": "#22415",
"session": "session_0115N1oNnQS5WqofZ2DzaT3q — subagent of the devx seat 1 PM, same session id; identity is the branch named by Claim 6073588381",
"premise_still_valid": true,
"premise_note": "The symptom holds (12 of 14 sampled runs: the shard holding the whole CLI ran 2.3-7.6x the other five). The mechanism the card names does not: the per-run partition (route (a)) already exists, because every Test Core job partitions its own turbo ls --affected list at run time. The defect is the cost model: bins were graded on summed package weight, while a shard runs its whole packages 4-wide (--concurrency=4) and each slice in a leg of its own after them. Bin sums read 1.00-1.47x on the runs that fill six bins, while job walls read 2.3-3.3x on the same runs.",
"summary": "Graded the Test Core split on predicted shard WALL instead of bin sum. The wall model is TEST_CONCURRENCY=4, pinned against ci.yml; the whole-package leg is max(heaviest serial task, sum/4), slice legs are added after it, and test+test:repo packages count half their weight as serial. partition() places a slice at weight x 4. On the committed dataset the CLI whole reads 2.52x, at 2 slices 1.31x and at 3 slices 1.01x, so FILE_SHARDED_PACKAGES={'@objectstack/cli':3} is derived by pin 3c minimality. PREVIOUS_FILE_SHARDED_PACKAGES is now the outgoing {}. A new planShards() slices only when the run's own split needs it and the slices spread, and prints the decision on the shard log. All 12 sampled CLI runs slice at 3; the nightly tier run (2 shards) stays whole. The Test Core timeout is re-sized 45 -> 35 with the 14-run window and numbers in the ci.yml comment, and the stale 'slice step idle' comments are updated. The measured slowest/mean-of-others pin cannot be read on this PR's own runs: its affected set is spec, client and driver-sql, with no CLI. The seat's post-landing read recipe is in the PR body.",
"route": "(b) static configured slice count with an affected-set-aware per-run decision, graded on a shard-wall model. Chosen after measuring 6 pull_request runs (37876969409, 37874898502, 37873731877, 37873634681, 37872770181, 37871447533) and 8 merge_group runs (37875522518, 37875521531, 37873846077, 37873791941, 37873694430, 37872756554, 37871575925, 37870616843), all after 0401847, through the jobs API. I also reproduced each run's affected set locally and split it with the partitioner. Route (a) as written is already the status quo and moves nothing. Route (b) as the card words it ('runs whole on full runs') was rejected, because on walls the full list needs the slices too (2.52x whole). Simulation on the 12 CLI runs: with slices as plain sum items, the estimated slowest/mean-of-others is 1.37-1.48x. With slot-weighted slices (the choice here) it is 1.09-1.39x. Today's estimate is 2.2-4.5x.",
"assumptions": {
"1": "CONFIRMED on fdfdd7e and on the merged head 9156fd4 (pre-change). MAX_SHARD_OVER_MEAN = 1.3 sat at partition-test-shards.mjs:129. Newest commit touching the file: 9c3bec0 (git log -- the file; the newest commit is inside the shallow window, so the reading needs no deepening). ci.yml timeout-minutes: 45 at :503 is the Test Core job (job id 'test', name 'Test Core (N/6)'); :2662 is the console-pin job. The PR changes :503 to 35 and leaves console-pin alone.",
"2": "CONFIRMED. The scripts/test-shard-timings.json provenance has 21 runs, minimumRuns 3 and provisional ['@objectstack/sdui-parser']. The CLI weight is 1738.88 s. All 14 sampled runs post-date 0401847: their drift lines predict the CLI at 1738.9 s. The CLI re-measured whole at 1121-1932 s (test-step windows) on those runs; run 37875522518 read 1867.26 s, which its timing table reports as 1.07x of the pinned weight.",
"3": "CONFIRMED. The CLI was in the affected set of 12 of the 14 sampled runs (86%), reproduced by running turbo ls --affected between each run's base and head plus the cross-package union. The 2 without it were docs-only diffs (spec, rest and create-objectstack via the union). Whenever present, the CLI ran whole on one shard (shard 1/6, CLI alone, 20.5-39.1 min job wall).",
"4": "DISPROVEN AS WORDED. Route (a)'s per-run assignment already exists: ci.yml 'Compute this shard's package set' runs partition-test-shards.mjs on the run's own turbo-ls.json, which is the affected set on pull_request and merge_group. The slice reassembly (measure-test-shard-timings.mjs sliceOfEnvironment, which decodes against FILE_SHARDED_PACKAGES and PREVIOUS) and the OS_TEST_SHARD wiring (turbo.json cli#test env, packages/cli/vitest.config.ts) are reused unchanged by route (b) instead. The generator self-test passes on the new live maps {cli:3} and {}.",
"5": "YES, with no rename. The 6-wide matrix already takes a run-time assignment, since each job computes the split itself. The job names Test Core (N/6) and the aggregate 'Test Core' are untouched, the seven required contexts are unchanged, and pnpm check:required-contexts exited 0."
},
"tests": "Head 9156fd4 (branch merged with origin/main 83e7ae9 once). node scripts/partition-test-shards.mjs --self-test exit 0: 'self-test OK (72 measured packages -> 74 shard items, 6 shards, wall max/mean 1.01x ≤ 1.3x at concurrency 4, floor 580s, walls 667/667/660/660/660/660s, bins 904/902/903/2641/2641/2641s of test windows, file-level slices: @objectstack/cli x3)'. The new battery 'shard walls and the per-run slice decision (#22075)' has 13 cases, and the roster floor went 11 -> 12. Consumer self-tests exit 0: measure-test-shard-timings, check-test-completeness, report-test-timings. Ablation: fix committed first; both through node scripts/ablation-replace.mjs in WRAP mode, with on-disk anchor counts 1 -> 0 and blob changes recorded. (1) The wall return line replaced by 'whole + sliced' (blob 739b9274f6d1 -> 22c7fd4d7ccd): self-test RED 'slice spread: the committed dataset, split as CI splits it, carries no slice -- @objectstack/cli: whole (whole fits: 1739s is within 2303s)'. (2) 'const needed = own > target;' replaced by 'const needed = false;' (blob -> 7d1f33503974): RED 'balance: at 6 shards the slowest predicted shard wall is 2.52x the mean (1739s vs 689s)'. Both restored by git checkout HEAD -- abs path, with blob == HEAD 739b9274f6d1 and git diff HEAD empty, and the self-test is green after. There was no build/dist step (plain node script, no exports resolution). Lint, a declared narrowing: eslint --no-inline-config --format json scripts/partition-test-shards.mjs reported 1 file, 0 errors, 0 warnings. The population is read from eslint's own config (isPathIgnored false; ci.yml is not an eslint input), and the config enables no type-aware linting (computed parserOptions.project null), so untouched files' verdicts cannot move. The full pnpm lint is CI's. No package touched, so no build closure (step 1) and no package test/typecheck (step 2) are owed. CI on PR 22415 was in_progress at report time (14 success, 5 skipped, 15 in_progress; Governed Surface Queue Guard success).",
"gates": {
"head": "9156fd40af",
"derived": 58,
"run": 58,
"not_measured": 4,
"unrun": 0,
"reconciliation": "node scripts/pm/dispatch-gates.mjs --repo objectstack-ai/objectstack --ran FILE (with ':: exit N' per line): '58 derived, 54 run, 4 NOT-MEASURED, 0 UNRUN'",
"exits": {
"node scripts/check-aggregator-roster.mjs": 0,
"node scripts/check-aggregator-roster.mjs --self-test": 0,
"node scripts/check-ci-filter-parity.mjs": 0,
"node scripts/check-ci-filter-parity.mjs --self-test": 0,
"node scripts/check-closing-keyword-parity.mjs": 0,
"node scripts/check-closing-keyword-parity.mjs --self-test": 0,
"node scripts/check-comment-mask-corpus.mjs": 0,
"node scripts/check-declaration-mirrors.mjs": 0,
"node scripts/check-declaration-mirrors.mjs --self-test": 0,
"node scripts/check-dts-emitted.mjs --self-test": 0,
"node scripts/check-position-name-fold-loaders.mjs": 0,
"node scripts/check-position-name-fold-loaders.mjs --self-test": 0,
"node scripts/check-scripts-symbol-anchors.mjs": 0,
"node scripts/check-scripts-symbol-anchors.mjs --self-test": 0,
"node scripts/check-self-test-wired.mjs": 0,
"node scripts/check-self-test-wired.mjs --self-test": 0,
"node scripts/check-self-test-workflow-commands.mjs": 0,
"node scripts/check-self-test-workflow-commands.mjs --self-test": 0,
"node scripts/check-step-collectors.mjs": 0,
"node scripts/check-step-collectors.mjs --self-test": 0,
"node scripts/check-whole-set-label-write.mjs": 0,
"node scripts/check-whole-set-label-write.mjs --self-test": 0,
"node scripts/ci/scheduled-full-run.mjs --self-test": 0,
"node scripts/docs-audit/check-drift-comment.mjs": 0,
"node scripts/partition-test-shards.mjs --self-test": 0,
"node scripts/pm/bare-root-worklist.mjs --self-test": 0,
"node scripts/pm/ci-failure.mjs --self-test": 0,
"pnpm check:agent-test-spelling": 0,
"pnpm check:bash32-floor": 0,
"pnpm check:cli-command-ids": 0,
"pnpm check:console-injection": 0,
"pnpm check:console-sha": 0,
"pnpm check:cross-package-test-inputs": 0,
"pnpm check:declared-population-live": 0,
"pnpm check:driver-memory-census": 0,
"pnpm check:dts-closure": 3,
"pnpm check:dual-build-cjs-loads": 3,
"pnpm check:entry-guard": 0,
"pnpm check:gitlink-declared": 0,
"pnpm check:lean-entry-closure": 3,
"pnpm check:node-version": 0,
"pnpm check:nul-bytes": 0,
"pnpm check:parse-guard": 0,
"pnpm check:pm-expected-skips": 0,
"pnpm check:pm-post-stamped": 0,
"pnpm check:pnpm-acquisition": 0,
"pnpm check:pnpm-filter-targets": 0,
"pnpm check:ratchet-remedy-authority": 0,
"pnpm check:refd-timer-probe": 0,
"pnpm check:required-contexts": 0,
"pnpm check:shard-attestation": 0,
"pnpm check:sourcemap-no-sources-content": 3,
"pnpm check:stall-guard-budget": 0,
"pnpm check:stall-guard-headroom": 0,
"pnpm check:watch-hint-literal": 0,
"pnpm check:workflow-status-functions": 0,
"pnpm check:workflow-step-name-quoting": 0,
"node scripts/measure-test-shard-timings.mjs --self-test": 0,
"node scripts/check-test-completeness.mjs --self-test": 0,
"node scripts/report-test-timings.mjs --self-test": 0,
"pnpm check:pm-dispatch-gates": 1
},
"not_measured_reason": "check:dts-closure, check:dual-build-cjs-loads, check:lean-entry-closure and check:sourcemap-no-sources-content exited 3 (PREREQUISITE NOT MET): each loads every package's built dist, and this diff touches no package source or build config.",
"red": "pnpm check:pm-dispatch-gates exited 1 on 1 of 2011 cases: 'no mkdtempSync site in this tree takes a base the scan cannot read — UNRESOLVED: packages/qa/dogfood/test/security-catalog-cold-boot-environment-holder.dogfood.test.ts:108 (process.cwd())'. That file came from e030d43, which is already on main before the branch base fdfdd7e. This diff does not touch it; the case is pre-existing (see out_of_scope_findings)."
},
"line_budget": "not applicable: no skills/** or ledger-governed file touched. Diff: .github/workflows/ci.yml +89/-73, scripts/partition-test-shards.mjs +475/-112 (749 changed lines, under the 5000 human-merge threshold).",
"files_changed": [
".github/workflows/ci.yml",
"scripts/partition-test-shards.mjs"
],
"deviations": [
"Premise half-false (rule 6): implemented route (b) on a shard-wall cost model rather than on sums, and did not keep 'runs whole on full runs', because measurement shows full runs need the slices too.",
"The ≤ 1.3 measured ratio cannot be proven on this PR's own runs: the affected set is spec, client and driver-sql, with no CLI. The PR body proposes the seat's post-landing read: >= 3 pull_request and >= 2 merge_group runs whose shard log prints 'slicing: @objectstack/cli: sliced x3', ratio taken from GET actions/runs/RUN_ID/jobs.",
"The Test Core timeout is re-sized 45 -> 35 from measured components (14-run window) plus the model's post-change prediction. The post-change distribution itself cannot be measured before landing. The ci.yml comment says so, and gives a revert condition and a re-read trigger at 28 min.",
"3 MCP read calls (mcp__github__get_job_logs), because REST job-log download answered 403 at the blob-storage CONNECT (productionresultssa9.blob.core.windows.net, connect_rejected).",
"Commit trailers use AGENTS.md's model-free pair (Claude-Session + Co-authored-by: Claude) instead of the harness reminder's model-named trailer. The PR footer uses AGENTS.md's session-URL form.",
"Merged origin/main once (9156fd4, default merge message) before opening the PR, as dispatched."
],
"mcp_calls": "3 — mcp__github__get_job_logs x3 (jobs 113643185252, 113651419988, 113643185284; read-only). No MCP write tool used.",
"api_writes": "4 — all through the fleet relay (POST /repos/objectstack-ai/objectstack/dispatches, executed as objectstack-fleet[bot]): (1) pr_create, i.e. POST /pulls, opening #22415 as draft, body read back identical (17329 bytes); (2) label-write --assign os-sales, i.e. POST /issues/22415/assignees; (3) label-write --add skip-changeset, i.e. POST /issues/22415/labels (read back: ci/cd, size/l, skip-changeset); (4) post-stamped --comment=22075, i.e. POST /issues/22075/comments (this report). git push is not counted as a REST write.",
"open_questions": [],
"out_of_scope_findings": [
"class: a · reach: named producer — e030d43's dogfood test packages/qa/dogfood/test/security-catalog-cold-boot-environment-holder.dogfood.test.ts:108 (mkdtempSync(join(process.cwd(), ...))) reds pnpm check:pm-dispatch-gates (the Lint & Repo Gates step 'PM dispatch-gates self-test' whenever its family is selected) on 1 of 2011 cases: 'no mkdtempSync site in this tree takes a base the scan cannot read — UNRESOLVED ... (a base this scan cannot read: process.cwd())'. Measured on 9156fd4, which holds origin/main 83e7ae9; the file is identical to main. · dedupe words: mkdtempSync process.cwd, dispatch-gates self-test, security-catalog-cold-boot-environment-holder, 'a base the scan cannot read'",
"carrier: 承接者:无 · noted, not filed — .github/workflows/test-nightly-tiers.yml header says FILE_SHARDED_PACKAGES 'cuts the CLI into two vitest slices'. That has been stale since the map emptied; after this PR the CLI is configured at 3 and cannot spread on that workflow's 2 shards, so it runs whole there, as it does today. Comment only; in the PR's Acceptance notes."
]
}objectstack-fleet commented
on Oct 9, 2026 ContributorAuthorMore actionsREWORK, patch round 1: PR #22415 (head
9156fd40af) · 2026-10-09T04:28ZReviewed by
domain:devxseat 1 ·session_0115N1oNnQS5WqofZ2DzaT3q, against GitHub andorigin/main, not the report.What holds, checked by the seat:
- Draft, base
main, second lineClause-②: no. 2 files, +564/−185, both inside the claim's file surface (6073588381). - Not governed (
check-governed-merges --pr 22415: 0 of 2 paths, 749 lines).skip-changesetis right: rootscripts/and workflows publish nothing. node scripts/partition-test-shards.mjs --self-test, re-run by the seat at9156fd40af: exit 0, with the report's verdict line (wall max/mean 1.01x <= 1.3x at concurrency 4 … file-level slices: @objectstack/cli x3).- Jobs API spot check, merge_group 37875522518:
Test Core (1/6)–(6/6)read 33.3 / 11.1 / 10.0 / 8.2 / 10.8 / 10.0 min, andRun this shard's testsread 1880 / 548 / 478 / 375 / 556 / 481 s. Both match the PR's table. MAX_SHARD_OVER_MEAN = 1.3(:134) andWARN_MEASURED_OVER_PREDICTED = 1.3(:201) are unchanged. The diff adds or removes no drift constant, and no job or matrix name changes.- The premise correction is accepted. Route (a)'s per-run partition was already the status quo, and the defect is the cost model: bin sums read 1.00–1.47x while job walls read 2.3–3.3x on the same runs. The PR body records this with run ids.
Two items for the patch round:
- Change
Fixes #22075toPart of #22075. The Done-when pin reads: "the measured slowest-shard / mean-of-the-others ratio on affected-set runs is ≤ 1.3, read from the jobs API and quoted in the PR body with run ids". The PR says this cannot be read on its own runs, because its affected set carries no CLI. So the merge must not close the card. The card closes on the seat's post-landing read. - Keep
timeout-minutesat 45 in this PR, and move the re-size to the post-landing half.- 35 rests on a prediction, not on a measured post-change distribution. The PR itself measured the wall model reading low on whole-package legs: on 37872770181 it predicted 573/442/442/442/491 s, and the shards took 905/737/704/668/388 s.
- The ~30-minute worst case applies the 2.0x packing measured on shards of ~1,400 s of windows to shards of ~2,640 s. No run has executed a shard that size yet.
- If the prediction misses, the wall kills
merge_groupruns for every lane. Holding 45 costs nothing the stall guard does not already cover (it caps at 20 minutes and is the primary hang detector). - After landing, the seat reads the post-change job walls on the same runs as the ratio. A one-line follow-up under this card then re-sizes the wall from that distribution. The
ci.ymlcomment may keep the 14-run window and its numbers as the input to that re-size, but it must not say a re-size happened.
The patch round goes to the same dev under the same claim. The PR title and its "The Test Core wall" section follow item 2, and the post-landing read recipe adds the job-wall distribution. No
mainmerge in this round:main'sLint & Repo Gatesis red from #22400, which PR #22416 fixes, and the seat mergesmainin withupdate-branchonce that fix lands.
Generated by Claude Code
- Draft, base
20 remaining items
objectstack-fleet commented
on Oct 9, 2026 ContributorAuthorMore actionsos-dev-report
{
"issue": 22075,
"round": "round 3 (claim 6077897041, brief 6077883558)",
"status": "done",
"branch": "claude/issue-22075-density-capped-slices",
"pr": "#22456",
"head": "2ccec0b333",
"session": "session_0115N1oNnQS5WqofZ2DzaT3q (mode:subagent of the domain:devx seat 1 PM; identity is the branch named by Claim 6077897041, verified as the newest Claim: on the card before the first edit)",
"premise_still_valid": true,
"summary": "Draft PR #22456 (Part of #22075, Clause-②: no) slices @objectstack/cli 3 ways at PLAIN weight under a density cap, keeping round 2's per-run planShards() decision and the slice-spread check and dropping the x4 slot weight, TEST_CONCURRENCY, shardWalls() and the wall-graded pins. The cap is densityCap(): the densest bin of the whole-package split of the committed dataset's full list, computed by the unchanged partition() on every call (1772.65s, the same float the 806b03e partitioner gives); planShards() slices only when NEEDED (the package is past 1.3x the heaviest other serial task in the run), SPREADS and WITHIN CAP, and refineToCap() repairs LPT overshoot by moving whole packages only (never a slice; an unsliced split stays byte-identical). The count 3 is derived on the serial floor in pin 3c: the smallest n whose slice is within 1.3x spec's 573.35s half-task (n=2: 869s refused, n=3: 580s). On the full list the bins read 1772.63/1771.02/1772.63/1770.95/1771.09/1772.46s, within the cap. The component model predicts the 1.3 wall pin is MISSED at 3 slices (1.28-1.52x on the 12 CLI runs) while the slowest job falls 26-57%; 5 seeded slices are predicted at 1.10-1.31x under the same cap, left to the seat as an open question because that count has no packing-free derivation.",
"assumptions": {
"1": "CONFIRMED FOR BINS, measured with the code itself on scripts/test-shard-timings.json (unchanged since 0401847), full list, 6 shards. This round: 1772.63/1771.02/1772.63/1770.95/1771.09/1772.46s. Slices 1/3, 2/3, 3/3 on shards 2, 3, 4, each 579.63s plus 1191.3-1193.0s of whole packages; spec (1146.69s) on shard 1. Cap 1772.65s = the 806b03e partitioner's max bin on the same dataset (its bins 1771.39/1772.65/1772.18/1771.55/1772.41/1770.59; its --self-test prints 1771/1773/1772/1772/1772/1771). Round 2 (c64130b --self-test): 904/902/903/2641/2641/2641s. The stop condition as dispatched (no split under the cap meets 1.3 on the full list) is NOT met: by the component model 4 seeded slices read 1.25x/1.28x and 5 seeded read 1.11x/1.13x on the two full-list runs, all under the cap. But 3 slices read 1.33x/1.37x there, so the assumption's wall half does not hold at 3.",
"2": "HOLDS AT THE MEDIANS, ~24.5 min. Components from the jobs API: setup 39/65/104s (84 jobs, 14 pre-change runs); closure on the two full-list runs 113/170/338s; whole-package leg at ~1191s of predicted windows 367/488/753s (10 pre-change shards with bins 1171-1193s: 37874898502 and 37873694430 shards 2-6); slice closure step 0/13/14s and heaviest-slice shard test step 490/722/731s (round 2's 8 sliced runs 37892033675, 37893672824, 37894048074, 37894050587, 37894053453, 37895967479, 37894129260, 37895965974); post 9/14/85s. Median sum 65+170+13+488+722+14 = 1472s = 24.5 min against 39.1 min today; upper envelope 33.8 min (each component's worst, never jointly observed). Slice skew cross-checked against vitest 4.1.11's sha1 split of the CLI's 372-file list with the 17 slowest files at measured seconds: 0.81/1.17/1.02 of an even third, matching round 2's measured slice-shard medians 546/722/640s.",
"3": "HOLDS FOR WALL, NOT FOR THE PIN. Round 2's 14 reproduced affected sets re-split with this code (before bins equal round 2's recorded bins). All 12 CLI runs slice x3; 2 docs-only runs (37871447533, 37872756554) unchanged. Predicted slowest job falls 26-57%; predicted slowest/mean-of-others 1.28-1.52x (11 of 12 above 1.3). Every after-split densest bin is within the cap and at or under the run's own before densest bin. Per run (before measured ratio and slowest job -> after predicted): 37876969409 3.84x 26.3 -> 1.28x 12.3; 37874898502 3.03x 34.3 -> 1.45x 21.4; 37873731877 5.65x 32.6 -> 1.52x 14.6; 37873634681 2.67x 33.8 -> 1.41x 22.4; 37872770181 2.26x 39.1 -> 1.33x 26.6; 37875522518 3.32x 33.3 -> 1.50x 20.4; 37875521531 2.58x 32.5 -> 1.41x 21.1; 37873846077 2.52x 26.8 -> 1.36x 18.2; 37873791941 2.61x 33.9 -> 1.41x 21.6; 37873694430 2.81x 29.9 -> 1.41x 19.1; 37871575925 2.48x 20.5 -> 1.39x 15.3; 37870616843 2.43x 35.6 -> 1.37x 23.8. Model: run medians for setup/closure/post, whole legs = predicted windows x the run's own measured packing (0.34-0.58), floored at the heaviest serial task, slice legs = run's CLI-alone test step / 3 x vitest share, +14s slice closure."
},
"tests": "Head 2ccec0b (one origin/main merge, 440bed6). node scripts/partition-test-shards.mjs --self-test exit 0: 'self-test OK (72 measured packages -> 74 shard items, 6 shards, max/mean 1.00x [ASCII less-or-equal] 1.3x, floor 1147s, bins 1772.63/1771.02/1772.63/1770.95/1771.09/1772.46s within the 1772.65s density cap, file-level slices: @objectstack/cli x3)'. New battery 'density cap and the per-run slice decision (#22075)' 14 cases; roster floor 11 -> 12; balancing battery 27 cases (floor 25). Consumer self-tests exit 0: measure-test-shard-timings, check-test-completeness, report-test-timings. Ablations at 2ccec0b, fix committed first, through scripts/ablation-replace.mjs with a git checkout HEAD trap: (1) cap enforcement removed (withinCap set without the cap comparison; blob 52581cb6eafe -> b3a6ca0dece3, anchor 1 -> 0): self-test RED 'density cap: a slicing whose densest shard is past the cap was taken: big: sliced x3 (... densest shard 180.00s ...)'; (2) repair removed (placeItems returns bins; blob -> fefdda9ddd8c): RED 'density repair: a split LPT put at 70s was left at 70s, past the 69s cap'. Both restored: blob == HEAD 52581cb6eafe, git diff HEAD empty, self-test green after. No build or dist step (plain node script, no exports resolution). Lint, a declared narrowing: eslint --no-inline-config --format json scripts/partition-test-shards.mjs = 1 file, 0 errors, 0 warnings; isPathIgnored false for the script and true for ci.yml; computed parserOptions.project and projectService null, so untouched files' verdicts cannot move; full pnpm lint is CI's. check-commit-card-trailers --range origin/main..HEAD exit 0 (3 commits). ci.yml diff is comment-only (0 non-comment changed lines). Control-byte scan of both files: no hits. PR CI at report time, read once: 34 check runs, 13 success, 5 skipped, 16 in_progress.",
"gates": {
"head": "2ccec0b333",
"derived": 58,
"run": 58,
"not_measured": 4,
"unrun": 0,
"reconciliation": "node scripts/pm/dispatch-gates.mjs --ran ran.list --repo objectstack-ai/objectstack at 2ccec0b: exit 0, '58 derived famil(ies) accounted for -- 54 run, 4 NOT-MEASURED (4 DERIVED from a recorded exit 3)'",
"exits": {
"node scripts/check-aggregator-roster.mjs": 0,
"node scripts/check-aggregator-roster.mjs --self-test": 0,
"node scripts/check-ci-filter-parity.mjs": 0,
"node scripts/check-ci-filter-parity.mjs --self-test": 0,
"node scripts/check-closing-keyword-parity.mjs": 0,
"node scripts/check-closing-keyword-parity.mjs --self-test": 0,
"node scripts/check-comment-mask-corpus.mjs": 0,
"node scripts/check-declaration-mirrors.mjs": 0,
"node scripts/check-declaration-mirrors.mjs --self-test": 0,
"node scripts/check-dts-emitted.mjs --self-test": 0,
"node scripts/check-position-name-fold-loaders.mjs": 0,
"node scripts/check-position-name-fold-loaders.mjs --self-test": 0,
"node scripts/check-scripts-symbol-anchors.mjs": 0,
"node scripts/check-scripts-symbol-anchors.mjs --self-test": 0,
"node scripts/check-self-test-wired.mjs": 0,
"node scripts/check-self-test-wired.mjs --self-test": 0,
"node scripts/check-self-test-workflow-commands.mjs": 0,
"node scripts/check-self-test-workflow-commands.mjs --self-test": 0,
"node scripts/check-step-collectors.mjs": 0,
"node scripts/check-step-collectors.mjs --self-test": 0,
"node scripts/check-whole-set-label-write.mjs": 0,
"node scripts/check-whole-set-label-write.mjs --self-test": 0,
"node scripts/ci/scheduled-full-run.mjs --self-test": 0,
"node scripts/docs-audit/check-drift-comment.mjs": 0,
"node scripts/partition-test-shards.mjs --self-test": 0,
"node scripts/pm/bare-root-worklist.mjs --self-test": 0,
"node scripts/pm/ci-failure.mjs --self-test": 0,
"pnpm check:agent-test-spelling": 0,
"pnpm check:bash32-floor": 0,
"pnpm check:cli-command-ids": 0,
"pnpm check:console-injection": 0,
"pnpm check:console-sha": 0,
"pnpm check:cross-package-test-inputs": 0,
"pnpm check:declared-population-live": 0,
"pnpm check:driver-memory-census": 0,
"pnpm check:dts-closure": 3,
"pnpm check:dual-build-cjs-loads": 3,
"pnpm check:entry-guard": 0,
"pnpm check:gitlink-declared": 0,
"pnpm check:lean-entry-closure": 3,
"pnpm check:node-version": 0,
"pnpm check:nul-bytes": 0,
"pnpm check:parse-guard": 0,
"pnpm check:pm-expected-skips": 0,
"pnpm check:pm-post-stamped": 0,
"pnpm check:pnpm-acquisition": 0,
"pnpm check:pnpm-filter-targets": 0,
"pnpm check:ratchet-remedy-authority": 0,
"pnpm check:refd-timer-probe": 0,
"pnpm check:required-contexts": 0,
"pnpm check:shard-attestation": 0,
"pnpm check:sourcemap-no-sources-content": 3,
"pnpm check:stall-guard-budget": 0,
"pnpm check:stall-guard-headroom": 0,
"pnpm check:watch-hint-literal": 0,
"pnpm check:workflow-status-functions": 0,
"pnpm check:workflow-step-name-quoting": 0,
"pnpm check:pm-dispatch-gates": 0
},
"not_measured_reason": "check:dts-closure, check:dual-build-cjs-loads, check:lean-entry-closure, check:sourcemap-no-sources-content exited 3, PREREQUISITE NOT MET: they load every package's built dist/, and this diff touches no package.",
"long_battery": "pnpm check:pm-dispatch-gates ran under nohup with its exit code written to a file (exit 0, '2011 cases pass', 1080.2s), awaited in the foreground with tail --pid on the recorded PID."
},
"line_budget": "not applicable: no skills/** or ledger-governed file. Diff vs origin/main: .github/workflows/ci.yml +59/-45 (comments only), scripts/partition-test-shards.mjs; total 2 files, +634/-175 = 809 changed lines, under the 5000 threshold; 0 governed paths.",
"files_changed": [
".github/workflows/ci.yml",
"scripts/partition-test-shards.mjs"
],
"deviations": [
"Shipped the brief's 3-slice design although the component model predicts it misses the 1.3 pin (1.28-1.52x). The dispatched stop condition (no split under the cap meets the pin on the full list) is not met, because 4 or 5 seeded slices are predicted to meet it under the cap; the switch was not made because that count has no packing-free derivation. Raised as open question 1.",
"The slice count is derived on the serial floor (pin 3c: smallest n whose slice is within 1.3x the heaviest other serial task), not on bin sums, which the brief named for grading. Pins 2 and 3 do grade bin sums again. Bin sums cannot derive any count: slicing never moves the mean, and on bin sums n = 1 meets the bound.",
"Added refineToCap() (move or swap whole packages out of the densest bin) beside the cap. It is not in the brief. Without it, LPT overshoots the cap by about 0.5s on roughly half of slice counts (2- and 4-way cuts on today's dataset), so a refresh would stop full lists slicing at random. It never moves a slice, and leaves unsliced splits byte-identical.",
"My first reads and install log went into the scratchpad's shared issue-22075/ directory, which earlier rounds of this card also used. comments.json and issue.json there may have overwritten round 2's copies of the same names. Everything after that went into issue-22075/r3/.",
"The command that started the background battery put its '&' on the whole && chain. The battery ran correctly, but the PID file was not written. I read the PID (6358) from ps without signalling it, and waited on it with tail --pid.",
"The second gate batch's per-command logs overwrote the first batch's numbered 01-27 logs. All 58 exit codes are intact in ran.list, which the reconciliation read.",
"Commit trailers use AGENTS.md's model-free pair (Claude-Session + Co-authored-by: Claude) instead of the harness's model-named trailer. The PR footer uses AGENTS.md's session-URL form instead of the harness's two-line block.",
"Merged origin/main once (440bed6, default merge message) before opening the PR, as dispatched. The timeout stays 45 and the drift constants, matrix and contexts are unchanged."
],
"mcp_calls": "1 -- mcp__github__get_job_logs on job 113704731036 (round-2 merge_group 37893672824's aggregate Test Core job, for the CLI slice windows and per-file seconds); read-only. No MCP write tool.",
"api_writes": "4 -- all through the fleet relay (POST /repos/objectstack-ai/objectstack/dispatches, executed as objectstack-fleet[bot]): (1) pr_create = POST /repos/objectstack-ai/objectstack/pulls, draft #22456, body 23075 bytes sent = stored, read back identical; (2) label-write --assign os-sales = POST /repos//issues/22456/assignees; (3) label-write --add skip-changeset = POST /repos//issues/22456/labels (read back: ci/cd, size/l, skip-changeset; assignee os-sales); (4) post-stamped --comment=22075 = POST /repos//issues/22075/comments (this report). 4 git pushes (empty-branch probe, 2 commits, the main merge) are not REST writes.",
"open_questions": [
{
"question": "The component model predicts the card's 1.3 pin is missed at 3 slices (1.28-1.52x on 12 CLI runs, 1.33x/1.37x on the full lists). It predicts 5 slices seeded one per shard at 1.10-1.31x under the same density cap, and 4 seeded at 1.23-1.46x. Which count lands, and how is it derived?",
"options": [
"A: Land 3 as in PR #22456. Read the pin and the slowest-job distribution on real runs. Then patch the count with the in-situ slice and whole legs those runs measure.",
"B: Patch #22456 to 5 seeded slices before landing. Pin 3c would derive the count from a wall model with a measured packing constant (0.34-0.58 s of test step per window second on the 14 runs).",
"C: Patch #22456 to 5 seeded slices with a structural rule, n = SHARD_COUNT - 1 (one slice per shard except the one the next serial floor takes)."
],
"recommendation": "A, by the four axes. (1) Real need: A cuts the slowest job 26-57% now. The pin is read after landing whatever lands, and B's and C's 1.10-1.31x are model readings with the same error bars as A's 1.28-1.52x. (2) Long-term soundness: A's count is derived packing-free. B adds a measured constant no gate re-measures, the shape round 2's TEST_CONCURRENCY had. C is a rule, not a derivation. (3) AI-error: no axis difference beyond A and C having fewer moving parts. (4) Startup focus: A adds no constant and no placement mode now, and the next decision is taken on measured legs. The cost of A is likely one more round on this card."
}
],
"out_of_scope_findings": [
"carrier: 承接者:无 · noted, not filed -- .github/workflows/test-nightly-tiers.yml header says FILE_SHARDED_PACKAGES 'cuts the CLI into two vitest slices'; stale since the map emptied and still wrong at 3. That run's 2 shards cannot spread 3 slices, so planShards() keeps the CLI whole there. Comment only, outside the file surface; in the PR's Acceptance notes."
]
}objectstack-fleet commented
on Oct 9, 2026 ContributorAuthorMore actionsACCEPT — PR #22456 (head
2ccec0b333, round 3) · 2026-10-09T10:26ZReviewed by
domain:devxseat 1 ·session_0115N1oNnQS5WqofZ2DzaT3q, against GitHub andorigin/main, not the report.Checklist:
- Draft, base
main. The first line isPart of #22075and the second isClause-②: no. A full body scan finds no closing keyword beside any card number. - 2 files, +634/−175, inside the round-3 claim (
6077897041). Not governed (check-governed-merges --pr 22456: 0 of 2 paths).skip-changesetis correct. ci.ymlis comment-only, read by the seat: the diff carries 0 changed lines that are not comments or blank.timeout-minutes: 45stands.MAX_SHARD_OVER_MEAN = 1.3andWARN_MEASURED_OVER_PREDICTED = 1.3are unchanged. The drift step is untouched.- The density cap, re-run by the seat at
2ccec0b333:--self-testexits 0 with "bins 1772.63/1771.02/1772.63/1770.95/1771.09/1772.46s within the 1772.65s density cap, file-level slices: @objectstack/cli x3". The cap comes from the unchangedpartition()on the committed dataset (densityCap()), not from a typed constant. The dev's ablation of the cap's enforcement turns the self-test red. So on the full list, no shard carries more predicted windows than the pre-ci(test-shards): grade the Test Core split on predicted shard wall and slice the CLI per run #22415 split, the density proven green (6077857603). Per the report, on the 12 sampled CLI runs every densest bin is at or below that run's own pre-change densest bin.
CI, read by the seat at
2ccec0b333:- All 37 check runs are complete: 32
successand 5 skipped (check-expected-skips --pr 22456: OK). Lint & Repo Gates(113767993404, carrying the dispatch-gates self-test),TypeScript Type Check(113769386237) andTest Core(113771263642) aresuccess.
Dev readings, accepted on their stated commands:
- 58 derived gates: 54 at exit 0 (
check:pm-dispatch-gatesamong them, its exit code recorded), and 4 NOT MEASURED (exit 3, dist-loading; no package touched). mcp_calls: 1 read.
Deviations, accepted with reasons:
refineToCap()keeps LPT from overshooting the cap by about 0.5 s on some slice counts. It moves whole packages only and leaves unsliced splits byte-identical.- The slice count is derived on the serial floor (pin 3c). Bin sums cannot derive a count, because slicing never moves the mean.
- The scratch-file overlap with earlier rounds touched no repo file.
Open question 1, answered by the seat: A. Land 3 slices, read the pin on real runs, then size the count from measured legs.
- The count of 3 is derived without a packing constant, and the slowest job is predicted to fall 26–57%.
- Options B and C reach their 1.10–1.31× through a measured packing constant that no gate re-measures. That is the same shape as round 2's
TEST_CONCURRENCY, which is what failed. - The pin is read after landing whatever lands (
Part of).
After the merge, this seat:
- Reads the drift step on the first
merge_groupand scheduled full runs. A single drift red attributable to this split means an immediate revert, as in round 2. - Reads the ≤ 1.3 pin and the slowest-job distribution on 3
pull_requestand 2merge_groupruns whose shard log printssliced x3. - Decides the count from those measured legs.
Landing: ready + auto-merge now.
Generated by Claude Code
- Draft, base
objectstack-fleet commented
on Oct 9, 2026 ContributorAuthorMore actionsLanded, round 3: PR #22456 →
5919483472· 2026-10-09T11:04Zdomain:devxseat 1 ·session_0115N1oNnQS5WqofZ2DzaT3q. The seat verified this onorigin/mainby content.- Merged through the queue at 2026-10-09T11:03Z.
scripts/partition-test-shards.mjsand.github/workflows/ci.ymlonmainequal the ACCEPTed head2ccec0b333.FILE_SHARDED_PACKAGESreads{ '@objectstack/cli': 3 }. - Its merge-group run, 37917885267: all six
Test Core (N/6)jobssuccess, drift step included. Job walls were 9.1 / 3.5 / 2.0 / 1.3 / 1.5 / 1.2 min. Walls that short mean most of its tasks replayed from the turbo cache, so this run says nothing yet about density under load. The density test is the next full run in which most packages actually execute. - Closing-keyword check: this card is open as intended (
Part of). No other card was closed by the PR body.
What remains (this seat owns it), in order:
- The drift watch. At every fire the seat reads the drift step on new
merge_groupand scheduled full runs until at least two runs with mostly executed (not replayed) packages have passed. A single drift red attributable to this split means an immediate revert, as in round 2. - The pin. Read the ≤ 1.3 slowest-over-mean-of-others ratio and the slowest job wall on 3
pull_requestand 2merge_groupruns whose shard log printssliced x3. The seat posts them here with run ids. - The count. From those measured legs, decide whether 3 stays or a follow-up changes it (ACCEPT
6079057565, open question 1).
Labels in this act:
pm:dispatched→pm:blocked. The unlock condition is that the runs in items 1–2 exist; the seat checks for them at each fire.
Generated by Claude Code
- Merged through the queue at 2026-10-09T11:03Z.
objectstack-fleet commented
on Oct 9, 2026 ContributorAuthorMore actionsRound 3, drift watch: item 1 of
6079617893is met ·domain:devxseat 1 · 2026-10-09T11:06ZTwo heavy full runs since
5919483472passed the drift step on all six shards. Each run contains round 3: RESTcompare 5919483472...HEADreadsahead,behind_by 0.run group Test Core (1/6)–(6/6)job walls (min)slowest slowest ÷ mean of the others 37918648139 PR #22441 18.4 / 18.5 / 22.8 / 14.4 / 11.3 / 10.1 22.8 1.57 37919164432 PR #22448 17.1 / 17.7 / 17.2 / 20.4 / 12.8 / 14.2 20.4 1.29 - Before the change: the job holding the whole CLI read 20.5–39.1 min on 12 sampled runs, and 33–39 min on the heavy ones (round 2's window). These two runs' slowest jobs read 20.4 and 22.8 min.
- One light run (mostly turbo replays) also passed. It is not counted.
- Not yet the pin reading. The ≤ 1.3 pin (item 2) needs 3
pull_requestand 2merge_groupruns whose shard log confirmssliced x3, and these two have not had their slicing line read yet. The two ratios above (1.57 and 1.29) are early data, in the range the dev's model predicted for 3 slices (1.28–1.52×).
Still
pm:blockedon item 2. The seat reads it assliced x3runs accumulate, then decides the count (item 3).
Generated by Claude Code
objectstack-fleet commented
on Oct 9, 2026 ContributorAuthorMore actionsRound 3, pin read (interim): the
merge_grouphalf is in, thepull_requesthalf needs 2 more runs ·domain:devxseat 1 · 2026-10-09T12:28ZWhere slicing was read. The shard logs'
slicing:line sits above the 5,000-line tail the log API returns, so slicing is read from each run's aggregateTest Corejob. Its timing table marks@objectstack/clias "all 3 slices" in every run below. One shard log, run 37924680433, also printsslicing: @objectstack/cli: sliced x3 (… densest shard 1146.69s, within the 1772.65s density cap).The runs. These are green CI runs after
5919483472with the CLI sliced. The ratio is the slowestTest Core (N/6)job wall divided by the mean of the other five, from the jobs API.run event job walls 1/6–6/6 (min) slowest ratio 37922022714 merge_group 12.3 / 15.9 / 15.3 / 16.6 / 13.6 / 11.8 16.6 1.20 37922106947 merge_group 18.6 / 16.7 / 18.6 / 15.9 / 16.9 / 16.6 18.6 1.10 37924560197 merge_group 13.1 / 14.5 / 13.6 / 10.1 / 12.0 / 11.3 14.5 1.21 37924620787 merge_group 13.8 / 14.7 / 10.7 / 10.4 / 8.0 / 11.3 14.7 1.35 37924680433 merge_group 14.1 / 13.6 / 16.0 / 14.4 / 9.7 / 6.9 16.0 1.36 37926958001 pull_request 20.9 / 15.4 / 25.4 / 25.2 / 15.8 / 12.9 25.4 1.41 - Merge-group runs: the slowest job read 14.5–18.6 min, against 26.8–35.6 min on round 2's sampled
merge_groupruns before any change. The ratio read 1.10–1.36; 3 of 5 meet ≤ 1.3. - Pull-request runs: one so far, at 1.41, with spec running 1.30× its weight in that run. The pin needs 2 more.
- Drift step: green on every run above.
The seat reads the remaining
pull_requestruns at its next fires, then posts the count decision (item 3 of6079617893). Stillpm:blocked.
Generated by Claude Code
- Merge-group runs: the slowest job read 14.5–18.6 min, against 26.8–35.6 min on round 2's sampled
objectstack-fleet commented
on Oct 9, 2026 ContributorAuthorMore actionsRound 3, pin read (complete):
merge_groupmeets the pin on 3 of 5 runs,pull_requeston 1 of 4 ·domain:devxseat 1 · 2026-10-09T13:46ZThis continues the interim read
6080893252with the same method:- The ratio is the slowest
Test Core (N/6)job wall over the mean of the other five, from the jobs API. - Slicing is read from each run's aggregate
Test Coretiming table, where@objectstack/clishows "all 3 slices". - Every run below is green, drift step included.
New
pull_requestruns since the interim read:run PR job walls 1/6–6/6 (min) slowest ratio packages measured 37932302950 #22486 10.4 / 11.6 / 15.4 / 14.4 / 10.3 / 5.5 15.4 1.47 24 37934454127 #22489 10.8 / 11.8 / 14.0 / 10.5 / 5.4 / 7.5 14.0 1.52 8 37936325084 #22215 17.5 / 18.1 / 20.8 / 20.1 / 11.5 / 13.7 20.8 1.29 70 The full pin set:
merge_group, 5 runs (6080893252): ratios 1.10–1.36, and 3 of 5 meet ≤ 1.3. The slowest job ran 14.5–18.6 min, against 26.8–35.6 min before round 3.pull_request, 4 runs: ratios 1.41 (37926958001), 1.47, 1.52 and 1.29, so 1 of 4 meets ≤ 1.3. The slowest job ran 14.0–25.4 min.- The Done-when pin (≤ 1.3 on affected-set runs) is not met. The wall-time win holds, but balance on the smaller affected sets misses the pin.
Why the misses happen, measured on one run:
- In 37934454127, the aggregate table measured 8 packages: the CLI at 1855.27 s across its 3 slices, and 788.22 s for the other 7 together.
- An even third of the CLI is about 618 s, against a six-shard mean of about 441 s, and the measured slice skew only makes the heaviest third larger. In turbo-window seconds, the ratio at ×3 cannot fall below about 1.5 whatever the packing. Job walls add a fixed setup cost to every shard, which is why they read a little lower.
- 37932302950 (24 packages, the CLI at 1629.88 s) also misses, but its table lists only the top 10 packages. Those do not show whether a CLI third outweighs the mean shard there, so this read does not explain that miss.
- On the one full run (70 packages, 37936325084), the ratio reads 1.29.
Item 3 of
6079617893(the count decision) is now unblocked. It is open for this lane's next seat holder, because this seat closes its shift with this act. The options as the data reads:- (a) Raise the CLI's slice count, a one-line follow-up under this card. The count would be set from the smallest affected sets, where a slice must sit near the mean shard.
- (b) Keep ×3 and put a narrower pin to the maintainer, for example ≤ 1.3 on full runs only. That is a change to the card's Done-when, so it is not the seat's to make.
Labels stay
pm:blockedwith assigneeos-sales, and the hand-over is the ledger on #6023.
Generated by Claude Code
- The ratio is the slowest
objectstack-fleet commented
on Oct 9, 2026 ContributorAuthorMore actions待维护者裁决:CLI 测试切片停在 3 片收口,还是为小 PR 继续拆 ·
domain:devxseat 1 · 2026-10-09T14:28Zos-bill·session_01LYXc6ckoWuZyVZpWYizdMh。⛔ 不是认领。本卡在本次写入中由pm:blocked换成needs-user-decision,并同笔取消os-sales的 assign。
Release: the hold ofdomain:devxseat 1's previous holder (session_0115N1oNnQS5WqofZ2DzaT3q, accountos-sales; claim6077897041) is released into the decision box. Source: that holder's shift brief6082310919on #6023, relay item 1, verbatim "#22075: make the count decision from6082212420. Option (a) raises the slice count in a follow-up; option (b) asks the maintainer for a narrower pin." Next: the maintainer's answer; anydomain:devxseat re-claims frompm:queue.维护者速读
CI 测试分片第 3 轮已落地:合并队列上最慢的测试分片从 27–36 分钟降到 15–19 分钟。卡上的验收线是「最慢分片 ≤ 其余分片平均的 1.3 倍」,大 PR 和合并队列基本达标,但只改了几个包的小 PR 达不到(1.41–1.52 倍)。要达标只能把 CLI 测试再拆细,而每多拆一片,小 PR 就要多构建一次 CLI 的全部依赖。推荐 A:停在 3 片,按已实测的提速验收,做完最后一项(把 45 分钟的分片超时按实测下调)后关卡。 回复一行即可,例如「22075 A」。
一句话问题。 小 PR 上测试分片仍不均衡,但不均衡的原因是其余分片几乎没活干,而不是有人在等最慢的那个。要不要为这个比值继续拆?
Governing text
- 卡面 Done when 第 2 条:"Pin: the measured slowest-shard / mean-of-the-others ratio on affected-set runs is ≤ 1.3, read from the jobs API and quoted in the PR body with run ids"。
- 卡面 Done when 第 3 条:"The
Test Coreshard wall is re-sized from the measured distribution (the [temporary] raiseTest Core (N/6)timeout-minutes 30 → 45 while #16173's shard balance is unfixed — and un-censor the readings that #16173 needs #16445 raise was declared temporary)"(尚未做)。 scripts/partition-test-shards.mjspin 3c(第 3 轮落地):切片数取「全量列表上让 CLI 不再独自成为最慢串行任务」的最小 n,"a slice count a smaller one could replace is one no pin can hold"。全量列表上 n = 3 已满足,所以 4 片及以上被这条 pin 拒绝。- 卡的出处:维护者「CI 优化按照你的建议创建任务」。1.3 这条线是立卡席位的建议,不是维护者原话。
- 是否改协议:不改协议。A 撤掉本卡 Done when 第 2 条;B 改 pin 3c 的推导口径。
前提(每条带复核命令)
- 现状是 3 片。复核:
git show origin/main:scripts/partition-test-shards.mjs | grep -n -A2 "export const FILE_SHARDED_PACKAGES",读到'@objectstack/cli': 3。 - 4 片及以上被 pin 3c 拒绝。复核:
git show origin/main:scripts/partition-test-shards.mjs | grep -n "a slice count a smaller one could replace",命中 1 行(对照词MAX_SHARD_OVER_MEAN同文件必中)。 - 每多一片,承载它的分片多一条 turbo 测试腿,再加一次 CLI 依赖闭包构建。复核:
git show origin/main:scripts/partition-test-shards.mjs | grep -n "a build of the",命中:231。 - 实测(前任席位
6080893252、6082212420,席位未重测):合并队列 5 次比值 1.10–1.36(3 次达标),最慢分片 14.5–18.6 分钟,改前 26.8–35.6;PR 4 次比值 1.29–1.52(1 次达标),最慢分片 14.0–25.4 分钟。只跑 8 个包的那次(37934454127),CLI 的三分之一约 618 秒,而六个分片的平均只有约 441 秒,所以 3 片时比值低不过约 1.5,怎么排都一样。 - 分片超时仍是 45 分钟。复核:
git show origin/main:.github/workflows/ci.yml | grep -n "timeout-minutes: 45",:505是Test Core分片。 - 与 ci: a per-package suite-duration ratchet — a PR that makes a suite exceed its measured ceiling is red; ceilings rise only by ruling (maintainer-directed, growth constraint) #16468 的关联(未验证):首张天花板表周一 05:30 UTC 由刷新任务写入(
shard-timings-refresh.yml:164)。按 ci: a per-package suite-duration ratchet — a PR that makes a suite exceed its measured ceiling is red; ceilings rise only by ruling (maintainer-directed, growth constraint) #16468 的裁定 Q1 A,已有天花板不随刷新变动。CLI 的天花板按 3 片的窗口加总生成;此后改片数,加总窗口会多出每片的启动开销,天花板却不跟着变。
选项
选项 做什么 客户/团队能感知到的后果 A. 停在 3 片,按实测提速收口 撤掉「≤ 1.3 比值」这条验收线,以已实测的提速验收;派一个 dev 按实测分布下调 45 分钟的分片超时(Done when 第 3 条),然后关卡 合并队列最慢分片保持 15–19 分钟;小 PR 最慢分片 14–15 分钟,其余分片早早跑完;不再增加构建量 B. 为小 PR 继续拆 改 pin 3c,按最小受影响集推导片数(约 5–6 片),再测 3 次 PR、2 次合并队列 小 PR 的比值可能达标;但每次 CLI 受影响的 PR 都要在更多分片上构建 CLI 的全部依赖,总算力上升,墙钟收益未测,可能变慢;最好赶在周一首张天花板表之前落地 C. 每次运行现算分片(卡上路线 a) 按本次受影响的包现算分片 改动最大;第 2 轮已因分片密度问题回滚过一次 业务含义直译: A ≈ 收银台已从一条长队变成几条短队,剩下的空闲台就让它空着;B ≈ 为了让每个台的人数一样多,把一位大客户的购物车拆到更多台上,每个台都要多开一次机;C ≈ 每来一批客人就重新排一次台。
四轴(业务立场)
- ① 项目长远合理性: 两年后的样子应该是以墙钟和排队吞吐为目标,而不是追分片均衡比;主流做法(Bazel、Buildkite 的测试分片)也是按耗时历史分片,目标是总时长。小 PR 上这个比值量的是其余分片有多空。A 撤掉一条结构上达不到的线;B 为它增加切片与构建,方向和「零件只减不增」相反。
- ② 实际业务拉动: 拉动在合并队列吞吐和 PR 等待时间,这两项已实测大幅改善。比值不达标的小 PR,最慢分片本身只有 14–15 分钟。
- ③ 防 AI 犯错: A 把一条追不到的验收线从卡上拿掉,下一任席位不会再追一个结构上不可能的目标;B 的风险是比值好看了,墙钟反而变慢,而没有门禁读墙钟。
- ④ 创业阶段不扩散: A 不加零件;B 增加切片数、闭包构建和一次推导口径修改。
os-decision-facets ① 项目长远合理性:A 以墙钟与吞吐为目标、撤掉结构上达不到的比值线;B 为比值增加切片与构建,特例扩大。 ② 实际业务拉动:合并队列最慢分片已从 27–36 分钟降到 15–19 分钟;比值不达标的小 PR,最慢分片只有 14–15 分钟。 ③ 防 AI 犯错:A 不让后续席位追不可能的目标;B 可能让比值达标而墙钟变差,且无门禁读墙钟。 ④ 创业阶段不扩散:A 零新增;B 增加切片数、闭包构建与推导口径改动。 Prior rulings read: shard,slice,partition,timeout-minutes → 6 hits; ADR-0029 D6, ADR-0053 §3, ADR-0057 D2, ADR-0086 D5/D7, ADR-0087 D4 (all incidental: ownership, dates, authorization, upgrades); thread: 0 rulings of 23 comments推荐 A,回退 B。 自检:只看①选 A;②③④ 是否翻转:否。
置信缺口: B 的墙钟影响没有实测;合并队列只有 5 个样本;超时下调到多少,要由 dev 从实测分布推导,席位不预设数字。
裁后执行
- A: 本席派一个 dev 改
.github/workflows/ci.yml的Test Core分片timeout-minutes,按实测分布推导,注释带窗口与数字;PR 正文写明 Done when 第 2 条按本裁决撤销,Fixes #22075。在 ci: a per-package suite-duration ratchet — a PR that makes a suite exceed its measured ceiling is red; ceilings rise only by ruling (maintainer-directed, growth constraint) #16468 的修改落地之后串行派发(同一 job 的相邻文件)。 - B: 本席派一个 dev 改 pin 3c 的推导口径并实测;赶在周一首张天花板表之前落地,否则同 PR 带 ci: a per-package suite-duration ratchet — a PR that makes a suite exceed its measured ceiling is red; ceilings rise only by ruling (maintainer-directed, growth constraint) #16468 的天花板裁定条目。
- C: 本席先写重设计简报,再派发。
objectstack-fleet commented
on Oct 9, 2026 ContributorAuthorMore actionsRuling: letter A (stop at 3 slices; accept on the measured speed-up; re-size the shard timeout, then close) · maintainer, answering in
domain:devxseat 1's session chat · 2026-10-09T14:32Zdomain:devxseat 1 ·os-bill·session_01LYXc6ckoWuZyVZpWYizdMh. Presented as the decision6082946866(this card). The maintainer's answer, verbatim (the option label chosen in the session's question prompt): 「A 停在 3 片收口 (Recommended)」. Thread-read: 6082946866. Freshness: the body is unchanged since filing; no comment after the presentation.What the ruling settles
- A.
FILE_SHARDED_PACKAGESstays{ '@objectstack/cli': 3 }, with pin 3c as built. - Done when, clause 2 (the ≤ 1.3 slowest-over-mean ratio on affected-set runs) is withdrawn. The card is accepted on the measured speed-up instead. The slowest
merge_groupTest Corejob fell from 26.8–35.6 min to 14.5–18.6 min; the readings are6080893252and6082212420. - ⛔ Not taken: B (re-deriving pin 3c on the smallest affected set, about 5–6 slices, which adds a closure build per extra slice), and C (a per-run partition).
What remains on this card
- Done when, clause 3: re-size the
Test Coreshardtimeout-minutes(.github/workflows/ci.yml, the shard job, today 45) from the measured job-wall distribution after5919483472. The rationale comment carries the window and the numbers. The PR closes this card (Fixes #22075). - Serial: it is dispatched after ci: a per-package suite-duration ratchet — a PR that makes a suite exceed its measured ceiling is red; ceilings rise only by ruling (maintainer-directed, growth constraint) #16468's in-flight change lands, because both edit
.github/workflows/ci.yml'sTest Corejobs and this seat dispatches one dev at a time.
Labels in this act:
needs-user-decision→pm:queue.tooling·priority:p2·domain:devx·area:devpathare unchanged; no assignee.- A.
Filing gate: ③ a maintainer-directed task — the maintainer, after this seat's CI assessment in this session, verbatim: 「CI 优化按照你的建议创建任务」 — carrying ① a measured defect with a named fix site: the test-shard balance is derived on full-run package sums, but PR and merge_group runs execute the affected set, and on those runs the CLI shard is about 1.9× the other five and sets CI and merge-queue wall time. Filed by
domain:skillsseat 2 (seat post #19287,session_0181E4ZeZmWyknawnauxD2CE). ⛔ Not a claim. Thetoolingentry rule (triage-duties.md:34) is met by the maintainer's instruction quoted here; the guarded surface is the required contextTest Coreand the merge queue's wall time.Reader: triage first-touch →
domain:devx(the lane of #16173, #16454, #16464, #22014); the devx seat dispatches it. Sequence it after #22014 (open, dispatched: the shard-timings dataset rests on one run; #22022 landed the 3-run refresh), so the re-derivation reads a measured dataset.Dedupe: page-looped REST listings, closed included (
domain:devxsince 2026-09-07: 391;ci/cd: 63; every issue updated since 2026-10-01: 636;toolingsince 2026-09-07: 417;domain:skillssince 2026-09-23: 109; union 1,266) grepped forshard|partition|@objectstack/cli|slowest→ 23 hits. The nearest: #16173 (closednot_plannedunder ruling 202 B — the stale CLI entry, 672 s predicted vs 28m46s measured), #16445 (the "temporary"Test Corewall 30 → 45 min while #16173 was unfixed; still 45 today), #16454 / #16464 / #16222 (closed: publish the slowest packages, scheduled dataset refresh), #21758 / #21826 (the two re-derivations recorded inscripts/partition-test-shards.mjs), #22014 (open) and #16468 (open, blocked on #22014). None measures the imbalance on affected-set runs, which is this card.What is measured
pull_requestrun 37567293616 (branchclaude/issue-22044-…): wall 36.5 min, 169 job-min.Test Core (1/6)35.7 min (the "run shard tests" step 34.2); shards 6/4/3/5/2: 18.0 / 17.9 / 15.4 / 14.6 / 11.8 min. The slowest shard is 1.9× the mean of the other five (15.5).merge_grouprun 37570577136 (queue entry for perf(spec): a bundle that never reads the ADR-0087 conversion table stops keeping it, 226 KB gzip off the console first screen (#22044) #22048): wall 37.1 min;Test Core (1/6)36.2 min (tests 34.9); other shards 11.6–20.2.scripts/test-shard-timings.json, 72 packages, 9,781 s total;@objectstack/cli1,702.69 s is the heaviest item;@objectstack/spec1,134.86 s next.scripts/partition-test-shards.mjs, the comment block ending "n = 1 is the derived answer"): the bound is a RATIO,MAX_SHARD_OVER_MEAN = 1.3; on the committed dataset the CLI whole is 1.044× the mean (bins 2–6 at 1,614–1,617 s), so pin 3c refuses a{ '@objectstack/cli': 2 }slicing entry and the CLI stays whole. The arithmetic is right for a FULL run.pull_requestandmerge_groupruns execute the affected set against the base sha. The CLI is downstream of most packages, so it is in the affected set of most PRs and runs whole (about 1,700 s), while the other five bins run only their affected members and shrink to 700–1,200 s. The bound holds on paper and fails on nearly every PR and queue build.Test Corewall is still the 45 minutes [temporary] raiseTest Core (N/6)timeout-minutes 30 → 45 while #16173's shard balance is unfixed — and un-censor the readings that #16173 needs #16445 raised it to.Done when
pull_requestand 2merge_groupruns before choosing:OS_TEST_SHARDwiring already exist);expandSlices, the existing mechanism) when the affected set makes it the only heavy item, and runs whole on full runs.--check-driftkeeps its meaning.Test Coreshard wall is re-sized from the measured distribution (the [temporary] raiseTest Core (N/6)timeout-minutes 30 → 45 while #16173's shard balance is unfixed — and un-censor the readings that #16173 needs #16445 raise was declared temporary), with the rationale comment carrying the window and numbers.Generated by Claude Code