Repository navigation
ci(test-shards): grade the Test Core split on predicted shard wall and slice the CLI per run - #22415
Merged
objectstack-fleet[bot] merged 5 commits intoOct 9, 2026
Conversation
…d slice the CLI per run A shard runs its whole packages in one turbo run at --concurrency=4 and each file-level slice in a leg of its own after it, so a bin's summed weight is not its wall time. Graded on sums the split read 1.00-1.47x while the CLI shard ran 2.3-7.6x the other shards' mean. The bound (1.3x) now grades a wall model, the CLI is configured at 3 slices (the smallest count that meets it on the committed dataset), and each run slices it only when its own split needs it and the slices spread. Claude-Session: https://claude.ai/code/session_0115N1oNnQS5WqofZ2DzaT3q Co-authored-by: Claude <noreply@anthropic.com>
…sured distribution The whole CLI no longer runs alone on one shard (it is cut into 3 slices when a run carries it), so the 45-minute wall raised for that shard is re-sized from the 14-run window after the 21-run dataset refresh, with the window, the numbers and the revert condition in the comment. The comments that described the slice steps as idle are brought up to date. Claude-Session: https://claude.ai/code/session_0115N1oNnQS5WqofZ2DzaT3q Co-authored-by: Claude <noreply@anthropic.com>
…fected-set-shard-balance
…ed after landing 35 rested on a prediction of the post-change walls, and the wall model reads low on whole-package legs. The wall stays at 45 and the comment keeps the 14-run window and the predicted figures as the input to the re-size the post-landing reading owes, with the revert condition restated. Claude-Session: https://claude.ai/code/session_0115N1oNnQS5WqofZ2DzaT3q Co-authored-by: Claude <noreply@anthropic.com>
Contributor
Author
|
Generated by Claude Code |
objectstack-fleet
Bot
deleted the
claude/issue-22075-affected-set-shard-balance
branch
October 9, 2026 06:32
This was referenced Oct 9, 2026
os-sales
pushed a commit
that referenced
this pull request
Oct 9, 2026
… wall and slice the CLI per run (#22415)" This reverts commit c64130b. Why: PR #22415 packs about 2000-3000 s of predicted windows onto each whole-package Test Core shard. The dataset's per-package windows were measured on shards carrying about 1,400 s, and under 4-wide concurrency each package's window inflates on the denser shards (objectql 1.84x, plugin-security 1.80x, rest 1.61x, plugin-auth 1.87x). That carries the shard past the "Check this shard's timing drift" step's 1.5x red while the test step itself passes. Red full runs on that partition model: - main push 37894048074 (e02833c; shards 4/6 and 6/6; 6/6 read 3992.4s measured vs 2467.1s predicted = 1.62x) - merge_group builds 37894050587, 37893672824, 37894053453 and 37892033675 The revert restores .github/workflows/ci.yml and scripts/partition-test-shards.mjs byte-for-byte to their content at f05649f, the reverted commit's single parent. A contention-aware redesign is the next round, not a hotfix. Claude-Session: https://claude.ai/code/session_0115N1oNnQS5WqofZ2DzaT3q Co-authored-by: Claude <noreply@anthropic.com>
This was referenced Oct 9, 2026
This was referenced Oct 9, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of #22075
Clause-②: no
Test Core's slowest shard set CI and merge-queue wall time because the split was graded on the wrong quantity. Measurement showed the card's mechanism is half right: the shard assignment is already computed per run from the affected set, and what was wrong is the cost model. A bin's summed package weight is not its wall time, because a shard runs its whole packages four at a time (
--concurrency=4) and each file-level slice in a leg of its own after them. This PR grades the split on predicted shard wall, configures@objectstack/cliat 3 slices (the smallest count that meets 1.3x on the committed dataset), and lets each run slice it only when its own split needs it. The Test Core wall stays at 45 minutes: re-sizing it is owed after landing, from the measured post-change job walls. This PR is thereforePart ofthe card. The card's ratio pin and that re-size are both read after merge.Done when (quoted verbatim from the card)
What measurement found
Route (a) is already the status quo. Every
Test Core (N/6)job runsscripts/ci/select-shard-packages.sh(which writesturbo ls --affectedforpull_requestandmerge_group) and thenpartition-test-shards.mjson that file, so the split is already computed at run time from the run's own affected list with the dataset as weights. Only the slicing refusal (FILE_SHARDED_PACKAGES, empty) was derived on full-run sums.The defect is the cost model, and it shows on full runs too. I reproduced each sampled run's affected set locally (
turbo ls --affectedbetween the run's base and head, plus the cross-package union) and split it with the partitioner. On the runs whose affected set fills six bins, the bin sums read 1.00-1.47x of the mean, inside the bound. The jobs API reads 2.3-3.3x for the same runs. One example, merge_group 37875522518: bin sums were 1739/1397/1399/1403/1400/1398 s, and theRun this shard's testssteps took 1880/548/478/375/556/481 s. Shard 2/6 summed 1036.6 s of package windows, and its turboTime: 9m8.239sequals its heaviest single task (@objectstack/spectest:repo, 547.04 s). Shard 1/6 was the CLI alone, 1867.26 s.The measured runs
These are all completed, successful CI runs after the 21-run dataset (0401847), from 2026-10-09T00:03Z to 02:57Z: 6
pull_requestand 8merge_group. Walls come from the jobs API; bins come from this script on each run's reproduced affected set.spec,rest,create-objectstack). On those, the slowest shard is@objectstack/spec, which cannot be sliced, so this change leaves them untouched.Route chosen: (b), graded on shard walls
Wall model (
TEST_CONCURRENCY = 4,shardWalls()). A shard's predicted wall is the whole-package leg plus every slice leg after it. The whole-package leg is max(heaviest serial task, summed weight / 4). A package that runstestandtest:repoconcurrently counts half its weight as its serial task, because the dataset holds only their sum. At concurrency 1 this formula equals the bin sum, so the old model is the special case. The pre-existing slice-count fixtures pass 1 explicitly and keep their arithmetic.Placement (
partition()). A slice costs its bin weight × 4, because it holds the runner for its whole window after the whole-package leg. Whole packages still cost their weight, so a split with no slices places exactly as before.Derivation, committed dataset (pins 2/3/3c, now graded on walls):
3 is therefore the smallest count that meets the bound, and pin 3c holds it there.
PREVIOUS_FILE_SHARDED_PACKAGESbecomes the outgoing map{}.Per run (
planShards()). A run slices a configured package only when two things hold: its whole suite is the item past the bound (heavier than 1.3x the larger of the run's mean shard wall and any other package's serial task), and its slices spread over the run's shard count. Otherwise it runs whole. The decision prints in "Compute this shard's package set", for exampleslicing: @objectstack/cli: sliced x3 (whole would be 1739s against 818s).Why not a free per-run count: the generator decodes a slice from the sha256 of
OS_TEST_SHARDagainst the counts the map names. A per-run derivation also gives 3 on all 12 CLI runs above, because 4 slices never spread there."Runs whole on full runs" (route (b)'s wording) does not survive measurement. On walls, the full list needs the slices as much as any PR does (2.52x whole). The nightly tier run (CLI alone, 2 shards) cannot spread 3 slices, so it keeps running the CLI whole, as it does today.
Assumption 5 (can the matrix take a run-time assignment without renaming the required contexts?): yes. The matrix already takes one. The job names (
Test Core (N/6), and the aggregateTest Core) are untouched, andpnpm check:required-contextsis green.The ≤ 1.3 pin
In code: pins 2/3/3c grade the committed dataset on walls: 1.01x at 6 shards, floor 580 s (one CLI slice) against a 663 s mean wall. A new self-test battery has 13 cases. It reproduces the measured shape (bin sums inside the bound, one serial suite at ~5x the other walls) and checks that the sum model keeps it whole while the wall model slices it and meets the bound. It also covers the per-run cases (package absent, another unsliceable task as the floor, slices that cannot spread) and pins the concurrency against ci.yml's
--concurrency=4.Measured on this PR's own runs: not provable, and I'm saying so. This PR changes
scripts/partition-test-shards.mjsandci.yml. Its affected set is@objectstack/spec,@objectstack/clientand@objectstack/driver-sql(cross-package union), with no CLI, so its CI and its queue run exercise no slice. On the 12 CLI runs above, the predicted wall ratio after the change is 1.01-1.21x.How the seat reads it after landing. Take the first ≥ 3
pull_requestruns and ≥ 2merge_groupruns whose "Compute this shard's package set" step printsslicing: @objectstack/cli: sliced x3. FromGET /repos/objectstack-ai/objectstack/actions/runs/RUN_ID/jobs, read two things on the same runs:Test Core (N/6)job wall divided by the mean of the other five, per run.Post the run ids, the ratios and the slowest-job walls on the card. The one-line
timeout-minutesre-size under this card is sized from that distribution. The 2 docs-only runs above show that the ratio cannot be reached when an affected set has fewer heavy items than shards (pin 3's floor). Those runs should be read as "slowest = spec", not as a breach.Before / after wall, one PR run and one queue run
The "after" half is a model reading. It cannot be measured on this PR's runs (see above), and it is what the seat's post-landing read replaces. One known bias, measured: the model is a lower bound on whole-package legs. On 37872770181 the five whole-package shards predicted 573/442/442/442/491 s and actually stepped 905/737/704/668/388 s.
The Test Core wall: held at 45, re-size owed after landing
timeout-minutesstays at 45. Removing the whole-CLI shard removes the reason for the old raise, but a wall is re-sized from a measured distribution. The post-change distribution does not exist until runs execute the new split, and a miss would killmerge_groupruns for every lane. The rationale comment inci.ymlrecords the inputs for the re-size and states that none of them has been applied. Window: the 14 runs above, 37870616843 to 37876969409.pnpm check:stall-guard-budgetis green at 45: the cap is 20 min against a 45-min budget, leaving 25 min of slack.--check-driftkeeps its meaningThere is no code change on that path. Both the 1.5x red and the 1.3x warning keep their values. On affected-set runs the ratio is still one shard's executed windows against their own prediction: a shard carrying a CLI slice is predicted a third of the CLI's weight, from the slice count its own summary records (the
OS_TEST_SHARDdigest), not from the config. The drift batteries (9 + 12 cases) are unchanged and green. Slice skew, estimated from the 367-file list the CLI runs and the 18 slowest CLI files of run 37875522518, is 0.87/1.14/0.99 of an even third. A slice-carrying shard should therefore stay under the 1.3x warning.What #16468 will read differently
--check-driftdoes throughpredictedSecondsFor). Or it can sum the parts across shards:report-test-timings.mjsalready does this and marks a partial set.Run this shard's testsstep.Required contexts and coverage
--shardpartition of its file list, so their union is the whole suite.check-test-completenessstill grades every scheduled item, and its self-test is green.scripts/test-shard-timings.jsonis not touched.Gates (head 6e8b8f6; patch round 1, no
origin/mainmerge this round)node scripts/pm/dispatch-gates.mjs --commands --repo objectstack-ai/objectstackre-derived the same 58 commands on 6e8b8f6. All 58 ran, with exits captured before any pipe.--ranreconciliation: 58 derived, 54 run, 4 NOT-MEASURED, 0 UNRUN.check:dts-closure,check:dual-build-cjs-loads,check:lean-entry-closureandcheck:sourcemap-no-sources-content. All exited 3 (PREREQUISITE NOT MET) because they load every package's builtdist/, and this diff touches no package source or build config.pnpm check:pm-dispatch-gatesexited 1 on 1 of 2011 cases, the same as in round 0: "no mkdtempSync site in this tree takes a base the scan cannot read". It namespackages/qa/dogfood/test/security-catalog-cold-boot-environment-holder.dogfood.test.ts:108. That file came from e030d43, which is already onmain, and this diff does not touch it.node scripts/partition-test-shards.mjs --self-test,pnpm check:stall-guard-budget(cap 20 min against a 45-min budget, 25 min of slack),pnpm check:stall-guard-headroomandpnpm check:required-contexts.measure-test-shard-timings.mjs,check-test-completeness.mjsandreport-test-timings.mjs. They readpartition-test-shards.mjs, which is byte-identical at 6e8b8f6.eslint --no-inline-config --format json scripts/partition-test-shards.mjsreported 1 file, 0 errors, 0 warnings.ci.ymlis not an eslint input.parserOptions.projectis null, so this diff cannot move any untouched file's verdict.pnpm lintis CI's.Ablation (run at 9156fd4;
scripts/partition-test-shards.mjsis byte-identical at the current head; fix committed first; restored bygit checkout HEAD, proven blob == HEAD andgit diff HEADempty)Both ablations went through
node scripts/ablation-replace.mjsin WRAP mode, which confirms the anchor hits 1 → 0 and records the blob change before running the command.Math.max(serial, whole / concurrency) + slicedwithwhole + sliced(blob 739b9274f6d1 → 22c7fd4d7ccd). The self-test went RED:slice spread: the committed dataset, split as CI splits it, carries no slice -- @objectstack/cli: whole (whole fits: 1739s is within 2303s). That is the old reading, which kept the CLI whole. Restored to 739b9274f6d1.const needed = own > target;withconst needed = false;(blob 739b9274f6d1 → 7d1f33503974). The self-test went RED:balance: at 6 shards the slowest predicted shard wall is 2.52x the mean (1739s vs 689s), past the 1.3x bound. Restored to 739b9274f6d1, and the self-test is green again on the restored tree.Acceptance notes
test-nightly-tiers.yml's header still saysFILE_SHARDED_PACKAGES"cuts the CLI into two vitest slices". That has been stale since the map emptied, and after this PR the CLI is configured at 3, which cannot spread on that workflow's 2 shards, so it runs whole there. Comment only, outside this card's file surface. Noted, not filed.timeout-minutesre-size is the post-landing half of this card. Its inputs are in theci.ymlcomment and in the section above.skip-changeset: rootscripts/and workflows publish nothing.Generated by Claude Code