Repository navigation
ci(test-shards): slice the CLI 3 ways at plain weight under a density cap - #22456
Merged
objectstack-fleet[bot] merged 3 commits intoOct 9, 2026
Merged
Conversation
… cap Cut @objectstack/cli into 3 file-level slices whenever a run's own list makes it the serial floor, place the slices at their plain weight so every bin keeps the whole-package split's density, and refuse any slicing that would put a shard past the densest bin of the whole-package split of the committed dataset's full list (derived per call, 1772.65s today). The slice count is derived on the serial floor: the smallest n whose slice sits within 1.3x the heaviest other serial task. Claude-Session: https://claude.ai/code/session_0115N1oNnQS5WqofZ2DzaT3q Co-authored-by: Claude <noreply@anthropic.com>
…re comments The comments that described the whole CLI alone on shard 1/6 now describe the 3-way slicing at plain weight, the density cap that bounds it, and the held 45-minute wall whose re-size is owed after landing. Comments only: no step, matrix, timeout or required context changes. Claude-Session: https://claude.ai/code/session_0115N1oNnQS5WqofZ2DzaT3q Co-authored-by: Claude <noreply@anthropic.com>
…nsity-capped-slices
objectstack-fleet
Bot
deleted the
claude/issue-22075-density-capped-slices
branch
October 9, 2026 11:03
This was referenced Oct 9, 2026
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of #22075
Clause-②: no
Round 3 of the card. Round 2 (PR #22415) cut
@objectstack/cliinto slices but weighted each slice at four times its window, which packed the remaining whole packages onto three shards at ~2641 s of predicted windows each; the drift step redded on full runs and the change was reverted (PR #22435,806b03e2ae). This round keeps round 2's per-run slicing decision and slice-spread check, drops the slot weight and the wall-graded bound, and places the three CLI slices at their plain weight under a density cap: no shard may carry more predicted windows than the densest bin of the whole-package split of the committed dataset's full list, computed by the same pre-change code on every call (1772.65 s today).This PR is
Part ofthe card: the card's ratio pin is read on real runs after landing, and the Test Core wall re-size is owed after that. See "The 1.3 pin" below for the prediction, which says the pin is likely missed at this slice count.Done when (quoted verbatim from the card)
Density on the committed dataset (the hard constraint)
All three rows split the same file,
scripts/test-shard-timings.json(unchanged since040184752c), as CI splits the full list (--exclude @objectstack/dogfood, 6 shards), each with that round's own partitioner code:806b03e2aec64130bafanode scripts/partition-test-shards.mjs --self-testof806b03e2ae(it printsbins 1771/1773/1772/1772/1772/1771s), and its exact maximum, 1772.6499999999999 s, read frompartition()andbalanceOf()of that file.densityCap()in this PR returns the same float: it runs the unchangedpartition()on the unsliced items, so the cap is the pre-change computation, derived on every call and never typed.@objectstack/spec(1146.69 s) is on shard 1.What changed
All in
scripts/partition-test-shards.mjs, plus comments in theTest Corejob of.github/workflows/ci.yml(comments only: no step, matrix, timeout or context change).FILE_SHARDED_PACKAGES = { '@objectstack/cli': 3 }.PREVIOUS_FILE_SHARDED_PACKAGESbecomes the outgoing map,{}.The count is derived on the serial floor (
sliceCountProblems(), pin 3c, both halves). A slice runs as a serial leg of its own, so the question is not the bins' sum (slicing never moves the mean, and on bin sums the whole CLI already "fit") but whether the package is the run's serial floor on its own. n is the smallest count whose slice sits within 1.3x the heaviest other serial task on the full list. That task is@objectstack/spec's, half its 1146.69 stest+test:reposum, 573.35 s (measured: itstest:reporan 547.04 s on merge_group 37875522518). So the limit is 745.35 s: the CLI whole is 1738.88 s (2.33x), at 2 slices 869.44 s (1.52x), at 3 slices 579.63 s (1.01x). 3 is the answer, and the pin refuses both a smaller count ("Raise it") and a larger one ("Lower it").Placement is unchanged LPT (
partition()is byte-identical): a slice is one bin item at its plain weight.The density cap (
densityCap(),densityCapOf()): the whole-package split's densest bin for the dataset's full list, minus CI's exclusions, at the run's shard count.Repair, not refusal, of LPT noise (
refineToCap(),placeItems()). LPT can overshoot the cap by where its last small items fall. On today's dataset the CLI cut 2 or 4 ways lands at 1773.23 s and 1773.16 s against 1772.65 s, while 3 lands at 1772.63 s. A cap that a refresh could breach by that noise would stop the full list slicing on roughly every other refresh. So a sliced split that overshoots moves or swaps whole packages (never a slice) out of its densest bin while that strictly lowers it. It is deterministic and terminates (each move lowers the sum of squared bin totals). A split with no slice is the plain LPT split, byte for byte. On today's dataset the repair is not needed (3 slices land at 1772.63 s), and it brings the 2-way cut to 1772.55 s.The per-run decision (
planShards(), kept from round 2, re-graded). A run slices the CLI only when all three hold, and prints the decision in "Compute this shard's package set":Example line:
slicing: @objectstack/cli: sliced x3 (whole 1739s against 745s, 1.3x the heaviest other serial task; densest shard 1767.91s, within the 1772.65s density cap).Dropped from round 2: the x4 slot weight in
partition(),TEST_CONCURRENCY,shardWalls()and the wall-graded pins 2/3. Pins 2 and 3 grade bin sums again, as before ci(test-shards): grade the Test Core split on predicted shard wall and slice the CLI per run #22415.Self-test: a new battery,
density cap and the per-run slice decision (#22075), has 14 cases, and the roster floor goes 11 to 12. Pin 3c gains the "Raise it" and two-task-floor cases (the balancing battery now registers 27 cases against its floor of 25).Assumption 1: the full list balances at ~1772 s per shard (confirmed for bins)
Confirmed as worded: the bins are in the density table above. Three slice shards each carry one 579.63 s slice plus 1191.3 to 1193.0 s of whole packages, and every bin is within the cap. The wall half of the assumption is in the two sections that follow, and that half does not hold at 3 slices.
Assumption 2: slice-shard job wall from components (about 24.5 min, within the stated 20-25)
The components were measured through the jobs API (
GET /repos/objectstack-ai/objectstack/actions/runs/RUN_ID/jobs, step timestamps), as round 2 did.Build this shard's dependency closureBuild the sliced package's dependency closureRun this shard's testsstep of the slice 2/3 shard in round 2's 8 sliced runs. It is an upper bound for the slice alone, since round 2 put a small whole-package leg in front of it.BaseSequencer.shard). On the CLI's 372-file list (vitest list --filesOnly,OS_TEST_TIERS=queue), with the 17 slowest CLI files at their measured seconds (run 37893672824's timing table) and the rest spread evenly, the three slices carry 0.81 / 1.17 / 1.02 of an even third. Round 2's measured slice-shard steps have medians of 546 / 722 / 640 s, the same order.Assumption 3: per-run before/after, round 2's 14 runs
turbo ls --affectedbetween each run's base and head, plus the cross-package union). They are re-split here with this PR's code. The "before" densest bins equal round 2's recorded bins.@objectstack/spec, which cannot be sliced.The 1.3 pin: predicted to be missed at 3 slices
The prediction. The table above reads 1.28-1.52x on the 12 CLI runs: 11 of 12 above 1.3, including both full lists (1.33x and 1.37x).
The reason is structural, not noise.
What would meet the pin under the same cap, by the same model (not built here): spreading the CLI's serial work over more shards. The slices are seeded one per shard before the whole packages, so they spread even where they are lighter than other packages. At plain weight under the cap, LPT does not spread 4 or 6 slices on the full list.
Five seeded slices read 1.10-1.31x on all 12 runs, and every bin stays within the cap. Six put a slice on
@objectstack/spec's shard and get worse on small sets.Why it is not in this PR. Its count has no packing-free derivation. The serial floor gives 3; 5 comes out only of a wall model with a measured packing constant that no gate re-measures, which is the shape round 2's
TEST_CONCURRENCYhad. It is a decision for the seat, recorded in the dev report. The model's own limits apply to every column: walls are predicted from per-run components, not measured.How the seat reads the pin after landing
pull_requestand 2 or moremerge_groupruns on or after this PR's merge commit whose "Compute this shard's package set" step printsslicing: @objectstack/cli: sliced x3.GET /repos/objectstack-ai/objectstack/actions/runs/RUN_ID/jobs, take each run's slowestTest Core (N/6)job wall divided by the mean of the other five. The pin is 1.3 or less.timeout-minutesre-size.@objectstack/spec) is read as "slowest = spec", not as a breach. That is pin 3's floor; slicing does not reach it.Run this shard's testsstep against the whole shards', and the slice legs on the shard log. Those are the in-situ packing and slice legs a 4- or 5-way count would be derived from.The Test Core wall: held at 45
timeout-minutesstays 45, and the comment says the re-size is owed after landing, from the measured post-change job walls (read 3 above).pnpm check:stall-guard-budgetis green: cap 20 min against a 45-min budget.--check-driftkeeps its meaning on slice-carrying shardsOS_TEST_SHARDdigest), not from the config. The ratio is still one shard's executed windows against their own prediction.Test Core (6/6)), gives at most 1.20x, under the 1.3x warning. The other three shards keep today's density, and today's readings.What #16468 will read differently
#16468 is
needs-user-decisionand not in flight. None of it is built here.predictedSecondsFordoes), or sum the parts across shards (report-test-timings.mjsalready does, marking a partial set).Required contexts and coverage
pnpm check:required-contextsexits 0.--shardpartition of the CLI's file list, so their union is the whole suite.check-test-completenessgrades every scheduled item, and its self-test is green.scripts/test-shard-timings.jsonis not touched.Gates (head
2ccec0b333, after oneorigin/mainmerge)node scripts/pm/dispatch-gates.mjs --commands --repo objectstack-ai/objectstackderived 58 commands on2ccec0b333(2 paths against the merge base440bed63e).pnpm check:pm-dispatch-gatesbattery ran in the background with its exit code written to a file, and I waited for it in the foreground: exit 0, "2011 cases pass".--ranreconciliation: 58 derived, 54 run, 4 NOT-MEASURED, 0 unrun. The 4 NOT MEASURED arecheck:dts-closure,check:dual-build-cjs-loads,check:lean-entry-closureandcheck:sourcemap-no-sources-content. Each exited 3, PREREQUISITE NOT MET: they load every package's builtdist/, and this diff touches no package.node scripts/partition-test-shards.mjs --self-test(its ASCII less-or-equal sign spelled out here):self-test OK (72 measured packages -> 74 shard items, 6 shards, max/mean 1.00x LESS-OR-EQUAL 1.3x, floor 1147s, bins 1772.63/1771.02/1772.63/1770.95/1771.09/1772.46s within the 1772.65s density cap, file-level slices: @objectstack/cli x3);pnpm check:required-contexts,pnpm check:stall-guard-budgetandpnpm check:stall-guard-headroom.measure-test-shard-timings.mjs(it decodes slices against the new maps{cli: 3}and{}),check-test-completeness.mjsandreport-test-timings.mjs.node scripts/check-commit-card-trailers.mjs --range origin/main..HEAD: exit 0 on the 3 branch commits.eslint --no-inline-config --format json scripts/partition-test-shards.mjsreported 1 file, 0 errors, 0 warnings.isPathIgnoredis false for the script, and true forci.yml, which is not an eslint input.parserOptions.projectand noprojectService, so no type-aware rule runs and this diff cannot move an untouched file's verdict.pnpm lintis CI's.Ablations (the fix committed first; each through
node scripts/ablation-replace.mjs)Each ablation was guarded by a trap that restored the file with
git checkout HEAD. The tool confirmed the anchor went 1 to 0 and the blob changed before it ran the self-test, then proved the restore: blob == HEAD52581cb6eafe, andgit diff HEADempty. The script is a plain node file with no build ordist/step.withinCapto "a densest bin exists AND it is at mostcap" becameconst withinCap = densest !== null;(blob52581cb6eafetob3a6ca0dece3). The self-test went RED:density cap: a slicing whose densest shard is past the cap was taken: big: sliced x3 (... densest shard 180.00s ...). Restored.placeItems()' return becamereturn bins;(blob tofefdda9ddd8c). The self-test went RED:density repair: a split LPT put at 70s was left at 70s, past the 69s cap. Restored. The self-test is green on the restored tree.Acceptance notes
test-nightly-tiers.yml's header still saysFILE_SHARDED_PACKAGES"cuts the CLI into two vitest slices". That is outside this card's file surface. That run splits the tier packages over 2 shards, where 3 slices cannot spread, so it keeps running the CLI whole (planShards()says so). Comment only; noted, not filed.timeout-minutesre-size stays the post-landing half of the card.skip-changeset: rootscripts/and workflows publish nothing.Generated by Claude Code