Skip to content

ci(test-shards): slice the CLI 3 ways at plain weight under a density cap - #22456

Merged
objectstack-fleet[bot] merged 3 commits into
mainfrom
claude/issue-22075-density-capped-slices
Oct 9, 2026
Merged

objectstack-fleet[bot] merged 3 commits into
mainfrom
claude/issue-22075-density-capped-slices

Conversation

@objectstack-fleet

Copy link
Copy Markdown
Contributor

Part of #22075
Clause-②: no

Round 3 of the card. Round 2 (PR #22415) cut @objectstack/cli into slices but weighted each slice at four times its window, which packed the remaining whole packages onto three shards at ~2641 s of predicted windows each; the drift step redded on full runs and the change was reverted (PR #22435, 806b03e2ae). This round keeps round 2's per-run slicing decision and slice-spread check, drops the slot weight and the wall-graded bound, and places the three CLI slices at their plain weight under a density cap: no shard may carry more predicted windows than the densest bin of the whole-package split of the committed dataset's full list, computed by the same pre-change code on every call (1772.65 s today).

This PR is Part of the card: the card's ratio pin is read on real runs after landing, and the Test Core wall re-size is owed after that. See "The 1.3 pin" below for the prediction, which says the pin is likely missed at this slice count.

Done when (quoted verbatim from the card)

Density on the committed dataset (the hard constraint)

All three rows split the same file, scripts/test-shard-timings.json (unchanged since 040184752c), as CI splits the full list (--exclude @objectstack/dogfood, 6 shards), each with that round's own partitioner code:

split code bins of predicted test windows, shard 1..6 (s) densest bin
pre-#22415 (CLI whole) 806b03e2ae 1771.39 / 1772.65 / 1772.18 / 1771.55 / 1772.41 / 1770.59 1772.65 (the cap)
round 2, PR #22415 (slices weighted x4) c64130bafa 904 / 902 / 903 / 2641 / 2641 / 2641 2641 (1.49x the cap)
this round (3 slices at plain weight) this PR 1772.63 / 1771.02 / 1772.63 / 1770.95 / 1771.09 / 1772.46 1772.63, within the cap
  • The pre-change row is node scripts/partition-test-shards.mjs --self-test of 806b03e2ae (it prints bins 1771/1773/1772/1772/1772/1771s), and its exact maximum, 1772.6499999999999 s, read from partition() and balanceOf() of that file. densityCap() in this PR returns the same float: it runs the unchanged partition() on the unsliced items, so the cap is the pre-change computation, derived on every call and never typed.
  • Slices 1/3, 2/3 and 3/3 land on shards 2, 3 and 4. Each of those shards carries one 579.63 s slice plus 1191.3 to 1193.0 s of whole packages. @objectstack/spec (1146.69 s) is on shard 1.
  • The self-test pins it: the committed dataset's full list, split as CI splits it, puts no shard past the cap, and that split carries the slices.

What changed

All in scripts/partition-test-shards.mjs, plus comments in the Test Core job of .github/workflows/ci.yml (comments only: no step, matrix, timeout or context change).

  • FILE_SHARDED_PACKAGES = { '@objectstack/cli': 3 }. PREVIOUS_FILE_SHARDED_PACKAGES becomes the outgoing map, {}.

  • The count is derived on the serial floor (sliceCountProblems(), pin 3c, both halves). A slice runs as a serial leg of its own, so the question is not the bins' sum (slicing never moves the mean, and on bin sums the whole CLI already "fit") but whether the package is the run's serial floor on its own. n is the smallest count whose slice sits within 1.3x the heaviest other serial task on the full list. That task is @objectstack/spec's, half its 1146.69 s test + test:repo sum, 573.35 s (measured: its test:repo ran 547.04 s on merge_group 37875522518). So the limit is 745.35 s: the CLI whole is 1738.88 s (2.33x), at 2 slices 869.44 s (1.52x), at 3 slices 579.63 s (1.01x). 3 is the answer, and the pin refuses both a smaller count ("Raise it") and a larger one ("Lower it").

  • Placement is unchanged LPT (partition() is byte-identical): a slice is one bin item at its plain weight.

  • The density cap (densityCap(), densityCapOf()): the whole-package split's densest bin for the dataset's full list, minus CI's exclusions, at the run's shard count.

  • Repair, not refusal, of LPT noise (refineToCap(), placeItems()). LPT can overshoot the cap by where its last small items fall. On today's dataset the CLI cut 2 or 4 ways lands at 1773.23 s and 1773.16 s against 1772.65 s, while 3 lands at 1772.63 s. A cap that a refresh could breach by that noise would stop the full list slicing on roughly every other refresh. So a sliced split that overshoots moves or swaps whole packages (never a slice) out of its densest bin while that strictly lowers it. It is deterministic and terminates (each move lowers the sum of squared bin totals). A split with no slice is the plain LPT split, byte for byte. On today's dataset the repair is not needed (3 slices land at 1772.63 s), and it brings the 2-way cut to 1772.55 s.

  • The per-run decision (planShards(), kept from round 2, re-graded). A run slices the CLI only when all three hold, and prints the decision in "Compute this shard's package set":

    1. NEEDED: whole, it is past 1.3x the heaviest serial task of every other package in the run.
    2. SPREADS: its slices land on distinct shards.
    3. WITHIN CAP: the sliced split keeps every shard within the cap.

    Example line: slicing: @objectstack/cli: sliced x3 (whole 1739s against 745s, 1.3x the heaviest other serial task; densest shard 1767.91s, within the 1772.65s density cap).

  • Dropped from round 2: the x4 slot weight in partition(), TEST_CONCURRENCY, shardWalls() and the wall-graded pins 2/3. Pins 2 and 3 grade bin sums again, as before ci(test-shards): grade the Test Core split on predicted shard wall and slice the CLI per run #22415.

  • Self-test: a new battery, density cap and the per-run slice decision (#22075), has 14 cases, and the roster floor goes 11 to 12. Pin 3c gains the "Raise it" and two-task-floor cases (the balancing battery now registers 27 cases against its floor of 25).

Assumption 1: the full list balances at ~1772 s per shard (confirmed for bins)

Confirmed as worded: the bins are in the density table above. Three slice shards each carry one 579.63 s slice plus 1191.3 to 1193.0 s of whole packages, and every bin is within the cap. The wall half of the assumption is in the two sections that follow, and that half does not hold at 3 slices.

Assumption 2: slice-shard job wall from components (about 24.5 min, within the stated 20-25)

The components were measured through the jobs API (GET /repos/objectstack-ai/objectstack/actions/runs/RUN_ID/jobs, step timestamps), as round 2 did.

component source min / median / max
setup (steps 1-9) 84 Test Core jobs of the 14 pre-change runs 39 / 65 / 104 s
Build this shard's dependency closure the 12 jobs of the two full-list runs (37872770181, 37870616843) 113 / 170 / 338 s
whole-package leg at the slice shards' density, ~1191 s of predicted windows the 10 pre-change whole-package shards with bins of 1171-1193 s (37874898502 shards 2-6, 37873694430 shards 2-6) 367 / 488 / 753 s
Build the sliced package's dependency closure the 24 slice-shard jobs of round 2's 8 sliced runs (below) 0 / 13 / 14 s
slice leg, heaviest slice (2/3) the Run this shard's tests step of the slice 2/3 shard in round 2's 8 sliced runs. It is an upper bound for the slice alone, since round 2 put a small whole-package leg in front of it. 490 / 722 / 731 s
post (steps 13 on) 84 jobs 9 / 14 / 85 s
  • Round 2's sliced runs: 37892033675, 37893672824, 37894048074, 37894050587, 37894053453, 37895967479, 37894129260 and 37895965974.
  • The full-list slice shard, at the medians: 65 + 170 + 13 + 488 + 722 + 14 = 1472 s, about 24.5 min. Today's worst is 39.1 min, the CLI-alone job of 37872770181. The upper envelope is about 33.8 min: every component's worst, from different runs, never observed together.
  • The whole-package shards of the same list keep today's density, so they keep today's walls: 10.6-20.8 min on the two full-list runs.
  • A cross-check of the slice leg against vitest's own split. vitest 4.1.11 shards by the sha1 of each file's root-relative path (BaseSequencer.shard). On the CLI's 372-file list (vitest list --filesOnly, OS_TEST_TIERS=queue), with the 17 slowest CLI files at their measured seconds (run 37893672824's timing table) and the rest spread evenly, the three slices carry 0.81 / 1.17 / 1.02 of an even third. Round 2's measured slice-shard steps have medians of 546 / 722 / 640 s, the same order.

Assumption 3: per-run before/after, round 2's 14 runs

  • Source. The affected sets are round 2's reproductions (turbo ls --affected between each run's base and head, plus the cross-package union). They are re-split here with this PR's code. The "before" densest bins equal round 2's recorded bins.
  • Before. The jobs API's measured job walls.
  • After. Predicted from each run's own components:
    • setup, closure and post: the run's medians;
    • each whole-package leg: the shard's predicted windows times the run's own measured packing (the sum of its whole-package shards' test steps over the sum of their predicted windows, 0.34-0.58), floored at the shard's heaviest serial task;
    • each slice leg: a third of the run's own CLI-alone test step, times the slice's vitest share above;
    • each slice shard: plus 14 s for its slice closure step.
run event pkgs before: densest bin before: slowest / mean of others (jobs API) before: slowest job after: plan after: densest bin (cap 1772.65 s) after: slowest / mean of others (predicted) after: slowest job (predicted)
37876969409 pull_request 12 1739 s 3.84x 26.3 min CLI x3 1147 s 1.28x 12.3 min
37874898502 pull_request 38 1739 s 3.03x 34.3 min CLI x3 1282 s 1.45x 21.4 min
37873731877 pull_request 11 1739 s 5.65x 32.6 min CLI x3 1147 s 1.52x 14.6 min
37873634681 pull_request 56 1739 s 2.67x 33.8 min CLI x3 1575 s 1.41x 22.4 min
37872770181 pull_request 75 1769 s 2.26x 39.1 min CLI x3 1768 s 1.33x 26.6 min
37871447533 pull_request 3 1147 s 4.64x 9.5 min no CLI: split unchanged 1147 s 4.64x (unchanged) 9.5 min (unchanged)
37875522518 merge_group 48 1739 s 3.32x 33.3 min CLI x3 1461 s 1.50x 20.4 min
37875521531 merge_group 67 1739 s 2.58x 32.5 min CLI x3 1714 s 1.41x 21.1 min
37873846077 merge_group 48 1739 s 2.52x 26.8 min CLI x3 1461 s 1.36x 18.2 min
37873791941 merge_group 67 1739 s 2.61x 33.9 min CLI x3 1714 s 1.41x 21.6 min
37873694430 merge_group 33 1739 s 2.81x 29.9 min CLI x3 1274 s 1.41x 19.1 min
37872756554 merge_group 3 1147 s 7.55x 13.8 min no CLI: split unchanged 1147 s 7.55x (unchanged) 13.8 min (unchanged)
37871575925 merge_group 31 1739 s 2.48x 20.5 min CLI x3 1190 s 1.39x 15.3 min
37870616843 merge_group 75 1769 s 2.43x 35.6 min CLI x3 1768 s 1.37x 23.8 min
  • Slicing still helps whenever the CLI is in the set. All 12 CLI runs slice, and the predicted slowest job falls 26-57% (39.1 to 26.6 min on the full list).
  • No run's densest bin rises. Every after-split is within the cap and at or under the run's own before-split densest bin.
  • The 2 docs-only runs are untouched. Their floor is @objectstack/spec, which cannot be sliced.

The 1.3 pin: predicted to be missed at 3 slices

  • The prediction. The table above reads 1.28-1.52x on the 12 CLI runs: 11 of 12 above 1.3, including both full lists (1.33x and 1.37x).

  • The reason is structural, not noise.

    • Under the cap, the full list's six bins all sit at ~1772 s.
    • A shard without a slice runs its 1772 s of whole packages four at a time (measured packing 0.34-0.58 s of test step per second of predicted window).
    • A slice shard runs 1191 s that way, and then one ~630 s CLI third with nothing to overlap it.
    • So a slice shard's leg is ~1.6-1.8x a whole shard's. Three of the six shards carry one, and the slowest lands ~1.3-1.5x the others' mean.
  • What would meet the pin under the same cap, by the same model (not built here): spreading the CLI's serial work over more shards. The slices are seeded one per shard before the whole packages, so they spread even where they are lighter than other packages. At plain weight under the cap, LPT does not spread 4 or 6 slices on the full list.

    run 3 (this PR) 4, seeded 5, seeded 6, seeded
    37876969409 1.28x, 12.3 min 1.23x, 12.3 min 1.23x, 12.3 min 1.70x, 15.9 min
    37874898502 1.45x, 21.4 min 1.35x, 20.1 min 1.14x, 17.4 min 1.10x, 17.1 min
    37873731877 1.52x, 14.6 min 1.26x, 12.7 min 1.31x, 13.1 min 1.60x, 15.4 min
    37873634681 1.41x, 22.4 min 1.39x, 22.2 min 1.15x, 19.0 min 1.11x, 18.7 min
    37872770181 1.33x, 26.6 min 1.25x, 25.4 min 1.11x, 23.2 min 1.07x, 22.5 min
    37875522518 1.50x, 20.4 min 1.40x, 19.0 min 1.14x, 16.0 min 1.11x, 15.9 min
    37875521531 1.41x, 21.1 min 1.32x, 20.1 min 1.14x, 18.0 min 1.09x, 17.4 min
    37873846077 1.36x, 18.2 min 1.31x, 17.3 min 1.10x, 14.9 min 1.13x, 15.6 min
    37873791941 1.41x, 21.6 min 1.37x, 21.2 min 1.14x, 18.3 min 1.07x, 17.6 min
    37873694430 1.41x, 19.1 min 1.46x, 19.5 min 1.19x, 16.4 min 1.15x, 16.2 min
    37871575925 1.39x, 15.3 min 1.35x, 14.7 min 1.14x, 12.7 min 1.30x, 14.2 min
    37870616843 1.37x, 23.8 min 1.28x, 22.6 min 1.13x, 20.4 min 1.08x, 19.8 min

    Five seeded slices read 1.10-1.31x on all 12 runs, and every bin stays within the cap. Six put a slice on @objectstack/spec's shard and get worse on small sets.

  • Why it is not in this PR. Its count has no packing-free derivation. The serial floor gives 3; 5 comes out only of a wall model with a measured packing constant that no gate re-measures, which is the shape round 2's TEST_CONCURRENCY had. It is a decision for the seat, recorded in the dev report. The model's own limits apply to every column: walls are predicted from per-run components, not measured.

How the seat reads the pin after landing

  1. Take the first 3 or more pull_request and 2 or more merge_group runs on or after this PR's merge commit whose "Compute this shard's package set" step prints slicing: @objectstack/cli: sliced x3.
  2. From GET /repos/objectstack-ai/objectstack/actions/runs/RUN_ID/jobs, take each run's slowest Test Core (N/6) job wall divided by the mean of the other five. The pin is 1.3 or less.
  3. On the same runs, take each run's slowest Test Core job wall. That distribution is the input to the owed timeout-minutes re-size.
  4. A run without the CLI (docs-only, the floor is @objectstack/spec) is read as "slowest = spec", not as a breach. That is pin 3's floor; slicing does not reach it.
  5. Two more readings make the next decision measurable: the slice shards' Run this shard's tests step against the whole shards', and the slice legs on the shard log. Those are the in-situ packing and slice legs a 4- or 5-way count would be derived from.

The Test Core wall: held at 45

timeout-minutes stays 45, and the comment says the re-size is owed after landing, from the measured post-change job walls (read 3 above). pnpm check:stall-guard-budget is green: cap 20 min against a 45-min budget.

--check-drift keeps its meaning on slice-carrying shards

  • There is no change on that path. The 1.5x red and the 1.3x warning keep their values.
  • A slice-carrying shard is predicted its slice at a third of the CLI's entry, from the slice count its own run summary records (the OS_TEST_SHARD digest), not from the config. The ratio is still one shard's executed windows against their own prediction.
  • The density cap is what keeps those predictions honest. The whole-package windows on every shard are now measured at, or under, the density the dataset was measured at. A slice shard runs 1191 s of whole packages, less dense than today's 1772 s, so its whole windows should read at or under their entries, and the slice runs alone, as the CLI did when it was measured.
  • The estimate for the heaviest slice shard is (1191 x r + 1.17 x 579.6) / 1772.6, where r is the measured/predicted ratio of its whole-package windows. Taking r up to 1.21, the worst whole-package shard reading at full density before ci(test-shards): grade the Test Core split on predicted shard wall and slice the CLI per run #22415 (main push 37888117205, Test Core (6/6)), gives at most 1.20x, under the 1.3x warning. The other three shards keep today's density, and today's readings.
  • The drift batteries (9 + 12 cases) are unchanged and green.

What #16468 will read differently

#16468 is needs-user-decision and not in flight. None of it is built here.

  • A per-package ceiling reads the same per-package windows the drift step reads. With density held at the proven level, those windows stay comparable to the dataset they come from. That is the property the cap exists for.
  • The CLI's window now arrives as three slice windows on three shards' summaries, each about a third of the whole, instead of one window on one shard. A ceiling check would compare a slice against its share (the ceiling divided by the slice count the summary records, as predictedSecondsFor does), or sum the parts across shards (report-test-timings.mjs already does, marking a partial set).
  • The dataset still holds each package's whole cost, and the generator still reassembles slices within a run before recording one.

Required contexts and coverage

  • No job, step, matrix or context name changed, so the seven required contexts are untouched. pnpm check:required-contexts exits 0.
  • Every affected package still runs. The three slices are vitest's own --shard partition of the CLI's file list, so their union is the whole suite. check-test-completeness grades every scheduled item, and its self-test is green.
  • scripts/test-shard-timings.json is not touched.

Gates (head 2ccec0b333, after one origin/main merge)

  • node scripts/pm/dispatch-gates.mjs --commands --repo objectstack-ai/objectstack derived 58 commands on 2ccec0b333 (2 paths against the merge base 440bed63e).
  • All 58 ran, each with its exit code captured before any pipe. The 18-minute pnpm check:pm-dispatch-gates battery ran in the background with its exit code written to a file, and I waited for it in the foreground: exit 0, "2011 cases pass".
  • --ran reconciliation: 58 derived, 54 run, 4 NOT-MEASURED, 0 unrun. The 4 NOT MEASURED are check:dts-closure, check:dual-build-cjs-loads, check:lean-entry-closure and check:sourcemap-no-sources-content. Each exited 3, PREREQUISITE NOT MET: they load every package's built dist/, and this diff touches no package.
  • The other 54 exited 0. They include:
    • node scripts/partition-test-shards.mjs --self-test (its ASCII less-or-equal sign spelled out here): self-test OK (72 measured packages -> 74 shard items, 6 shards, max/mean 1.00x LESS-OR-EQUAL 1.3x, floor 1147s, bins 1772.63/1771.02/1772.63/1770.95/1771.09/1772.46s within the 1772.65s density cap, file-level slices: @objectstack/cli x3);
    • pnpm check:required-contexts, pnpm check:stall-guard-budget and pnpm check:stall-guard-headroom.
  • The consumer self-tests are green too: measure-test-shard-timings.mjs (it decodes slices against the new maps {cli: 3} and {}), check-test-completeness.mjs and report-test-timings.mjs.
  • node scripts/check-commit-card-trailers.mjs --range origin/main..HEAD: exit 0 on the 3 branch commits.
  • Lint, a declared narrowing. eslint --no-inline-config --format json scripts/partition-test-shards.mjs reported 1 file, 0 errors, 0 warnings.
    • The population comes from eslint's own config: isPathIgnored is false for the script, and true for ci.yml, which is not an eslint input.
    • The computed config for the file has no parserOptions.project and no projectService, so no type-aware rule runs and this diff cannot move an untouched file's verdict.
    • The full pnpm lint is CI's.

Ablations (the fix committed first; each through node scripts/ablation-replace.mjs)

Each ablation was guarded by a trap that restored the file with git checkout HEAD. The tool confirmed the anchor went 1 to 0 and the blob changed before it ran the self-test, then proved the restore: blob == HEAD 52581cb6eafe, and git diff HEAD empty. The script is a plain node file with no build or dist/ step.

  1. The cap's enforcement removed. The line that sets withinCap to "a densest bin exists AND it is at most cap" became const withinCap = densest !== null; (blob 52581cb6eafe to b3a6ca0dece3). The self-test went RED: density cap: a slicing whose densest shard is past the cap was taken: big: sliced x3 (... densest shard 180.00s ...). Restored.
  2. The repair removed. placeItems()' return became return bins; (blob to fefdda9ddd8c). The self-test went RED: density repair: a split LPT put at 70s was left at 70s, past the 69s cap. Restored. The self-test is green on the restored tree.

Acceptance notes

  • test-nightly-tiers.yml's header still says FILE_SHARDED_PACKAGES "cuts the CLI into two vitest slices". That is outside this card's file surface. That run splits the tier packages over 2 shards, where 3 slices cannot spread, so it keeps running the CLI whole (planShards() says so). Comment only; noted, not filed.
  • The 5-way seeded count above is the next lever under the cap. It needs a decision on how the count is derived, and it is recorded in the dev report. It is not this PR.
  • The Test Core timeout-minutes re-size stays the post-landing half of the card.
  • skip-changeset: root scripts/ and workflows publish nothing.

Generated by Claude Code

claude added 3 commits October 9, 2026 09:27
… cap

Cut @objectstack/cli into 3 file-level slices whenever a run's own list
makes it the serial floor, place the slices at their plain weight so every
bin keeps the whole-package split's density, and refuse any slicing that
would put a shard past the densest bin of the whole-package split of the
committed dataset's full list (derived per call, 1772.65s today). The
slice count is derived on the serial floor: the smallest n whose slice sits
within 1.3x the heaviest other serial task.

Claude-Session: https://claude.ai/code/session_0115N1oNnQS5WqofZ2DzaT3q
Co-authored-by: Claude <noreply@anthropic.com>
…re comments

The comments that described the whole CLI alone on shard 1/6 now describe
the 3-way slicing at plain weight, the density cap that bounds it, and the
held 45-minute wall whose re-size is owed after landing. Comments only: no
step, matrix, timeout or required context changes.

Claude-Session: https://claude.ai/code/session_0115N1oNnQS5WqofZ2DzaT3q
Co-authored-by: Claude <noreply@anthropic.com>
@objectstack-fleet objectstack-fleet Bot added the skip-changeset PR has no user-facing published change; bypasses the changeset gate label Oct 9, 2026
@objectstack-fleet
objectstack-fleet Bot marked this pull request as ready for review October 9, 2026 10:27
@objectstack-fleet
objectstack-fleet Bot enabled auto-merge October 9, 2026 10:27
@objectstack-fleet
objectstack-fleet Bot added this pull request to the merge queue Oct 9, 2026
Merged via the queue into main with commit 5919483 Oct 9, 2026
39 checks passed
@objectstack-fleet
objectstack-fleet Bot deleted the claude/issue-22075-density-capped-slices branch October 9, 2026 11:03
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci/cd size/l skip-changeset PR has no user-facing published change; bypasses the changeset gate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants