Skip to content

ci(test-shards): grade the Test Core split on predicted shard wall and slice the CLI per run - #22415

Merged
objectstack-fleet[bot] merged 5 commits into
mainfrom
claude/issue-22075-affected-set-shard-balance
Oct 9, 2026
Merged

objectstack-fleet[bot] merged 5 commits into
mainfrom
claude/issue-22075-affected-set-shard-balance

Conversation

@objectstack-fleet

@objectstack-fleet objectstack-fleet Bot commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Part of #22075
Clause-②: no

Test Core's slowest shard set CI and merge-queue wall time because the split was graded on the wrong quantity. Measurement showed the card's mechanism is half right: the shard assignment is already computed per run from the affected set, and what was wrong is the cost model. A bin's summed package weight is not its wall time, because a shard runs its whole packages four at a time (--concurrency=4) and each file-level slice in a leg of its own after them. This PR grades the split on predicted shard wall, configures @objectstack/cli at 3 slices (the smallest count that meets 1.3x on the committed dataset), and lets each run slice it only when its own split needs it. The Test Core wall stays at 45 minutes: re-sizing it is owed after landing, from the measured post-change job walls. This PR is therefore Part of the card. The card's ratio pin and that re-size are both read after merge.

Done when (quoted verbatim from the card)

What measurement found

Route (a) is already the status quo. Every Test Core (N/6) job runs scripts/ci/select-shard-packages.sh (which writes turbo ls --affected for pull_request and merge_group) and then partition-test-shards.mjs on that file, so the split is already computed at run time from the run's own affected list with the dataset as weights. Only the slicing refusal (FILE_SHARDED_PACKAGES, empty) was derived on full-run sums.

The defect is the cost model, and it shows on full runs too. I reproduced each sampled run's affected set locally (turbo ls --affected between the run's base and head, plus the cross-package union) and split it with the partitioner. On the runs whose affected set fills six bins, the bin sums read 1.00-1.47x of the mean, inside the bound. The jobs API reads 2.3-3.3x for the same runs. One example, merge_group 37875522518: bin sums were 1739/1397/1399/1403/1400/1398 s, and the Run this shard's tests steps took 1880/548/478/375/556/481 s. Shard 2/6 summed 1036.6 s of package windows, and its turbo Time: 9m8.239s equals its heaviest single task (@objectstack/spec test:repo, 547.04 s). Shard 1/6 was the CLI alone, 1867.26 s.

The measured runs

These are all completed, successful CI runs after the 21-run dataset (0401847), from 2026-10-09T00:03Z to 02:57Z: 6 pull_request and 8 merge_group. Walls come from the jobs API; bins come from this script on each run's reproduced affected set.

run event packages CLI bin-sum max/mean Test Core job walls 1..6 (min) slowest / mean of others run wall (min) after: plan after: predicted wall max/mean
37876969409 pull_request 12 yes 2.32x 26.3 / 14.0 / 7.2 / 6.8 / 3.1 / 3.2 3.84x 27.1 CLI x3 1.21x
37874898502 pull_request 38 yes 1.37x 34.3 / 13.9 / 10.0 / 11.8 / 11.4 / 9.4 3.03x 37.0 CLI x3 1.06x
37873731877 pull_request 11 yes 2.43x 32.6 / 11.0 / 5.6 / 7.0 / 2.5 / 2.7 5.65x 33.4 CLI x3 1.21x
37873634681 pull_request 56 yes 1.11x 33.8 / 18.8 / 10.6 / 9.7 / 12.6 / 11.7 2.67x 36.0 CLI x3 1.02x
37872770181 pull_request 75 yes 1.00x 39.1 / 20.8 / 17.9 / 18.5 / 18.6 / 10.6 2.26x 42.0 CLI x3 1.01x
37871447533 pull_request 3 no 4.10x 9.5 / 4.5 / 1.4 / 1.5 / 1.5 / 1.3 4.64x 11.1 no slice 4.05x
37875522518 merge_group 48 yes 1.19x 33.3 / 11.1 / 10.0 / 8.2 / 10.8 / 10.0 3.32x 34.0 CLI x3 1.01x
37875521531 merge_group 67 yes 1.02x 32.5 / 12.4 / 13.8 / 13.6 / 11.9 / 11.2 2.58x 33.3 CLI x3 1.02x
37873846077 merge_group 48 yes 1.19x 26.8 / 11.5 / 9.1 / 12.4 / 11.7 / 8.4 2.52x 27.6 CLI x3 1.01x
37873791941 merge_group 67 yes 1.02x 33.9 / 14.5 / 12.6 / 14.8 / 14.0 / 9.1 2.61x 34.7 CLI x3 1.02x
37873694430 merge_group 33 yes 1.37x 29.9 / 11.9 / 10.6 / 10.8 / 10.2 / 9.9 2.81x 30.8 CLI x3 1.06x
37872756554 merge_group 3 no 4.10x 13.8 / 3.9 / 1.5 / 0.9 / 1.3 / 1.6 7.55x 15.2 no slice 4.05x
37871575925 merge_group 31 yes 1.47x 20.5 / 10.7 / 6.5 / 6.9 / 8.8 / 8.4 2.48x 21.3 CLI x3 1.08x
37870616843 merge_group 75 yes 1.00x 35.6 / 18.0 / 16.0 / 15.8 / 12.4 / 11.1 2.43x 36.5 CLI x3 1.01x
  • The CLI was in the affected set of 12 of the 14 runs (86%). The two runs without it are a 3-package docs diff (spec, rest, create-objectstack). On those, the slowest shard is @objectstack/spec, which cannot be sliced, so this change leaves them untouched.
  • The "after" columns are the new split's decision and its predicted wall ratio, computed on the same reproduced affected sets.

Route chosen: (b), graded on shard walls

  • Wall model (TEST_CONCURRENCY = 4, shardWalls()). A shard's predicted wall is the whole-package leg plus every slice leg after it. The whole-package leg is max(heaviest serial task, summed weight / 4). A package that runs test and test:repo concurrently counts half its weight as its serial task, because the dataset holds only their sum. At concurrency 1 this formula equals the bin sum, so the old model is the special case. The pre-existing slice-count fixtures pass 1 explicitly and keep their arithmetic.

  • Placement (partition()). A slice costs its bin weight × 4, because it holds the runner for its whole window after the whole-package leg. Whole packages still cost their weight, so a split with no slices places exactly as before.

  • Derivation, committed dataset (pins 2/3/3c, now graded on walls):

    model CLI whole at 2 at 3
    bin sums (concurrency 1) 1.00x meets 1.00x meets 1.00x meets
    shard walls (4) 2.52x 1.31x 1.01x meets

    3 is therefore the smallest count that meets the bound, and pin 3c holds it there. PREVIOUS_FILE_SHARDED_PACKAGES becomes the outgoing map {}.

  • Per run (planShards()). A run slices a configured package only when two things hold: its whole suite is the item past the bound (heavier than 1.3x the larger of the run's mean shard wall and any other package's serial task), and its slices spread over the run's shard count. Otherwise it runs whole. The decision prints in "Compute this shard's package set", for example slicing: @objectstack/cli: sliced x3 (whole would be 1739s against 818s).

  • Why not a free per-run count: the generator decodes a slice from the sha256 of OS_TEST_SHARD against the counts the map names. A per-run derivation also gives 3 on all 12 CLI runs above, because 4 slices never spread there.

  • "Runs whole on full runs" (route (b)'s wording) does not survive measurement. On walls, the full list needs the slices as much as any PR does (2.52x whole). The nightly tier run (CLI alone, 2 shards) cannot spread 3 slices, so it keeps running the CLI whole, as it does today.

Assumption 5 (can the matrix take a run-time assignment without renaming the required contexts?): yes. The matrix already takes one. The job names (Test Core (N/6), and the aggregate Test Core) are untouched, and pnpm check:required-contexts is green.

The ≤ 1.3 pin

  • In code: pins 2/3/3c grade the committed dataset on walls: 1.01x at 6 shards, floor 580 s (one CLI slice) against a 663 s mean wall. A new self-test battery has 13 cases. It reproduces the measured shape (bin sums inside the bound, one serial suite at ~5x the other walls) and checks that the sum model keeps it whole while the wall model slices it and meets the bound. It also covers the per-run cases (package absent, another unsliceable task as the floor, slices that cannot spread) and pins the concurrency against ci.yml's --concurrency=4.

  • Measured on this PR's own runs: not provable, and I'm saying so. This PR changes scripts/partition-test-shards.mjs and ci.yml. Its affected set is @objectstack/spec, @objectstack/client and @objectstack/driver-sql (cross-package union), with no CLI, so its CI and its queue run exercise no slice. On the 12 CLI runs above, the predicted wall ratio after the change is 1.01-1.21x.

  • How the seat reads it after landing. Take the first ≥ 3 pull_request runs and ≥ 2 merge_group runs whose "Compute this shard's package set" step prints slicing: @objectstack/cli: sliced x3. From GET /repos/objectstack-ai/objectstack/actions/runs/RUN_ID/jobs, read two things on the same runs:

    1. The ratio pin: the slowest Test Core (N/6) job wall divided by the mean of the other five, per run.
    2. The input for the follow-up wall re-size: the post-change Test Core job-wall distribution, meaning the slowest Test Core job wall of each run, together with its range across the runs.

    Post the run ids, the ratios and the slowest-job walls on the card. The one-line timeout-minutes re-size under this card is sized from that distribution. The 2 docs-only runs above show that the ratio cannot be reached when an affected set has fewer heavy items than shards (pin 3's floor). Those runs should be read as "slowest = spec", not as a breach.

Before / after wall, one PR run and one queue run

run before: run wall before: slowest Test Core job after (predicted, same affected set)
pull_request 37872770181 (full 75-package list) 42.0 min 39.1 min (CLI alone) CLI x3; predicted walls 667/667/659/659/659/659 s; heaviest slice ≈ 12.2 min of tests
merge_group 37875522518 (48 packages) 34.0 min 33.3 min (CLI alone) CLI x3; predicted walls 591/583/583/581/581/583 s

The "after" half is a model reading. It cannot be measured on this PR's runs (see above), and it is what the seat's post-landing read replaces. One known bias, measured: the model is a lower bound on whole-package legs. On 37872770181 the five whole-package shards predicted 573/442/442/442/491 s and actually stepped 905/737/704/668/388 s.

The Test Core wall: held at 45, re-size owed after landing

timeout-minutes stays at 45. Removing the whole-CLI shard removes the reason for the old raise, but a wall is re-sized from a measured distribution. The post-change distribution does not exist until runs execute the new split, and a miss would kill merge_group runs for every lane. The rationale comment in ci.yml records the inputs for the re-size and states that none of them has been applied. Window: the 14 runs above, 37870616843 to 37876969409.

  • Input 1 (measured before the change). The job holding the whole CLI took 20.5-39.1 min (12 runs). Every other Test Core job took at most 20.8 min (72 jobs). The closure build took at most 5.6 min, and fixed setup at most 2.4 min.
  • Input 2 (predicted, not measured).
    • The heaviest CLI slice should take at most ~12.2 min of tests, or ~24 min of job. vitest's hash split puts 1.14x of an even third on slice 2/3.
    • The busiest whole-package shard of a full run will carry ~2640 s of windows, a size no run has executed. At the 2.0x packing measured on ~1400 s shards, that is at most ~30 min of job.
    • The wall model reads low on whole-package legs. On 37872770181 it predicted 573/442/442/442/491 s, and the shards took 905/737/704/668/388 s. So these figures are a floor for the post-landing reading, not a substitute for it.
  • The re-size. Size the wall from the slowest-job distribution in the post-landing read above, using the existing 80% margin rule. The revert condition is restated: back to 30 only when the slowest job stays at or under 24 min on every scheduled run for a week.
  • pnpm check:stall-guard-budget is green at 45: the cap is 20 min against a 45-min budget, leaving 25 min of slack.

--check-drift keeps its meaning

There is no code change on that path. Both the 1.5x red and the 1.3x warning keep their values. On affected-set runs the ratio is still one shard's executed windows against their own prediction: a shard carrying a CLI slice is predicted a third of the CLI's weight, from the slice count its own summary records (the OS_TEST_SHARD digest), not from the config. The drift batteries (9 + 12 cases) are unchanged and green. Slice skew, estimated from the 367-file list the CLI runs and the 18 slowest CLI files of run 37875522518, is 0.87/1.14/0.99 of an even third. A slice-carrying shard should therefore stay under the 1.3x warning.

What #16468 will read differently

  • The CLI's per-run test window now arrives as three slice windows in three shards' turbo summaries, each about a third of the whole, instead of one window on one shard. A per-shard ceiling check has two options. It can compare a slice against its share (the ceiling divided by the slice count the summary records, as --check-drift does through predictedSecondsFor). Or it can sum the parts across shards: report-test-timings.mjs already does this and marks a partial set.
  • The dataset still holds the whole cost per package, and the generator still reassembles slices within a run before recording them.
  • The shard step's timing changes on the three shards that carry a slice: they run the whole-package leg and then the slice leg in the same Run this shard's tests step.
  • Nothing of ci: a per-package suite-duration ratchet — a PR that makes a suite exceed its measured ceiling is red; ceilings rise only by ruling (maintainer-directed, growth constraint) #16468's ceiling check is built here.

Required contexts and coverage

  • No job, step or matrix name changed, so the seven required contexts are untouched.
  • Every affected package still runs. The CLI's three slices are vitest's --shard partition of its file list, so their union is the whole suite. check-test-completeness still grades every scheduled item, and its self-test is green.
  • scripts/test-shard-timings.json is not touched.

Gates (head 6e8b8f6; patch round 1, no origin/main merge this round)

  • node scripts/pm/dispatch-gates.mjs --commands --repo objectstack-ai/objectstack re-derived the same 58 commands on 6e8b8f6. All 58 ran, with exits captured before any pipe.
    • --ran reconciliation: 58 derived, 54 run, 4 NOT-MEASURED, 0 UNRUN.
    • The 4 NOT MEASURED are check:dts-closure, check:dual-build-cjs-loads, check:lean-entry-closure and check:sourcemap-no-sources-content. All exited 3 (PREREQUISITE NOT MET) because they load every package's built dist/, and this diff touches no package source or build config.
    • pnpm check:pm-dispatch-gates exited 1 on 1 of 2011 cases, the same as in round 0: "no mkdtempSync site in this tree takes a base the scan cannot read". It names packages/qa/dogfood/test/security-catalog-cold-boot-environment-holder.dogfood.test.ts:108. That file came from e030d43, which is already on main, and this diff does not touch it.
    • Every other command exited 0. That includes node scripts/partition-test-shards.mjs --self-test, pnpm check:stall-guard-budget (cap 20 min against a 45-min budget, 25 min of slack), pnpm check:stall-guard-headroom and pnpm check:required-contexts.
  • Round 0 (head 9156fd4) also ran these consumer self-tests, all green: measure-test-shard-timings.mjs, check-test-completeness.mjs and report-test-timings.mjs. They read partition-test-shards.mjs, which is byte-identical at 6e8b8f6.
  • Lint, a declared narrowing (round 0; the script is unchanged since): eslint --no-inline-config --format json scripts/partition-test-shards.mjs reported 1 file, 0 errors, 0 warnings.
    • The population comes from eslint's own config: the file is not ignored, and ci.yml is not an eslint input.
    • Computed parserOptions.project is null, so this diff cannot move any untouched file's verdict.
    • The full pnpm lint is CI's.

Ablation (run at 9156fd4; scripts/partition-test-shards.mjs is byte-identical at the current head; fix committed first; restored by git checkout HEAD, proven blob == HEAD and git diff HEAD empty)

Both ablations went through node scripts/ablation-replace.mjs in WRAP mode, which confirms the anchor hits 1 → 0 and records the blob change before running the command.

  1. Wall model reverted to bin sums. I replaced the wall formula's return line Math.max(serial, whole / concurrency) + sliced with whole + sliced (blob 739b9274f6d1 → 22c7fd4d7ccd). The self-test went RED: slice spread: the committed dataset, split as CI splits it, carries no slice -- @objectstack/cli: whole (whole fits: 1739s is within 2303s). That is the old reading, which kept the CLI whole. Restored to 739b9274f6d1.
  2. Per-run decision neutered. I replaced const needed = own > target; with const needed = false; (blob 739b9274f6d1 → 7d1f33503974). The self-test went RED: balance: at 6 shards the slowest predicted shard wall is 2.52x the mean (1739s vs 689s), past the 1.3x bound. Restored to 739b9274f6d1, and the self-test is green again on the restored tree.

Acceptance notes

  • test-nightly-tiers.yml's header still says FILE_SHARDED_PACKAGES "cuts the CLI into two vitest slices". That has been stale since the map emptied, and after this PR the CLI is configured at 3, which cannot spread on that workflow's 2 shards, so it runs whole there. Comment only, outside this card's file surface. Noted, not filed.
  • A seventh shard would lower the predicted maximum from 667 s to the 580 s floor. Past seven nothing moves. I did not do this: it changes the matrix, and it is not this card.
  • The wall model's whole-package term is a lower bound (the bias is measured above). If post-landing readings show whole-package shards consistently above the slice shards, the next lever is the slot weight given to a slice, not the bound.
  • The Test Core timeout-minutes re-size is the post-landing half of this card. Its inputs are in the ci.yml comment and in the section above.
  • skip-changeset: root scripts/ and workflows publish nothing.

Generated by Claude Code

claude added 3 commits October 9, 2026 03:48
…d slice the CLI per run

A shard runs its whole packages in one turbo run at --concurrency=4 and each
file-level slice in a leg of its own after it, so a bin's summed weight is not
its wall time. Graded on sums the split read 1.00-1.47x while the CLI shard ran
2.3-7.6x the other shards' mean. The bound (1.3x) now grades a wall model, the
CLI is configured at 3 slices (the smallest count that meets it on the
committed dataset), and each run slices it only when its own split needs it
and the slices spread.

Claude-Session: https://claude.ai/code/session_0115N1oNnQS5WqofZ2DzaT3q
Co-authored-by: Claude <noreply@anthropic.com>
…sured distribution

The whole CLI no longer runs alone on one shard (it is cut into 3 slices when a
run carries it), so the 45-minute wall raised for that shard is re-sized from
the 14-run window after the 21-run dataset refresh, with the window, the
numbers and the revert condition in the comment. The comments that described
the slice steps as idle are brought up to date.

Claude-Session: https://claude.ai/code/session_0115N1oNnQS5WqofZ2DzaT3q
Co-authored-by: Claude <noreply@anthropic.com>
…ed after landing

35 rested on a prediction of the post-change walls, and the wall model reads
low on whole-package legs. The wall stays at 45 and the comment keeps the
14-run window and the predicted figures as the input to the re-size the
post-landing reading owes, with the revert condition restated.

Claude-Session: https://claude.ai/code/session_0115N1oNnQS5WqofZ2DzaT3q
Co-authored-by: Claude <noreply@anthropic.com>
@objectstack-fleet

Copy link
Copy Markdown
Contributor Author

Lint & Repo Gates is red on 6e8b8f68e7, and the red is not this PR's · domain:devx seat 1 · session_0115N1oNnQS5WqofZ2DzaT3q · 2026-10-09T04:53Z


Generated by Claude Code

@objectstack-fleet objectstack-fleet Bot changed the title ci(test-shards): grade the Test Core split on predicted shard wall, slice the CLI per run, re-size the shard wall 45 -> 35 ci(test-shards): grade the Test Core split on predicted shard wall and slice the CLI per run Oct 9, 2026
@objectstack-fleet
objectstack-fleet Bot marked this pull request as ready for review October 9, 2026 05:57
@objectstack-fleet
objectstack-fleet Bot enabled auto-merge October 9, 2026 05:57
@objectstack-fleet
objectstack-fleet Bot added this pull request to the merge queue Oct 9, 2026
Merged via the queue into main with commit c64130b Oct 9, 2026
35 checks passed
@objectstack-fleet
objectstack-fleet Bot deleted the claude/issue-22075-affected-set-shard-balance branch October 9, 2026 06:32
os-sales pushed a commit that referenced this pull request Oct 9, 2026
… wall and slice the CLI per run (#22415)"

This reverts commit c64130b.

Why: PR #22415 packs about 2000-3000 s of predicted windows onto each
whole-package Test Core shard. The dataset's per-package windows were
measured on shards carrying about 1,400 s, and under 4-wide concurrency
each package's window inflates on the denser shards (objectql 1.84x,
plugin-security 1.80x, rest 1.61x, plugin-auth 1.87x). That carries the
shard past the "Check this shard's timing drift" step's 1.5x red while
the test step itself passes.

Red full runs on that partition model:
- main push 37894048074 (e02833c; shards 4/6 and 6/6; 6/6 read
  3992.4s measured vs 2467.1s predicted = 1.62x)
- merge_group builds 37894050587, 37893672824, 37894053453 and
  37892033675

The revert restores .github/workflows/ci.yml and
scripts/partition-test-shards.mjs byte-for-byte to their content at
f05649f, the reverted commit's single parent. A contention-aware
redesign is the next round, not a hotfix.

Claude-Session: https://claude.ai/code/session_0115N1oNnQS5WqofZ2DzaT3q
Co-authored-by: Claude <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci/cd size/l skip-changeset PR has no user-facing published change; bypasses the changeset gate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants