Repository navigation
chore(ci): refresh the Test Core shard-timings dataset - #20388
Conversation
Regenerated by .github/workflows/shard-timings-refresh.yml from the test-core-run-summary artifacts of 1 accumulated run(s) (36380128221), newest 36380128221 at eee0974. Generated, never hand-edited.
|
Claim: PM loop round 1 Why a seat claims a bot PR. The maintainer approved this PR and armed auto-merge on it (2026-09-30T05:51Z, How it lands. The dev opens a separate draft PR from the branch above that re-derives the pins so the self-test passes on both datasets: the one on Stamp: 2026-10-02T22:01Z · read against Generated by Claude Code |
|
os-dev-report Generated by Claude Code |
|
os-dev-report Generated by Claude Code |
ACCEPT — PR #21487 (head
|
…'s own slice-count derivation (objectstack-ai#21487) Unblocks objectstack-ai#20388 (the Test Core shard-timings refresh), which objectstack-ai#16465 and objectstack-ai#16468 wait on. This PR binds no card: objectstack-ai#20388, objectstack-ai#16465 and objectstack-ai#16468 all remain open. objectstack-ai#20388 lands by itself after this merges, on its own green, through the auto-merge the maintainer armed. ## What this changes `scripts/partition-test-shards.mjs --self-test` now passes on BOTH datasets: the `scripts/test-shard-timings.json` on `main`, and objectstack-ai#20388's (blob `84342b45`, byte-identical to its head `8d60ae86`, overlaid unstaged in a second worktree and never committed). - **`FILE_SHARDED_PACKAGES` is emptied** (the CLI's `2` is retired). The mechanism stays: the item grammar, `expandSlices`, the vitest file-count floor, the `OS_TEST_SHARD` wiring judge, the generator's slice reassembly, and the dormant CLI wiring in `turbo.json` and `packages/cli/vitest.config.ts`. - **Pin 3c is rewritten.** It no longer substitutes the hard-coded `CLI_MEASURED = 1231.52` bridge. It asks the committed dataset, for every map entry, whether `n - 1` slices would also meet the bound (`sliceCountProblems`), and refuses an `n` that `n - 1` could replace. Five fixtures hold each refusal in both directions: needed, retire, lower, unmeasured entry, and an entry below 2. - **Pin 3b is rewritten.** It grades the real split as before. It also grades a cut of the dataset's heaviest package in two, so the spread check is not vacuous while the map slices nothing. - **The mechanism's pins read a fixture map** through new optional `sliced` parameters (default: the live map) on `sliceCountFor`, `weighItems` and `driftReport`. The re-pointed cases are the vitest floor, the slice-share prediction, drift on a sliced package, observed-run-wins, `weighItems` slicing, and the wiring reader pointed at the CLI's real tree. - **Pin 6's inversion pair is replaced.** `example-todo` over `core` flipped in the refresh (35.30s vs 57.09s). The new pair is `plugin-pinyin-search` over `sdui-parser`: 2 vs 13 test files, 14.40s vs 1.62s on `main`'s dataset, 38.77s vs 4.54s on objectstack-ai#20388's. A new guard reads the dataset's own order, so a flipped pair asks for a new pair instead of reporting a weighing defect. - **`PREVIOUS_FILE_SHARDED_PACKAGES` (new, beside the map) holds the outgoing map, `{ '@objectstack/cli': 2 }`.** The generator's slice-digest matcher decodes a digest against both declared maps, exactly by hash, so a run summary written before this change still reads. Without it, the refresh lane refused every run on `main`; see the next section. A decoded slice is still summed within its run (an incomplete set contributes nothing), and a digest neither map names is still refused. The map has no reader once a day has passed, because run summaries are kept for one day (`retention-days: 1`). - **`scripts/measure-test-shard-timings.mjs` (outside the claimed surface, and why).** This is the matcher change above. Its live-map case also read `FILE_SHARDED_PACKAGES['@objectstack/cli']`, which reds under an empty map (measured: the unedited file exits 1). That case now covers both live maps, and a new case reads a pre-change run's two halves into their 733.33s whole while still refusing a count neither map names. The env-carried battery floor goes from 10 to 11. A stale "(2 for @objectstack/cli today)" parenthetical is dropped. Battery floors are raised to the measured counts: balancing pins 21 to 25, OS_TEST_SHARD wiring 9 to 11. File-level slice items stays at 20. The counts are the same on both datasets. ## The three bin readings, and the keep-or-retire choice Predicted bins at 6 shards against the 1.3x bound, using the partitioner's own `partition` / `balanceOf` / `expandSlices`: | dataset | case | bins (s) | max/mean | heaviest item | verdict | |:--|:--|:--|:--|:--|:--| | objectstack-ai#20388 (72 pkgs, 7430.00s) | **K** CLI sliced at 2 | 1391/1208/1208/1207/1208/1208 | 1.124x | spec 1391s | meets | | objectstack-ai#20388 | **R** CLI whole (733.33s) | 1391/1208/1206/1208/1208/1208 | 1.124x | spec 1391s | meets | | objectstack-ai#20388 | **R-worst** CLI whole at 1231.52s | 1391/1308/1307/1307/1308/1307 | 1.053x | spec 1391s | meets | | `main` (70 pkgs, 3997.47s) | K | 666/666/666/667/666/666 | 1.000x | spec 404s | meets | | `main` | R | 666/666/667/667/667/666 | 1.001x | cli 458s | meets | | `main` | R-worst (the old bridge) | 1232/708/707/709/708/707 | 1.549x | cli 1232s | breach | On objectstack-ai#20388's dataset, both R and R-worst meet the bound, and slicing changes no bin's maximum (spec is the heaviest item in all three cases). Solving C ≤ (1.3/6)(6696.67 + C), the CLI fits whole until about 1852s, which is 1.5x its worst reading. The file's own rule makes n the smallest count that meets the bound. That count is 1, so the entry is retired. Four axes: - **Real need:** measured. Slicing buys nothing in the predicted maximum on the refreshed dataset. It costs a duplicated closure build and a sequential slice leg on two shards. In objectstack-ai#20388's run 36380128221, "Build the sliced package's dependency closure" took 6m15s on `Test Core (5/6)` and 6m43s on `(6/6)`, and 5/6 was the longest job at 19m23s. - **Long-term:** keeping `2` would leave a count no pin can justify. The old counterfactual cannot be rewritten on any measured basis, because R-worst meets. Retirement keeps the mechanism, and pin 3 (the floor) is the live trigger to slice again. - **AI-error resistance:** the new 3c reds on a slice count that is larger than needed. A future author cannot leave slicing configured without a measurement behind it. - **Startup focus:** fewer live moving parts and no new gate. 3c replaces a case inside the existing self-test, and the retirement is immediate. ## The refresh lane across the change: red, then green - **Red at `c614a094`.** This PR's `pull_request` rehearsal of `shard-timings-refresh.yml` (run 37074888579, job 111062437870) refused all ten eligible hourly runs on `main`, 37010481060 through 37070188866. It then ended with "The 0 eligible run(s) ... NOTHING was regenerated". - **Cause, reproduced locally** (the run's own log and artifacts are behind the egress policy here). A synthetic run was built from `main`'s dataset in the exact turbo 2.10.10 shape: six whole-leg summaries, plus two slice legs whose `OS_TEST_SHARD` digest is the sha256 of `1/2` or `2/2` (`d939926f…`, the measured format). It was fed to the lane's own two commands: the generator with `--run` / `--merge-into`, then `select-shard-timings-run.mjs --check-coverage`. - `aa463223` (the map `{cli: 2}`): exit 0, coverage OK. - `c614a094` (the empty map): exit 1, "matches no slice the partitioner can emit for it (none: FILE_SHARDED_PACKAGES does not slice it)". - `d049d353`: exit 0, coverage OK, and a `packages` map identical to `aa463223`'s. - **So the hypothesis holds, and its rival is falsified.** The 2026-09-30 green rehearsals predate the env carrier, so they cannot separate "the empty map refuses these slices" from "the env carrier never decoded on real summaries". The rehearsal at `d049d353` does: run 37075730446 (job 111065090886) is **green**. It accepted its first candidate, with no "dropping that run" warning, so `main`'s real env-carried summaries decode once the matcher knows the outgoing map. - **Why decode the outgoing map, not keep the 2 slices.** Read along the four axes: - Need: every run summary that exists today carries the CLI's two slices. That is the lane's only input until a day after the merge. - Long-term: the matcher stays a whitelist. It reads two declared maps by hash equality, never a guessed count, and any future map change in either direction gets the same one-day bridge by setting one constant in the same PR. - AI error: the refusal is unchanged for any digest neither map names (N4 below), and completeness is still judged within a run. - Startup: no new gate, and no grace window beyond the artifacts' own one-day life. - Keeping 2 would bring back a slice count that pin 3c refuses on both datasets, with no measured counterfactual to replace it. - **The next scheduled weekly refresh (Mon 2026-10-05 05:30Z), if this has merged by then: expected to succeed.** With one-day retention, its candidates are runs made under the new map: whole CLI, no slice digest. Measured on synthetic summaries at `d049d353`: a post-change run alone gives exit 0 (cli whole, 740s). A post-change run accumulated with a pre-change sliced run gives exit 0 (cli 599.08s, the median of 740 and 458.15). So a refresh that straddles the merge also reads. It is still a prediction: it depends on an eligible full-battery run existing, as every refresh does. ## Self-test on both datasets (head `d049d353`) ``` BEFORE main dataset (aa46322) exit 0 partition-test-shards: self-test OK (70 measured packages -> 71 shard items, 6 shards, max/mean 1.00x <= 1.3x, floor 404s, bins 666/666/666/667/666/666s) BEFORE objectstack-ai#20388 dataset (aa46322) exit 1 Error: slice derivation: UNSLICED, @objectstack/cli at 1231.52s is 1391s against a 1321s mean and now fits under 1.3x on its own ... Re-derive the slice count (or retire it) ... AFTER main dataset (d049d35) exit 0 partition-test-shards: self-test OK (70 measured packages -> 70 shard items, 6 shards, max/mean 1.00x <= 1.3x, floor 458s, bins 666/666/667/667/667/666s, file-level slices: none) AFTER objectstack-ai#20388 dataset (d049d35) exit 0 partition-test-shards: self-test OK (72 measured packages -> 72 shard items, 6 shards, max/mean 1.12x <= 1.3x, floor 1391s, bins 1391/1208/1206/1208/1208/1208s, file-level slices: none) ``` The BEFORE red reproduces objectstack-ai#20388's AFTER block on today's `main` (line 1075 now, 967 in that body). It also mislabels spec's 1391s as the CLI's, because the old counterfactual read the overall heaviest item. A second red was masked behind it. With the 3c throw muted on `aa463223`, pin 6 reds on objectstack-ai#20388's dataset with "The run is weighing test-file count again". That is a misdiagnosis: the pair had flipped. ## Mutation proof, one per rewritten or added pin Every leg was taken from committed `c614a094` through `scripts/ablation-replace.mjs`. The whole table plus N1 to N4 was then re-run from committed `d049d353` on both datasets: 48 legs, every one landed and restored. In each leg the anchor hit 1 and the replacement went 0 to 1. Each leg was restored to the HEAD blob with `git diff HEAD` empty, under a shell `trap` that re-proved the hashes in both worktrees at the end. The red line below is the self-test's first `Error:`, the same on both datasets unless shown. | # | mutation | first red | |:--|:--|:--| | M1 | re-add `'@objectstack/cli': 2` to the live map | `slice derivation, committed dataset (1 of 1 ...)`, "sliced 2 ways at 458.15s (main) / 733.33s (objectstack-ai#20388), but at 1 the split already meets 1.3x ... Retire the entry" | | M2 | `meetsBound` always meets | `a count of 2 that 1 cannot replace was refused` | | M3 | n-1 = 1 never judged | `a package that fits whole kept its slicing with no refusal` | | M4 | only n-1 = 1 judged | `a count of 3 where 2 meets the bound was accepted` | | M5 | unmeasured entry not refused | `an entry the dataset never measured was accepted` | | M6 | the "at least 2" floor dropped | `an entry of fewer than 2 slices was accepted` | | M7 | `expandSlices` stops slicing | `slice spread: cutting @objectstack/cli (main) / @objectstack/spec (objectstack-ai#20388) in two produced no ... pair to grade` | | M8 | pin 6 back to the old pair | **green on `main`** (the pair still holds there); on objectstack-ai#20388: `the dataset no longer measures @objectstack/example-todo slower than @objectstack/core (35.3s vs 57.09s) ... Pick a new inversion pair` | | M9 | `weighItems` weighs test-file count | `weight: @objectstack/plugin-pinyin-search weighed 2 and @objectstack/sdui-parser weighed 13 ...` | | M10 | every package takes the map's first count | `slice count: a package outside the slice map was sliced` | | M11 | `sliceCountFor` ignores `sliced` | `slice count: the configured package did not read its configured count` | | M12 | vitest floor refusal dropped | `slice floor: slicing below the test-file count was not refused (no throw)` | | M13 | floor refuses below 1000 files | the floor's own refusal for 500 files, thrown from the "plenty of test files" case | | M14 | prediction stops dividing | `prediction: a sliced package was charged its WHOLE dataset entry` | | M15 | `driftReport` ignores `sliced` | `drift: a sliced overshoot read 0.83x and was not reported as drift` | | M16 | observed slices ignored | `drift: an observed WHOLE run was charged a slice-sized prediction (1.67x)` | | M17 | `weighItems` ignores `sliced` | `weighItems: 1 package(s) produced 1 item(s), expected 2` | | M18 | wiring reader's default is not the live map | `slice wiring: judged 1 of 0 sliced package(s) ...` | | M19 | `packages/cli/vitest.config.ts` reads `OS_TEST_SHARD_UNREAD` | `slice wiring, the tree read through a fixture map ... never reads OS_TEST_SHARD into vitest's shard` | | M20 | measure: `samplesFromSummary`'s default `sliced` is not the live map | `env slice: a @objectstack/cli digest of 1/3, a count neither live map names, was not refused listing (1/2, 2/2) (got "no throw")` | | N1 | the matcher stops reading the previous map | the live-maps case: `@objectstack/cli ran with OS_TEST_SHARD set ... matches no slice the partitioner can emit for it, or emitted under the map it replaced (none: ...)` | | N2 | `samplesFromSummary`'s default `previous` is not the live one | the same refusal, from the live-maps case | | N3 | `samplesFromSummary` drops its `previous` argument | the across-change case: `a: cli ran with OS_TEST_SHARD set (digest d939926f…) ... matches no slice ...` | | N4 | the matcher decodes a count neither map names | `env slice: a @objectstack/cli digest of 1/3, a count neither live map names, was not refused listing (1/2, 2/2)` | ## Premise checks (the dispatch's A1 to A4) - **A1** holds on `aa463223`: the map was `{ '@objectstack/cli': 2 }`, pin 3b carried the 1231.52s/800.7s prose, and pin 3c carried `CLI_MEASURED = 1231.52`. - **A2** reproduced on current `main`; see BEFORE above. - **A3** recomputed. On `main`: 70 packages, 3997.47s, cli 458.15s, spec 403.65s. On objectstack-ai#20388: 72 packages, 7430.00s, `measuredAt` 2026-09-28, runs `[36380128221]`, spec 1391.38s (heaviest), cli 733.33s. The two new packages are `organizations` and `vitest-filter-preflight`. The package set equals today's workspace minus the CI-excluded `dogfood`. - **A4: well-formed.** Today's generator writes the identical shape: the same top-level and `provenance` keys, and byte-identical `note`, `mergeRule` and `refresh` strings. It accepts the file as a `--merge-into` target (carried `spec` on a HIT witness). `skippedAsCached`, `skippedIncompleteSlices` and `carriedOver` are all empty, and the CLI's 733.33s is a two-slice sum from one run. The 09-28 run carried its slices as a `--shard` passthrough, which today's generator still reads (`sliceOfCliArguments`). Its run-summary artifacts are no longer retained (the run lists only `test-core-timing-table` and `build-output`), so a byte-level re-generation is NOT MEASURED. ## Consequences outside this diff (known, deliberate) - **The refresh lane.** See its own section above: red at `c614a094`, green at `d049d353`. - **Nightly tiers.** `test-nightly-tiers.yml` partitions the one tier-owning package at 2 shards. Measured locally: shard 1/2 now carries `@objectstack/cli` whole, and 2/2 carries nothing ("No packages on this shard", exit 0). The run stays inside its 45-minute timeout by its own header's estimate. Its header prose and its 2-shard matrix are now stale; no carrier. - **The window between this merge and objectstack-ai#20388's.** `main` then splits on the stale weights with the CLI whole. Modeled with objectstack-ai#20388's weights standing in for real cost: the heaviest actual bin stays spec's (1775s sliced, 1773s whole). The CLI's bin is 1116s at 733.33s, or 1614s at 1231.52s. Merging `main` into objectstack-ai#20388 right after this lands keeps the window short. ## Gates (head `d049d353`) `node scripts/pm/dispatch-gates.mjs --commands --repo objectstack-ai/objectstack` (change set from git, 2 paths) derived 32 families. The list is the same on `c614a094` and `d049d353`. On both heads, all 32 ran with exit 0, recorded and reconciled by `--ran`: "32 derived, 32 run, 0 NOT-MEASURED, 0 UNRUN". `node scripts/report-test-timings.mjs --self-test` also ran, exit 0. The list includes the three self-tests (partition / measure / select-shard-timings-run), `check:cross-package-test-inputs`, `check:nul-bytes`, `check:pm-dispatch-gates` (1976 cases, 977s), `check-scripts-symbol-anchors` and `check-self-test-wired`. Lint, narrowed and proven. `eslint --no-inline-config --format json` on the 2 changed files reports 2 files, 0 errors and 0 warnings on both heads, and neither file is ignored. `--print-config` shows 2 per-file rules (`no-restricted-imports`, `comment-swallow/no-code-inside-block-comment`) and no `parserOptions.project` or `projectService`. Type-aware linting is off, so this diff cannot move any untouched file's verdict. The full `pnpm lint` is CI's. `skip-changeset`: root `scripts/` ships in no package's `files[]` (the root package is private). ## Acceptance notes - **Observation, not filed:** spec's 1391.38s is larger than every shard's measured test-step wall in the same run (max 1023s). That is consistent with the generator's documented fold, which sums the `test` and `test:repo` windows, and the windows overlapped. Spec is now the floor, at about 83% of its own breach point (about 1670s). objectstack-ai#16468's ceilings will read this number. - **Observed, not filed:** when the generator refuses every candidate, the refresh lane's final `::error::` blames a coverage shortfall ("a suite failed, a package was renamed or removed, or its slices could not be assembled") rather than the refusals, as run 37074888579 showed. No public door; no carrier. - **Stale prose outside this diff after the retirement.** The `ci.yml` slice-leg comments and the 1231.52s note at its drift step are for objectstack-ai#16465, which edits that file. `packages/cli/vitest.config.ts`'s "partition-test-shards.mjs slices this package" has no carrier. The `test-nightly-tiers.yml` header and matrix have no carrier. `shard-timings-refresh.yml`'s "pin 3c ... the day the CLI comes back under the bound" is still true in mechanism; no carrier. --- _Generated by [Claude Code](https://claude.ai/code/session_01HRYqpqGcWpJuJkDmbRF75w)_ --------- Co-authored-by: Claude <noreply@anthropic.com>
Refreshes
scripts/test-shard-timings.json, the balancing input for the Test Coreshard split. Opened automatically by
.github/workflows/shard-timings-refresh.yml.Every byte came out of
scripts/measure-test-shard-timings.mjs; nothing here washand-edited, and no bound, timeout or matrix entry was touched.
Source
Measured across 1 accumulated run(s) of the HOURLY
schedulerun of CI onmain— the full-battery run (#16467). Apushrun onmainis affected-only andis not a measurement of the workspace, so no push run feeds this file.
No single green run measures the whole workspace either — turbo's cache is namespaced
per shard and only main pushes write it, so a
package whose inputs have not changed is a HIT and the generator refuses hits rather
than recording a replay as a duration. Runs are therefore accumulated, each fenced by
its own
--rungroup, until every package the committed dataset holds is measuredagain; a package seen in several of them gets the median of those observations.
https://github.com/objectstack-ai/objectstack/actions/runs/36380128221
Newest run in the set:
36380128221, commiteee0974236c87a7232f9805a19b6e53e3e36da59— the date this refresh carries.Every run above had all six
Test Core (N/6)jobs concludesuccesswith its sixrun-summary artifacts still retained; runs that were cancelled, failed or had lost
their artifacts were rejected by name in the log before any of these were used.
All 72 package weights were measured in these runs; nothing was carried.
because a weekly lane cannot know which cards a given run ought to retire. If this is
the first refresh to land, retire those two by hand as part of merging it.
Measured per-shard suite time on the newest run in the set
Predicted bins, before and after
The partitioner's own pins RED on this refresh — read this before merging
This is the designed behaviour, not a defect in the refresh: the acceptance bound is
a ratio, and a package that has grown past what any six-way split can bin makes the
pins fail with the arithmetic in the message. The remedy the partitioner names is to
raise the file-level slice count for that package — ⛔ never to raise the bound, and
⛔ never to hand-edit this dataset. This workflow deliberately does neither: it
reports and stops, because both are decisions.
No checks will start on this PR by themselves
It was opened with the Actions
GITHUB_TOKEN, and GitHub's recursion guard means aPR opened that way triggers no workflow runs. Push any commit to the branch, or close
and reopen the PR, to start CI.
Refs #16464, #16173, #16222.