Skip to content

perf(gc): cut the copying minor's per-live-object trace cost (straight-line object scan, hoisted per-object facts, exact memos) - #11676

Merged
proggeramlug merged 8 commits into
mainfrom
perf/11549-trace-cost
Sep 30, 2026
Merged

proggeramlug merged 8 commits into
mainfrom
perf/11549-trace-cost

Conversation

@proggeramlug

@proggeramlug proggeramlug commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Part of #11549

This cuts the copying minor's per-live-object trace cost by about a third, and it does not change when collections run. On a fully-live 5 MB binary tree (131,076 objects promoted in place, btree n=20), the minor goes from 1,738 to 1,219 instructions per live object. Leaving out the first-cycle barrier-arming walk (#11668's subject, unchanged here), it goes from 1,472 to 953 (−35%). That is callgrind over gc_collect_minor_with_trigger_inner only: 227.8 M → 159.8 M.

It does not rescue #11645's small-tree rows. See "#11645 stacked on this" below.

What the profile showed (Phase 1)

I attributed every instruction of the minor to the innermost inlined perry-runtime function (callgrind --dump-instr plus addr2line -i), then grouped the functions by step. Per promoted object, btree n=20, main b0bf0ae7e vs this branch. The top 70 functions cover about 96% of the minor. Shared page-cache and TLS lookups are listed separately.

step main this PR
barrier-arming reconstruct walk (first minor only, #11668) 196 201
in-place promotion stamping and old-page index 123 135
worklist drain, mark bits, clear marks 113 114
per-slot decode and classification (tag, memo, page lookup, header plausibility) 252 195
remembering predicate per slot (the slot's own generation) 72 11
slot enumeration (rewrite walk, child-slot iterator, selection, shape mask, per-object mask, carrier note) 434 260
per-slot visit body and closure/call overhead 243 136
weak-holder fact 7 6
layout-read counters 15 6
page-cache and TLS lookups 206 109
other 12 9
total (top 70) 1,673 1,182

How that compares with the hypotheses:

  • (1) Per-type fast paths: yes, the biggest win. The cost was plumbing, not classification. Per slot, the path was two indirect calls (the descriptor visitor and the slot visitor) plus a 5.7 KB generic walk rebuilding a HeapChildSlotIterator per object. There was a SHAPE_LAYOUTS hash probe and mask clone per object, a per-object-mask filter, and a generation probe per object for the carrier note.
  • (2) Header bit for side-table entries: not feasible. gc_flags is fully used and _reserved is type-overloaded (bit 6 is OBJ_FLAG_NULL_PROTO on objects, which is why the residual-prototype registry already has no per-object bit for them). The side-table lookups are cheaper here than expected: per-object layout masks already have an address filter (26/object), and overflow fields hit an empty map. I did not add a bit.
  • (3) Cheap young-range check: not applicable. The nursery is not one range: blocks are registered in a page map. The avoidable part was classifications nobody used. Every slot paid a page-map probe of its own address for barrier_parent_needs_remembering, and on a promoting cycle skip_remembering makes that answer irrelevant (72 → 11). The raw keys word of every shaped receiver was also re-validated.
  • (4) Hoisting: yes (the remembering fact, the memo, the carrier note).
  • (5) O(tables) work: real, but not per live object, and not in this PR. On qs minors, whole-table passes take a large share. The closure box-capture prune is 16% of qs_parse minor instructions. The per-object layout owner prune is 12%, including a sort of the young-key log. The remembered-set rebuild is 12%. These scale with dead or young records, not with the live objects traced. dotenv's full collections are dominated by the valid-pointer-set build, the sweep and per-dead-object layout clears, not by tracing.

What changed

  1. Ordinary objects take a straight-line scan (gc/copying_object_scan.rs). For GC_TYPE_OBJECT, the drain builds the generic walk's slot sequence as a plan, in the same order, and visits it directly:

    • the residual-prototype edge;
    • the keys edge, with the old-carrier note;
    • the meta record;
    • the payload slots;
    • the overflow fields.

    The payload selection is heap_payload_slot_selection_from itself, so there is one copy of the mask logic. The scan declines only up front, before any side effect: during a full trace or a layout-scan trace, or if the type table no longer describes objects this way. After that it never hands an object back, so no edge is visited twice.

    Drift check: in test and debug-assertion builds, every object it scans is also enumerated by the generic walk, and the slot lists must match.

  2. The parent's remembering question is asked once per object (ParentRemembering: Never / Always / ExternalSlotsOnly). barrier_parent_needs_remembering(parent, external) is Old(parent) || (external && malloc_parent(parent)). Both parent terms are fixed while one object's slots are visited, so the slot is classified only for a malloc parent. skip_remembering folds into Never.

  3. Raw words consult the mark memo before classifying. The memo only holds an address that classified successfully this cycle. That classification cannot change within the cycle: the range stays registered until the from-space reset, and forwarding does not touch the header fields that plausibility reads. Test and debug builds re-derive this on every hit.

  4. SHAPE_LAYOUTS memoizes its last pointer-mask answer, exactly. The map is reachable mutably only through DerefMut, which clears the memo first. Every write path (insert, poison, entry, iter_mut) goes through it.

  5. The old-carrier note skips its generation probe when it would change nothing (both record flags are already set).

  6. The drain's slot walk is instantiated for its visitor (visit_gc_rewrite_slots_inline), and its classification is inlined. Every other caller keeps the shared dyn walk. gc_child_slots and the selection are inlined into both instantiations, so the full-trace path does not get slower.

Gate-maintenance edits:

  • gc_runtime_root_holders.json: re-audited the PASS1_MARKED pin for the one-line gc/mod.rs change, and pinned one #[cfg(test)] counter.
  • shape_descriptor_census.py: now reads the walk's one body and the plan, and has a new sabotage case.
  • SHAPE_LAYOUTS is still declared with its map type, so registry_lifetime_check.py still sees it.

Sabotage tests (each has a twin that must fail)

  • gc::tests::copying_object_scan:
    • the plain-object path is taken, and moves every child;
    • it still scans with the residual-prototype latch armed;
    • a plan that drops the top payload slot is refused by the cross-check;
    • a promoted receiver still notes its shape as old-carried, and claiming every shape already noted loses the note.
  • gc::tests::copy_slot_hoists: ParentRemembering agrees with barrier_parent_needs_remembering for old, young and malloc parents, and a forgetting fact disagrees.
  • gc::tests::shape_layout_table: the memo matches the map across insert, poison, remove and entry; a memo kept across a write answers stale.

Measurements

Setup:

  • perrymaster (Linux x86_64 Zen 4), --release, PERRY_NO_AUTO_OPTIMIZE=1, under /tmp/perry-bench-lock.d.
  • A/B base is main b0bf0ae7e. This branch was then rebased onto 74fed830f with no code change; see "Not run".
  • Instructions: perf stat -e instructions:u, median of 3.
  • Peak RSS: /usr/bin/time %M, median of 11, arms interleaved, THP on (host default) and off (MIMALLOC_ALLOW_THP=0).
  • Fixed-size probes report the total; packages run at their default sizes.
workload main instr this PR instr Δ instr Δ RSS THP on Δ RSS THP off RSS THP off KB (main → this PR) GCs minor/full (main ; this PR)
btree@3 24.92 M 24.92 M +0.00% -1.8% +0.9% 15,332 → 15,476 0/0 ; 0/0
btree@6 28.65 M 28.65 M +0.00% -1.8% -0.0% 16,440 → 16,436 0/0 ; 0/0
btree@10 33.64 M 33.64 M +0.00% -1.8% -0.3% 17,692 → 17,632 0/0 ; 0/0
btree@20 276.92 M 208.90 M -24.56% +0.1% +0.2% 30,372 → 30,436 1/0 ; 1/0
btree@40 301.82 M 233.81 M -22.54% +0.2% +0.0% 34,140 → 34,144 1/0 ; 1/0
01_nursery_churn 188.95 M 188.70 M -0.13% +0.3% +0.7% 24,028 → 24,204 1/1 ; 1/1
02_survivor_promotion 311.11 M 286.67 M -7.86% +10.4% +0.0% 25,024 → 25,028 1/1 ; 1/1
12_large_live_set 3600.54 M 3289.08 M -8.65% -5.2% -0.1% 105,480 → 105,380 5/1 ; 5/1
alloc@8000000 3518.79 M 3518.74 M -0.00% -18.5% +0.4% 21,232 → 21,308 67/0 ; 67/0
bare@8000000 113.09 M 113.08 M -0.00% +0.3% +0.9% 7,828 → 7,896 0/0 ; 0/0
retain@2000000 6308.68 M 5997.10 M -4.94% +5.2% +0.2% 110,536 → 110,800 9/1 ; 9/1
dotenv_parse 29.21 G 29.22 G +0.04% -2.7% +0.6% 31,776 → 31,980 0/7 ; 0/7
moment_parse_format 24.59 G 24.60 G +0.04% +1.2% +7.0% 40,652 → 43,504 1/33 ; 1/33
date-fns_format_add 19.78 G 19.79 G +0.04% -0.0% +0.3% 37,116 → 37,228 17/0 ; 17/0
validator_batch 104.13 G 102.53 G -1.54% +0.1% +0.0% 60,060 → 60,064 46/2 ; 46/2
qs_parse_nested 41.05 G 40.85 G -0.48% -0.4% -0.4% 199,448 → 198,692 11/0 ; 11/0
qs_stringify_nested 127.75 G 124.78 G -2.32% -1.6% -0.1% 502,112 → 501,700 24/10 ; 24/10
uuid_v4 6858.05 M 6852.51 M -0.08% -0.1% +0.1% 30,076 → 30,116 1/0 ; 1/0
jsonwebtoken_decode 1720.59 M 1717.68 M -0.17% +0.0% +0.2% 58,948 → 59,076 2/0 ; 2/0

Notes on the rows that are not better:

  • dotenv +0.04%, deterministic. Split with callgrind: the GC functions are +0.08 M (neutral). The rest is code outside the GC being inlined differently: the catch_js_throw instances in the regex paths, value_is_callable, js_try_end. A no-op control build (main plus 4 KiB of unused rodata) moves dotenv by −0.17 M, so the shift comes from the code change, not from the build. I could not steer it.
  • date-fns +0.04% (perf) is noise. Callgrind with the same runtime and generic user code gives −12.2 M (−0.06%).
  • moment is nondeterministic. Its GC schedule and instruction count vary from run to run: callgrind gives 7.408–7.470 G for main and 7.411–7.438 G for this branch at 3000 iterations. Its THP-off peak RSS is bimodal (38–41 MB vs 43–44 MB), depending on which schedule a run takes. Over 31 runs without PERRY_GC_DIAG, main hit the high mode 13 times and this branch 22 times. Over 30 runs with PERRY_GC_DIAG, main hit it 4 times and this branch once. I don't see a direction attributable to this change, but I have not proven there isn't one.
  • THP-on RSS rows (02 +10%, retain +5%, alloc −18%) are multimodal on this host: 02 ranges from 25 to 56 MB in both arms. The THP-off column is the stable one, and there every row is within ±1% except moment.

Stress and correctness

  • Seeded GC-schedule sweeps: 200 seeds each, PERRY_GC_PROTECT_FROMSPACE=1, [gc-fromspace-protect] retired_set checked armed on every run (unarmed = 0), output compared to Node 26.5.1.
  • RUST_TEST_THREADS=1 cargo test --release -p perry-runtime (codegen-units 16) on the final rebased head: 4804 passed, 0 failed, 5 ignored.
  • Gap A/B vs pristine main, each arm with its own release build (PERRY_SKIP_BUILD=1 PERRY_NO_AUTO_OPTIMIZE=1, Node 26.5.1), identical to main on every filter:
  • SKIP_COMPILE_GATES=1 scripts/run_lint_gates.sh: 105 of 107 script gates pass. The two failures are the known ones: public-baseline freshness, and cargo xwin missing on this host. The compile tier was not run. git diff --stat was clean afterwards. cargo fmt --check, check_file_size.sh, gc_runtime_root_holders.py and thread_exit_address_globals.py all pass.
  • No new env knob.

#11645 stacked on this

I built #11645's seven commits (#11668 + #11612 + the pacing) on top of this branch and measured them against the same main:

workload main instr #11645 stacked instr Δ instr Δ RSS THP on Δ RSS THP off RSS THP off KB (main → #11645 stacked) GCs minor/full (main ; #11645 stacked)
btree@3 24.92 M 142.97 M +473.84% +61.2% +65.7% 15,332 → 25,412 0/0 ; 1/0
btree@6 28.65 M 146.71 M +412.07% +56.9% +59.4% 16,440 → 26,204 0/0 ; 1/0
btree@10 33.64 M 151.69 M +350.89% +51.8% +54.7% 17,692 → 27,376 0/0 ; 1/0
btree@20 276.92 M 164.23 M -40.69% -8.0% -8.7% 30,372 → 27,720 1/0 ; 2/0
btree@40 301.82 M 189.27 M -37.29% -18.1% -18.6% 34,140 → 27,796 1/0 ; 4/0
01_nursery_churn 188.95 M 146.35 M -22.55% -26.0% -24.7% 24,028 → 18,084 1/1 ; 4/1
02_survivor_promotion 311.11 M 302.33 M -2.82% +7.9% -11.5% 25,024 → 22,140 1/1 ; 2/1
12_large_live_set 3600.54 M 3092.69 M -14.10% -8.6% +1.6% 105,480 → 107,220 5/1 ; 6/1
alloc@8000000 3518.79 M 3498.33 M -0.58% -8.3% -18.8% 21,232 → 17,244 67/0 ; 203/0
bare@8000000 113.09 M 113.08 M -0.00% +0.0% +0.6% 7,828 → 7,872 0/0 ; 0/0
retain@2000000 6308.68 M 5466.29 M -13.35% +4.7% -2.3% 110,536 → 108,048 9/1 ; 9/1
dotenv_parse 29.21 G 20.47 G -29.92% -12.7% -19.8% 31,776 → 25,492 0/7 ; 61/0
moment_parse_format 24.59 G 20.46 G -16.78% +0.9% -13.3% 40,652 → 35,232 1/33 ; 43/0
date-fns_format_add 19.78 G 19.56 G -1.13% -17.8% -29.1% 37,116 → 26,320 17/0 ; 71/0
validator_batch 104.13 G 97.61 G -6.26% -15.4% -25.3% 60,060 → 44,836 46/2 ; 211/0
qs_parse_nested 41.05 G 41.07 G +0.06% +0.8% -1.0% 199,448 → 197,496 11/0 ; 14/0
qs_stringify_nested 127.75 G 123.97 G -2.96% -1.8% -2.8% 502,112 → 488,192 24/10 ; 24/10
uuid_v4 6858.05 M 6834.14 M -0.35% -10.4% -20.9% 30,076 → 23,776 1/0 ; 6/0
jsonwebtoken_decode 1720.59 M 1721.60 M +0.06% -16.9% -21.4% 58,948 → 46,356 2/0 ; 17/0

Verdict: the stack still does not meet the owner rule on every row. This PR moves several of its misses:

  • binary-trees n=3/6/10: +734/+638/+543% → +474/+412/+351%;
  • 02_survivor_promotion: +4.3% → −2.8%;
  • the alloc loop: +0.53% → −0.58%;
  • qs_parse: +0.38% → +0.06%.

Still missing:

  • binary-trees n=3/6/10. They are +351–474% in instructions and +52–66% in RSS. The single in-place-promoting minor over the 131,076-object tree still costs 120.5 M instructions (callgrind: 919/object), against a whole program of 25–34 M on main. Its non-trace part alone is ~33 M (in-place stamping 135/object, drain and marks 114/object). So even a free trace would leave these rows at about +120%. This is the pacing decision, not the trace.
  • 12_large_live_set THP-off RSS +1.6% (106.5–107.5 MB vs 104.9–105.7 MB), unchanged by this PR.
  • qs_parse +0.06% and jsonwebtoken +0.06% in instructions. Both are at noise level: jsonwebtoken's samples overlap main's.

Not run

  • The perf/RSS A/B was measured at base b0bf0ae7e. The PR head is the same diff rebased onto 74fed830f, plus two type-level and lint-only commits: ShapeMaskMemo's generic parameter and the census update. The head passed cargo test -p perry-runtime and the lint gates, but the perf/RSS tables and the gap A/B were not re-run on it.
  • The full gap sweep, cargo test beyond perry-runtime (no codegen change), the lint compile tier, and the Windows cargo xwin check.
  • macOS/arm64: every number here is Linux x86_64.
  • Wall-clock and pause times.
  • The claude-code and OpenCode app workloads.
  • The side-table passes that dominate qs minors (hypothesis 5). They are named above as the next target.

Summary by CodeRabbit

  • Performance
    • Reduced copying minor-collection tracing cost from 1,472 to 953 instructions per live object on a fully live binary tree.
    • Peak memory usage remained within ±1% in most benchmarks; moment was an exception.
  • Reliability
    • Added validation to check that optimized garbage collection scans preserve expected object fields and traversal order.
    • Added checks for remembered references and object-layout lookups during copying minor collections.

@coderabbitai

coderabbitai Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

Important

Review skipped

Review was skipped as selected files did not have any reviewable changes.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: 71b539ee-fe94-4103-ac93-f01c7b69ca0b

📥 Commits

Reviewing files that changed from the base of the PR and between f88e4fa and 02ed058.

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: 7ce88637-c02a-4037-b85e-ff7009925301

📥 Commits

Reviewing files that changed from the base of the PR and between b50185b and f88e4fa.

📒 Files selected for processing (1)
  • scripts/gc_runtime_root_holders.json

Included review availability: This review used your included allowance. Your plan provides up to 8 included reviews per hour; 4 remain after this review.


📝 Walkthrough

Walkthrough

The copying nursery collector adds a specialized scan for eligible ordinary objects. It also adds memoized shape-layout lookup, inline traversal and classification paths, and per-parent remembering facts. Tests and audit checks cover scan behavior and helper correctness.

Changes

Copying Nursery Trace Cost

Layer / File(s) Summary
Memoized layouts and inline traversal
crates/perry-runtime/src/gc/{hot_tls.rs,layout.rs,layout/shape_layout_table.rs,layout_slot_visit.rs,copying_pointer_set.rs}, crates/perry-runtime/src/gc/tests/shape_layout_table.rs
Shape-layout lookups memoize pointer masks and clear the memo on mutable access. Layout visitors and pointer classification gain inline-capable paths. Tests check cached results across reads and map updates.
Plain-object scan and parent facts
crates/perry-runtime/src/gc/{copying.rs,copying_object_scan.rs,copying_parent_facts.rs,mod.rs}, crates/perry-runtime/src/object/shapes.rs, crates/perry-runtime/src/gc/tests/{copy_slot_hoists.rs,copying_object_scan.rs,mod.rs}, scripts/{gc_runtime_root_holders.json,shape_descriptor_census.py}, changelog.d/11676-gc-per-object-trace-cost.md
The collector attempts a planned scan for eligible ordinary objects and uses per-parent remembering facts for slot visits. Tests check slot coverage, child evacuation, and carrier notes. Census checks include the scan plan. The changelog reports benchmark results and lists excluded collection work.

Priority: ➖ Normal

Estimated code review effort: 4 (Complex) | ~45 minutes

Change: Refactor

Sequence Diagram(s)

sequenceDiagram
  participant CopyingNurseryCollector
  participant scan_plain_object
  participant PlainObjectPlan
  participant visit_slot_with_parent_facts
  CopyingNurseryCollector->>scan_plain_object: try scanning object
  scan_plain_object->>PlainObjectPlan: build and enumerate eligible object slots
  PlainObjectPlan->>visit_slot_with_parent_facts: visit each planned slot with parent facts
  visit_slot_with_parent_facts-->>scan_plain_object: report slot updates
  scan_plain_object-->>CopyingNurseryCollector: report whether the scan handled the object
Loading

Suggested reviewers: claude

Merge Risk: ⚪ Minimal · up to f88e4

The new scan remains confined to copying-minor collections, leaving the full-cycle census path unaffected. No actionable merge-blocking risk is established.

Architecture Summary

Architecture risk: 🔵 Low · up to f88e4

The change affects 3 systems.

Changed systems: crates, scripts, changelog.d

Architecture concerns
No architecture-level concerns identified.

Review details

Systems and components

  • observed — crates (service) was modified; 14 changed files map to changed impact.
  • observed — scripts (service) was modified; 2 changed files map to changed impact.
  • observed — changelog.d (service) was modified; 1 changed file maps to changed impact.

Before / after behavior

  • observed — Modified behavior in changelog.d/11676-gc-per-object-trace-cost.md: The changelog reports a reduction from 1,472 to 953 instructions per live object on btree n=20, gives additional benchmark changes, and notes peak RSS stayed within ±1% except for moment.
  • observed — Modified behavior in changelog.d/11676-gc-per-object-trace-cost.md: The changelog describes a straight-line ordinary-object slot scan that declines for full or layout-scan traces and incompatible type tables, with test/debug slot-list equivalence checks. It also reports one parent-remembering query per object, generation classification only when needed, mark-memo lookup before raw-word classification, a SHAPE_LAYOUTS memo cleared through DerefMut, skipping an unnecessary old-carrier generation probe when both flags are set, and inlining the drain visitor and classification.
  • observed — Modified behavior in changelog.d/11676-gc-per-object-trace-cost.md: The changelog excludes full collections and lists side-table passes not covered by the change: closure box-capture pruning, per-object layout-owner pruning and sorting, and remembered-set rebuilding.
  • observed — Modified behavior in crates/perry-runtime/src/gc/copying.rs: The import adds ParentRemembering alongside weak_holder_fact.
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 70.00% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 80 functions across 15 files. (1 skipped:… Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly and specifically summarizes the main performance changes to the copying minor collector. It is somewhat long but remains focused and understandable.
Description check ✅ Passed The description is highly complete. It explains the motivation, implementation changes, related issue, benchmarks, correctness tests, known limitations, and tests not run. Although it does not use eve…
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Full details: Docstring Coverage

Explanation

Docstring coverage is 70.00% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 80 functions across 15 files. (1 skipped: 1 unsupported.)

✨ Finishing Touches 💡 1
🛠️ Fix failing CI checks 💡
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@proggeramlug

Copy link
Copy Markdown
Contributor Author

Rebased onto main 7fa094cb4. New head is f88e4fa87. The only conflict was scripts/gc_runtime_root_holders.json. I took main's version and re-applied this PR's two edits: the PASS1_MARKED gc/mod.rs pin, re-derived and re-audited, and the PLAN_ATTEMPTS frontier entry. No code conflicts.

CI failures from run 36628176053: none are caused by this PR. Method: a pristine release build of main 7fa094cb4 against a release build of this branch, each with its own target directory and a pinned PERRY_RUNTIME_DIR, on perrymaster with Node 26.5.1.

failure main 7fa094cb4 this PR cause
lint "Class-id collision audit" fails, identical output fails pre-existing. A test-local const CLASS_ID from #11678 is read as a drifted mirror. Filed as #11691.
test_gap_node_redis_from_source fails: #<perry:private-member:64:#validateOptions> is not a function same output, byte for byte pre-existing. Bisected to #11667 (parent 2c20e92d6 passes, 9f9513717 fails). Filed as #11692.
test_gap_mongodb_from_source fails: TypeError: @@iterator is not a function at parseOptions/new MongoClient same output apart from one stack-frame address pre-existing, same bisect (#11667). Filed as #11692.
test_gap_sloppy_this_bound_once fails (parity_fail in the harness) fails pre-existing. The test came in with #11679. The Node oracle runs the .ts file as strict ESM and throws Cannot create property 'extra' on number '1'. Filed as #11693.

The two from-source tests use in-process fake servers, so no services are needed. I built and ran both arms by hand from test-files/ with the root npm ci deps (redis@6.1.0, mongodb@7.5.0). The harness reported compile_fail for them in the main worktree only, and I did not investigate why. The by-hand runs and the bisect are the evidence here.

Re-validation on f88e4fa87:

  • RUST_TEST_THREADS=1 cargo test --release -p perry-runtime: 4810 passed, 0 failed.
  • cargo fmt --check and check_file_size.sh: OK.
  • SKIP_COMPILE_GATES=1 run_lint_gates.sh: 104 of 107. The three failures are public-baseline, cargo xwin missing on this host, and the class-id audit (lint red on main: Class-id collision audit flags a test-local CLASS_ID from #11678 #11691, red on main).
  • Gap filters on the branch: gc 68/68, array 120/120, class 133 pass and 1 fail. The failure is test_gap_2159_defineproperty_class_prototype, which fails on main too. Main's arm had the same single parity failure. It also had 7 harness compile_fails in that worktree, all in the same environment-artifact category as the from-source tests, and all 7 pass on this branch. The branch arm has no failure that main's arm lacks.

Perf and RSS numbers in the description were measured at the older base and were not re-run after this rebase.

@proggeramlug

Copy link
Copy Markdown
Contributor Author

Rebased onto main 6a50907518 (#11694). New head is 02ed05807. The rebase was clean. On the new head, gc_runtime_root_holders.py, class_id_collisions.py (now green thanks to #11694) and cargo fmt --check pass. On the previous rebase (onto 23d563404), the full cargo test --release -p perry-runtime gave 4819 passed, 0 failed.

The three remaining reds from run 36653797552 all predate this PR. Each fails identically on pristine main, and each bisects to #11682 (b67ad0005) on perrymaster, with 7fa094cb4 passing and b67ad0005 failing. #11682's own PR CI showed the same three reds.

failure pristine main vs this PR bisect issue
perry-stdlib symbols_tests::thread_exit_releases_the_threads_symbol_side_table_entries same assertion ([false, false, false, true]) on 23d563404 and on this branch 7fa094cb4: 10/10 pass, b67ad0005: fails #11696
gc-call-effects macOS and Windows (UNSAFE drift in js_abort_* and js_error_is_error) main's own push run 36653831017 lists the identical 4 UNSAFE symbols #11682 regenerated only the linux table #11695
test_gap_fetch_handle_lifecycle parity_fail 4/4 on 23d563404 and 4/4 on this branch 7fa094cb4: 2/2 pass, b67ad0005: 2/2 fail. It is deterministic: the subclass body reads back as {} #11700

The full cargo test --release -p perry-stdlib on this branch fails only the test above: 241 passed, 1 failed.

@proggeramlug
proggeramlug merged commit 6cb7ed8 into main Sep 30, 2026
52 of 60 checks passed
@proggeramlug
proggeramlug deleted the perf/11549-trace-cost branch September 30, 2026 07:50
proggeramlug pushed a commit that referenced this pull request Sep 30, 2026
- check_file_size: gc/layout.rs was 2009 lines after #11676; move the
  cfg(test) probe layout_descriptor_reachable into layout/test_accessors.rs
  (pure relocation), 1989 lines now.
- class_id_collisions: static_shapes_tests.rs's new test-local
  ANON_CLASS_ID (0x0075_5eed) read as a drifted mirror of put_value.rs's
  ANON_CLASS_ID (0x8783_1001). Rename the test-local to
  REP_SEED_ANON_CLASS_ID; the two tests never shared an id.
- shape_descriptor_census: one keys_array declaration moved from
  object/alloc.rs (3->2) to object/alloc_plain.rs (2->3); total unchanged.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant