Skip to content

perf(gc): the first collection's barrier-arming walk skips wholly-nursery blocks - #11668

Merged
proggeramlug merged 2 commits into
mainfrom
perf/11549-first-minor-arming-walk
Sep 29, 2026
Merged

proggeramlug merged 2 commits into
mainfrom
perf/11549-first-minor-arming-walk

Conversation

@proggeramlug

@proggeramlug proggeramlug commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Part of #11549

Split out of the #11645 follow-up; #11645 now stacks on this PR. The change stands alone on main and does not depend on the nursery pacing.

What changes

At a thread's first collection, the lazily-armed write barrier (#7187) rebuilds the old→young remembered set by walking every arena object (arm_and_reconstruct_remembered_set_if_unarmed → rebuild_minor_old_to_young_remembered_set). The walk keeps a parent only if barrier_parent_needs_remembering says so, which means an Old-generation object or a malloc object. An object on a nursery block is neither, so the walk classified each one and threw it away. The first collection is also when the young generation is at its fullest. On gc_ratchet 02 that walk cost main 41 M instructions (12.9% of the program). On binary-trees at n=3, it was ~20 M for 131 k young objects.

The walk now skips every arena block that is wholly nursery:

The skip is exact, not a heuristic. Every object it skips would have been rejected. The walk runs to completion synchronously inside the arming call, so no mutator window can change a block's generation mid-walk. The budgeted full's remembered-set rebuild uses the same state type and is untouched.

RememberedReconstructCensus gains objects_walked, the live-subject counter the witness test reads.

Tests and sabotage

Numbers (perrymaster, Linux x86_64 Zen 4, --release, PERRY_NO_AUTO_OPTIMIZE=1, under /tmp/perry-bench-lock.d)

main = 10ece9958.

  • Instructions: perf stat -e instructions:u. For loops, a two-N differential per iteration (median of 3 at n1, median of 11 at n2). For fixed-size probes, the total (median of 11).
  • Peak RSS: /usr/bin/time %M on the binary itself (no perf wrapper), median of 11, arms interleaved, [min..max].

Two RSS columns, because transparent huge pages matter here. mimalloc's allow_thp defaults to on, so on a THP host (enabled = madvise here) most of the heap sits in 2 MB pages, and peak RSS moves in 2 MB steps that depend on layout. THP on is the host as-is. THP off (MIMALLOC_ALLOW_THP=0) is 4 KiB accounting, i.e. memory actually touched.

Noise floor: a no-op build (main plus 4 KiB of unused rodata) moves deterministic rows by ≤0.2% THP-on and ≤1.3% THP-off. On dotenv/moment, THP availability flips on some runs of the same binary: see the min column, 31 MB against a 52 MB median.

Instructions and THP-on peak RSS (A0 = this PR):

workload (THP on) main instr A0 instr Δinstr A0 main RSS KB med [min..max] A0 RSS KB med [min..max] ΔRSS A0 GCs minor/full
dotenv_parse 2.79 M 2.79 M -0.00% 52,476 [52,168..52,556] 52,888 [52,764..53,168] +0.8% 0/7 ; 0/7
moment_parse_format 2.33 M 2.32 M -0.72% 65,968 [65,844..68,244] 66,624 [66,348..68,576] +1.0% 1/33 ; 1/33
date-fns_format_add 0.94 M 0.94 M +0.02% 69,072 [68,936..69,340] 69,388 [69,216..69,560] +0.5% 17/0 ; 17/0
validator_batch 19.70 M 20.02 M +1.63% 91,208 [90,984..91,476] 91,628 [91,228..91,832] +0.5% 46/2 ; 46/2
qs_parse_nested 1.96 M 1.96 M +0.01% 232,676 [217,324..239,792] 227,192 [220,168..238,224] -2.4% 11/0 ; 11/0
qs_stringify_nested 5.93 M 5.94 M +0.13% 513,516 [509,064..519,556] 514,272 [503,528..517,404] +0.1% 24/10 ; 24/10
uuid_v4 0.03 M 0.03 M -0.68% 58,104 [30,012..59,136] 49,328 [29,828..59,036] -15.1% 1/0 ; 1/0
jsonwebtoken_decode 0.08 M 0.07 M -3.03% 94,288 [94,084..94,544] 94,200 [93,980..94,336] -0.1% 2/0 ; 2/0
alloc 463.6 463.6 +0.00% 47,852 [47,576..47,988] 47,984 [47,712..48,188] +0.3% 67/0 ; 67/0
bare 14.0 14.0 +0.00% 15,612 [15,480..15,832] 15,740 [15,484..15,804] +0.8% 0/0 ; 0/0
retain 3,348.7 3,348.6 -0.00% 150,084 [143,864..150,620] 150,252 [145,096..150,604] +0.1% 9/1 ; 9/1
btree 9.01 M 7.84 M -12.93% 60,344 [48,548..60,600] 60,160 [48,700..60,424] -0.3% 1/0 ; 1/0
btree@3 (fixed) 25.52 M 25.52 M -0.00% 29,612 [29,360..29,752] 29,456 [29,412..29,712] -0.5% 0/0 ; 0/0
btree@6 (fixed) 29.36 M 29.36 M -0.00% 31,616 [31,528..31,816] 31,536 [31,472..31,648] -0.3% 0/0 ; 0/0
btree@10 (fixed) 34.48 M 34.48 M +0.00% 33,648 [33,528..33,820] 33,640 [33,424..33,844] -0.0% 0/0 ; 0/0
btree@20 (fixed) 279.14 M 244.21 M -12.51% 58,252 [44,592..58,556] 58,132 [49,732..58,448] -0.2% 1/0 ; 1/0
01_nursery_churn@- (fixed) 200.53 M 165.22 M -17.61% 43,956 [43,896..44,256] 44,140 [44,000..44,472] +0.4% 1/1 ; 1/1
02_survivor_promotion@- (fixed) 319.25 M 285.05 M -10.71% 54,176 [53,632..54,484] 54,172 [53,764..54,296] -0.0% 1/1 ; 1/1
12_large_live_set@- (fixed) 3,625.17 M 3,590.99 M -0.94% 146,256 [139,792..146,460] 146,156 [138,832..146,652] -0.1% 5/1 ; 5/1

Noise: instruction spread on an identical binary is ≤0.1% on every row except validator (~5%), moment (~2%) and qs_stringify (~0.5%), so validator's +1.6% and qs_stringify's +0.13% are inside it. THP-on RSS moves ≤0.2% on a no-op build of the deterministic rows. The bare loop runs no GC at all, so none of this code is active there, and it still moves +0.8%, which bounds the file-page variance.

THP-off peak RSS (4 KiB accounting, 11 runs, interleaved):

workload (THP off) main KB med [min..max] A0 KB med [min..max] Δ A0
dotenv_parse@10000 31,628 [31,364..31,888] 32,144 [32,008..32,416] +1.6%
moment_parse_format@10000 40,708 [40,432..40,936] 41,212 [40,872..41,348] +1.2%
date-fns_format_add@20000 36,864 [36,364..37,224] 37,216 [36,968..37,548] +1.0%
validator_batch@5000 60,452 [60,136..60,684] 60,912 [60,296..61,304] +0.8%
qs_parse_nested@20000 194,460 [189,408..198,076] 196,224 [191,148..201,240] +0.9%
qs_stringify_nested@20000 478,216 [464,684..485,384] 480,920 [478,344..483,136] +0.6%
uuid_v4@200000 30,032 [29,752..30,292] 30,040 [29,792..30,284] +0.0%
jsonwebtoken_decode@20000 58,972 [58,456..59,356] 58,960 [58,816..59,156] -0.0%
alloc@8000000 21,328 [20,416..21,660] 21,248 [20,944..21,688] -0.4%
bare@8000000 7,748 [7,288..7,868] 7,860 [7,652..8,028] +1.4%
retain@2000000 110,704 [110,108..110,944] 110,660 [110,096..111,052] -0.0%
btree@3 15,568 [15,220..15,644] 15,384 [15,140..15,632] -1.2%
btree@6 16,396 [16,244..16,536] 16,292 [16,132..16,520] -0.6%
btree@10 17,688 [17,584..17,800] 17,724 [17,512..17,808] +0.2%
btree@20 30,036 [29,860..30,384] 30,300 [29,884..30,640] +0.9%
btree@40 33,820 [33,524..34,052] 33,816 [33,596..34,336] -0.0%
01_nursery_churn@- 23,828 [23,624..24,152] 24,060 [23,848..24,472] +1.0%
02_survivor_promotion@- 24,632 [24,132..25,156] 24,992 [24,260..25,256] +1.5%
12_large_live_set@- 105,400 [104,836..105,752] 105,604 [105,168..105,920] +0.2%

THP-off differences of up to ±1.6% are file-backed page variance, not heap. On dotenv, anonymous RSS is 18.01 MB against main's 18.02 MB, and minor faults are 6,118–6,152 against 6,161. The bare loop again moves +1.4% with no GC active.

An alternative I measured and dropped: shrinking the minor's transient lists

The #11645 follow-up's first item was the ~8 MB of collector bookkeeping in binary-trees at n=3's one minor. That bookkeeping is worklist and moved_headers growing by doubling, with realloc copies and the allocator holding the old buffers. I built a never-reallocating chunked moved_headers and a worklist sized once on promoting cycles. With THP off it cut that minor's peak RSS from 25.3 MB to 22.0 MB when the minor is the peak, and it was instruction-neutral once the rebuild loop stayed non-generic. On main it was a trade, so it is not in this PR.

  • The freed doubling buffers were not waste. They were recycled into the next 1 MB Eden blocks. With the lists shrunk, the mutator faulted the same pages itself: binary-trees at n=40 lost 1,004 collector faults and gained 957 in make.
  • With THP on, that moved memory into new huge pages: binary-trees at n=10→40 +7.0% peak RSS, gc_ratchet 01 +3.6%. Both reproduced, and both are well outside the no-op build's ±0.2%.

The code is on wip/11549-header-list-experiment for anyone who picks this up. The finding is that the doubling waste only matters where a minor over a fully-live young generation is the program's peak.

Correctness

Not run

Summary by CodeRabbit

  • Performance
    • Garbage-collection remembered-set reconstruction now skips memory regions fully covered by a nursery generation range, while continuing to scan other regions and malloc objects.
    • The change reduces the number of objects visited during reconstruction in tested cases. Reported benchmark results show no change to peak memory use.

@coderabbitai

coderabbitai Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: cc5c2c70-8d9f-4a16-9c50-94949d74aa40

📥 Commits

Reviewing files that changed from the base of the PR and between 7601742 and e26749d.

📒 Files selected for processing (2)
  • crates/perry-runtime/src/arena/mod.rs
  • crates/perry-runtime/src/gc/verify.rs

Included review availability: This review used your included allowance. Your plan provides up to 8 included reviews per hour; 5 remain after this review.


📝 Walkthrough

Walkthrough

The minor remembered-set reconstruction walk now skips arena blocks covered by a single registered nursery range. It records the number of objects walked. Tests cover normal skipping and a test-only mode that skips every block.

Changes

Nursery Arming Walk

Layer / File(s) Summary
Range classification and cursor skipping
crates/perry-runtime/src/arena/page_meta/mod.rs, crates/perry-runtime/src/arena/mod.rs, crates/perry-runtime/src/arena/walk.rs
uniform_heap_generation returns a generation when one registered range covers the full interval. ArenaObjectCursor::skip_blocks_where updates the skip set for block-index cursors using a predicate over each block’s allocated extent.
Arming walk integration and validation
crates/perry-runtime/src/gc/verify.rs, crates/perry-runtime/src/gc/barrier_arming.rs, crates/perry-runtime/src/gc/tests/barrier_arming.rs, changelog.d/11668-gc-arming-walk-skips-nursery.md
The reconstruction walk skips blocks classified as nursery and records the number of objects walked. Tests check that the walk skips nursery blocks while retaining remembered-set coverage, and that the test-only skip-all mode walks no objects and leaves the parent page uncovered. The changelog reports benchmark comparisons and unchanged peak RSS.

Priority: ➖ Normal

Estimated code review effort: 3 (Moderate) | ~20 minutes

Change: Refactor

Sequence Diagram(s)

sequenceDiagram
  participant Rebuild as rebuild_minor_old_to_young_remembered_set
  participant Cursor as ArenaObjectCursor
  participant SkipCheck as arming_walk_skips_block
  participant Generation as uniform_heap_generation
  participant Census as RememberedReconstructCensus
  Rebuild->>Cursor: Configure skip_blocks_where
  Cursor->>SkipCheck: Check block extent
  SkipCheck->>Generation: Classify the full interval
  Generation-->>SkipCheck: Return generation or None
  Cursor->>Rebuild: Visit non-skipped objects
  Rebuild->>Census: Record objects walked
Loading

Merge Risk: ⚪ Minimal · up to e2674

The optimization preserves remembered-set reconstruction while avoiding scans of wholly Nursery blocks.

Security Architecture Review

Security architecture risk: 🔵 Low · up to e2674

The change affects which heap objects are scanned before collection, so a mistaken skip could lose reference tracking. The skip requires a wholly nursery block, and the added tests check that an old parent’s reference is still recovered. No new externally callable control was identified.

Retained concerns
No architecture-level concerns identified.

Security review details

Security Blast Radius

  • inferred — A missed edge would affect GC reference tracking in the collecting runtime, but the changed skip is an internal GC decision rather than an added caller-controlled interface.

Security Findings and Attack Paths

  • observed — The control that forces every block to be skipped is test-gated; the production predicate requires a wholly Nursery range. No production route to the test override is shown.

Trust Boundaries and Controls

  • observed — The predicate requires complete Nursery range coverage, while the barrier is armed before reconstruction so stores occurring after arming remain subject to logging.
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 72.22% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 18 functions across 6 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description check ✅ Passed The description clearly explains the change, motivation, implementation, tests, benchmark results, related issue, and unrun checks. It does not follow every template heading or checklist item, but it …
Title check ✅ Passed The title is concise, specific, and accurately identifies the main change: skipping wholly nursery blocks during the first collection's barrier-arming walk.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
  • Fix all pre-merge checks with AI
✨ Finishing Touches 💡 2
📝 Generate docstrings 💡
  • Commit to this branch
  • Create a new PR
🛠️ Fix failing CI checks 💡
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants