Skip to content

HOLD (RSS trade): perf(regex): stop charging operation scratch to GC pressure; grow lent registers (bounded) - #11612

Closed
proggeramlug wants to merge 2 commits into
mainfrom
perf/11549-regex-scratch-pressure
Closed

proggeramlug wants to merge 2 commits into
mainfrom
perf/11549-regex-scratch-pressure

Conversation

@proggeramlug

Copy link
Copy Markdown
Contributor

Part of #11549

HOLD: do not land this alone. It is a compute-for-RSS trade on two package rows (moment/parse_format RSS +29%, and the dotenv-shaped microbench +32%). Under the owner's rule (minimize RSS and keep the best compute, never trade one for the other) that is not ready. The RSS cost is localised below: it is the standard 16 MB nursery, which the phantom pressure used to pre-empt. Recovering it takes the young-generation pacing work that #11549 lists as direction 2. This PR is the accounting half, measured and ready to pair with that work.

What changed

  1. Regex operation scratch is no longer reported to the collector (regex/perex_memory.rs). Buffer and Reservation cover the owned search path's match buffers, the compiler's node and range scratch, the KMP failure table and the replacer argument slots. They used to call gc_note_external_side_alloc and gc_note_external_side_free. Every one of them is freed by the operation that allocated it, when it returns or unwinds. No collection can reclaim a byte of it. But each release was added to GC_EXTERNAL_SIDE_DRAINED_SINCE_FULL, which old-reclaim holds as pressure until the next full collection. So a loop over a program with more than 32 registers ran a full mark-sweep every few hundred calls on phantom bytes. In PERRY_GC_DIAG output that shows as external_drained=33553968, arena_total=3145728, old_in_use=0, 81 times in one run.
    • What bounds this memory instead: every buffer, inline charge and reservation is checked against the operation's hard MemoryBudget limit (perex_api::SCRATCH_BYTES) before it exists. That check is unchanged.
    • Pressure accounting is kept for memory that grows with the subject. perex_replace_direct::Spans and the native piece records in perex_replace_storage still report, because their size follows the input rather than a fixed limit. Map, Set, JSON-tape and N-API external bytes are untouched.
    • StorageError::Abrupt is deleted. Its only producer was the collection that the note could trigger.
  2. The lent per-thread scratch cell grows its registers up to a fixed bound (regex/perex_runtime.rs). Registers used to be a fixed [usize; 32]. They are now a Vec grown on demand to LENT_REGISTERS = 1024, so the cell retains at most 8 KiB per thread whatever programs run. dotenv's LINE needs 42 registers, so with the fixed 32 every search built and dropped owned buffers. It now builds them once per thread. A register count is a property of the program, so the choice of path is made before any work is done, never through a failed attempt. A program past 1024 registers still takes the owned path, which the budget bounds and the operation frees.
    • Unchanged, and worth knowing: the cell's frames and undo vectors already grew on demand and were never reported. They are bounded only indirectly, by the per-operation Charge over the whole cell. I left that alone here.

Measurements (perrymaster, Linux x86_64; both arms built here with --release from 5af043c, differing only by this diff; PERRY_NO_AUTO_OPTIMIZE=1)

Instruction counts are perf stat -e instructions:u, a two-N differential with the median of 3 at each N, taken under /tmp/perry-bench-lock.d. The host load average was about 23. Instructions per iteration:

probe main this PR Δ
bare loop (control) 14.0 14.0 0
non-regex alloc loop (control) 469.2 469.2 0
RE.test uuid, 2-register program (control) 13,673 13,668 −0.04%
dotenv LINE.exec over 40 lines (42 registers) 1,704,257 1,077,210 −36.8%
s.replace(/(\d+)/g, fn) 106,239 103,614 −2.5%

Package workloads (scripts/package_bench.py run --arms perry --modes instr). Every run's output matched Node 26.5.1 byte for byte:

workload main this PR Δ
dotenv/parse 2,881,147 2,013,012 −30.1%
moment/parse_format 3,355,980 2,893,521 −13.8%
validator/batch 27,120,208 26,600,285 −1.9%
date-fns/format_add 1,144,781 1,131,493 −1.2%
dayjs/parse_format 1,817,833 1,813,692 −0.2%
uuid/v4, v5_parse, v7; jsonwebtoken/decode; qs ×2; dayjs/diff_startof; date-fns/diff_interval; moment/diff_duration; validator/sanitize; controls within ±0.3%

Collections (PERRY_GC_DIAG=1, one run at n2) and peak RSS (/usr/bin/time -f %M, median of 11 runs, with the [min..max] range):

workload GCs, main GCs, this PR peak RSS KB, main peak RSS KB, this PR
dotenv/parse 47 budgeted full + 7 sync full 31 copying minor 42,020 [31,736..46,256] 37,588 [37,044..39,664]
moment/parse_format 33 sync full 23 copying minor 38,624 [38,304..44,856] 49,912 [49,456..57,184]
validator/batch 21 budgeted full + 1 sync + 93 minor 107 minor 89,832 91,628
date-fns/format_add 1 budgeted full + 35 minor 35 minor 36,932 38,588
dotenv-shaped microbench 32 budgeted full 11 copying minor 40,708 53,944
date-fns/diff_interval, dayjs/parse_format, validator/sanitize, replace/uuid microbenches unchanged unchanged flat flat

main's RSS on dotenv is bimodal and depends on load, because the budgeted fulls are time-sliced. In an earlier, quieter set of 11 runs its median was 31.5 MB against this PR's 37.8 MB. The PR arm is stable in both sets.

Where the RSS goes, and what recovers it

Without the phantom fulls these loops pace the way every other allocating Perry program already does: the young generation fills to the 16 MB scavenge nursery cap before a copying minor runs. main was collecting them at about 3 MB of young data, using full collections that cost 12.7% of dotenv's wall time ([gc-time] share_permille=127, against 10 with this PR).

PERRY_GC_SCAVENGE_NURSERY_MB measures the lever directly, on this PR's binaries (instructions per iteration, RSS from one run):

nursery dotenv/parse moment/parse_format non-regex alloc loop
16 MB (default) 2,012,782 2,893,242 469.2
8 MB 2,014,742 (+0.1%), RSS 45.6 MB 2,894,266 (+0.04%), RSS 56.3 MB 474.9 (+1.2%)
4 MB 2,021,469 (+0.4%), RSS 35.3 MB 2,900,427 (+0.25%), RSS 40.1 MB 483.1 (+3.0%)

A 4 MB nursery would give both wins on the regex rows: −29.8% and −13.6% instructions at main's RSS. But a smaller global nursery costs +3% instructions on an allocation-bound loop. So lowering the nursery cap is itself a trade, and it does not belong in this PR.

Per minor, the fixed cost is about 310–380k instructions on the alloc loop and about 720k on dotenv. A symbolized profile at a 1 MB nursery puts most of the minor-side cost in RuntimeRootVisitor::visit_tagged_raw_addr (8.7% of the run) and array_tail_transition::{scan_table, prune_table} (5.3%). Cutting that fixed cost is what would make a smaller or work-aware nursery free, and it is the natural next step for direction 2.

Tests

  • New gc/tests/runtime_roots/perex_scratch_pressure.rs has three tests:
    • An operation Buffer leaves external_side_live_bytes and the drained counter untouched, and the budget still bounds it.
    • A 42+-register program runs 64 searches on the lent cell with zero owned-path searches and no external bytes. This uses a new #[cfg(test)] OWNED_SEARCHES counter.
    • A program past the lent bound (600 groups) takes the owned path on every call and still notes nothing.
  • Fails without the fix: with main's perex_memory.rs, tests 1 and 3 fail (left: (32768, 0) right: (0, 0); drained 200256 vs 161792). With LENT_REGISTERS = 32, test 2 fails (65 owned searches, not 1). All three pass with the fix.
  • Existing test updated: perex_execution::perex_host_buffers_account_overlap_failure_and_unwind asserted the old contract, that a buffer's bytes appear as external side bytes. It now asserts the budget's overlap accounting and an unchanged external reading.
  • RUST_TEST_THREADS=1 cargo test --release -p perry-runtime (codegen-units 16): 4701 passed, 0 failed, 5 ignored.
  • RUSTFLAGS="-D warnings" cargo check -p perry-runtime --all-targets (dev profile): clean. cargo fmt --all -- --check: clean. scripts/check_file_size.sh: OK.
  • Gap A/B against Node 26.5.1 (/opt/node-v26.5.1-linux-x64), PERRY_SKIP_BUILD=1 PERRY_NO_AUTO_OPTIMIZE=1, with PERRY_BIN and PERRY_RUNTIME_DIR pinned to each arm's own -p perry -p perry-runtime-static -p perry-stdlib-static release build. Filters were test_gap_ combined with each of regex, regexp, split, replace, string and gc: 139 distinct tests. Both arms: 134 PASS and the same 5 pre-existing COMPILE_FAILs (test_gap_regex_replace_dyn_regex_with_http, test_gap_11258_eventemitter_async_resource_subclass_gc, test_gap_9552_cross_thread_promise_survives_gc, test_gap_gc_http2_pending_event_callback_rooting, test_gap_gc_net_once_flags_rekey). Per-test results identical.
  • SKIP_COMPILE_GATES=1 scripts/run_lint_gates.sh: 97 of 100 script gates passed; the compile tier was not run. The 3 failures are pre-existing:
    • cargo xwin is not installed on this host.
    • "Public benchmark evidence freshness" is known-red on main.
    • gc_runtime_root_holders.py flags GC_EXTERNAL_SIDE_ALLOC_PENDING and GC_EXTERNAL_SIDE_LIVE_BYTES in gc/policy.rs, which this PR does not touch. It fails identically on 5af043c.
    • git diff --stat was clean afterwards.

Not run: the full gap sweep, cargo test --workspace, macOS, the default auto-optimize perry compile arm, wall-clock measurements (this host is shared and loaded), the Windows xwin check, and the compile tier of the lint gates.

No test outside perry-runtime's regex and GC test modules is expected to change.

… grow the lent registers to a fixed bound

Regex operation scratch (perex_memory::Buffer / Reservation: the owned
search path's match buffers, compile scratch, KMP failure tables, replacer
argument slots) is freed by the operation that allocated it and bounded by
that operation's MemoryBudget. Reporting it through
gc_note_external_side_alloc/free put every per-call release into
GC_EXTERNAL_SIDE_DRAINED_SINCE_FULL, which old-reclaim holds as pressure
until the next full, so a loop over a program with more than 32 registers
(dotenv's LINE has 42) ran a full mark-sweep every few hundred calls on
bytes no collection could free.

The lent per-thread scratch cell now grows its registers on demand up to
LENT_REGISTERS = 1024 (8 KiB per thread), so such programs stop building
per-call buffers at all. Programs past the bound keep the owned path.

Subject-proportional storage (replace Spans, native piece records) is
unchanged and still reported.

Part of #11549
@coderabbitai

coderabbitai Bot commented Sep 28, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@proggeramlug

Copy link
Copy Markdown
Contributor Author

The survival-aware nursery pacing this PR was waiting for is in draft #11645. It carries this PR's two commits (and #11634's) unchanged, so they can land together. It also repairs a gate that this PR alone turns red on current main: scripts/gc_runtime_root_holders.py flags GC_EXTERNAL_SIDE_ALLOC_PENDING / GC_EXTERNAL_SIDE_LIVE_BYTES as new rule-T holders once regex scratch stops reporting to them. #11645 classifies both as byte counters. #11645 is itself on HOLD: it improves every package row but leaves three GC-probe rows above main (see its body).

proggeramlug pushed a commit that referenced this pull request Sep 29, 2026
…rs inventory

#11612 stops regex operation scratch from reporting to
GC_EXTERNAL_SIDE_ALLOC_PENDING / GC_EXTERNAL_SIDE_LIVE_BYTES, so the
reachability walk no longer reaches them from a registered scanner and
gc_runtime_root_holders.py flags both as new rule-T holders. They are
usize byte counters and cannot hold a GC pointer.
proggeramlug pushed a commit that referenced this pull request Sep 29, 2026
…rs inventory

#11612 stops regex operation scratch from reporting to
GC_EXTERNAL_SIDE_ALLOC_PENDING / GC_EXTERNAL_SIDE_LIVE_BYTES, so the
reachability walk no longer reaches them from a registered scanner and
gc_runtime_root_holders.py flags both as new rule-T holders. They are
usize byte counters and cannot hold a GC pointer.
proggeramlug pushed a commit that referenced this pull request Sep 29, 2026
…rs inventory

#11612 stops regex operation scratch from reporting to
GC_EXTERNAL_SIDE_ALLOC_PENDING / GC_EXTERNAL_SIDE_LIVE_BYTES, so the
reachability walk no longer reaches them from a registered scanner and
gc_runtime_root_holders.py flags both as new rule-T holders. They are
usize byte counters and cannot hold a GC pointer.
proggeramlug pushed a commit that referenced this pull request Sep 30, 2026
…rs inventory

#11612 stops regex operation scratch from reporting to
GC_EXTERNAL_SIDE_ALLOC_PENDING / GC_EXTERNAL_SIDE_LIVE_BYTES, so the
reachability walk no longer reaches them from a registered scanner and
gc_runtime_root_holders.py flags both as new rule-T holders. They are
usize byte counters and cannot hold a GC pointer.
proggeramlug pushed a commit that referenced this pull request Sep 30, 2026
…rs inventory

#11612 stops regex operation scratch from reporting to
GC_EXTERNAL_SIDE_ALLOC_PENDING / GC_EXTERNAL_SIDE_LIVE_BYTES, so the
reachability walk no longer reaches them from a registered scanner and
gc_runtime_root_holders.py flags both as new rule-T holders. They are
usize byte counters and cannot hold a GC pointer.
proggeramlug pushed a commit that referenced this pull request Sep 30, 2026
…rs inventory

#11612 stops regex operation scratch from reporting to
GC_EXTERNAL_SIDE_ALLOC_PENDING / GC_EXTERNAL_SIDE_LIVE_BYTES, so the
reachability walk no longer reaches them from a registered scanner and
gc_runtime_root_holders.py flags both as new rule-T holders. They are
usize byte counters and cannot hold a GC pointer.
@proggeramlug

Copy link
Copy Markdown
Contributor Author

This lands via #11645, which carries its two commits unchanged (perf(regex): stop charging operation scratch to GC external pressure; grow the lent registers to a fixed bound and its changelog fragment changelog.d/11612-regex-scratch-not-gc-pressure.md). Please close this PR when #11645 merges, not before.

proggeramlug added a commit that referenced this pull request Sep 30, 2026
…nflux ladder (includes #11612) (#11645)

* perf(regex): stop charging operation scratch to GC external pressure; grow the lent registers to a fixed bound

Regex operation scratch (perex_memory::Buffer / Reservation: the owned
search path's match buffers, compile scratch, KMP failure tables, replacer
argument slots) is freed by the operation that allocated it and bounded by
that operation's MemoryBudget. Reporting it through
gc_note_external_side_alloc/free put every per-call release into
GC_EXTERNAL_SIDE_DRAINED_SINCE_FULL, which old-reclaim holds as pressure
until the next full, so a loop over a program with more than 32 registers
(dotenv's LINE has 42) ran a full mark-sweep every few hundred calls on
bytes no collection could free.

The lent per-thread scratch cell now grows its registers on demand up to
LENT_REGISTERS = 1024 (8 KiB per thread), so such programs stop building
per-call buffers at all. Programs past the bound keep the owned path.

Subject-proportional storage (replace Spans, native piece records) is
unchanged and still reported.

Part of #11549

* changelog: #11612 regex scratch is not GC pressure

* perf(gc): survival-aware nursery pacing: the ladder starts at a quarter of the base

The scavenge nursery cap's influx-driven ladder (#7377) now counts in
quarters of the base cap and powers on at a quarter of it: 4 MB instead
of 16 MB. A thread stays there while its minors find little alive and
climbs toward the unchanged 64 MB top while survivor influx stays above
4% of the cap, with the same debounced 4%/1% band.

A move now takes every step the one-step rule would take in a row on
the same reading, so a survivor-heavy program settles at the size it
settled at before without two extra minors per step. The landing is
inside the dead band, so it cannot oscillate.

The survivor target, the promoted-cohort floor, the allocation-census
seed point and the JSON-leaf gate stay keyed on the 16 MB base: they
measure lifetimes in bytes allocated, which a smaller Eden must not
shorten.

Part of #11549.

* chore(gc): classify the external-side byte counters in the root-holders inventory

#11612 stops regex operation scratch from reporting to
GC_EXTERNAL_SIDE_ALLOC_PENDING / GC_EXTERNAL_SIDE_LIVE_BYTES, so the
reachability walk no longer reaches them from a registered scanner and
gc_runtime_root_holders.py flags both as new rule-T holders. They are
usize byte counters and cannot hold a GC pointer.

* changelog: #11645 survival-aware nursery pacing

---------

Co-authored-by: Perry Bot <perry-bot@users.noreply.github.com>
@proggeramlug

Copy link
Copy Markdown
Contributor Author

Landed as part of #11645 (merged as 9e29f59).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant