Repository navigation
A gc shared library loads without starting the collector's markers under the loader lock - #521
Merged
Merged
Conversation
…der the loader lock Every JIT run against the default library hung at 0% CPU since #516, and the release workflow's "Test Compiler with Default Library" step with it. #516 calls __tslang_gc_enter first in every exported function of a -mm=gc shared library, and its first call ran GC_enable_threads. An exported function can run while the library is being loaded - the default library's exported Number.static_constructor is a global constructor, and top-level code can call an export - and on Windows that is inside DllMain, under the loader lock. GC_allow_register_threads starts the collector's parallel marker threads and waits for them to come up; a new thread cannot start until the loader lock is released, so LoadLibrary never returned. --emit=exe links the static default library, which runs no DllMain, so only the DLL hung. __tslang_gc_enter now enables threads only for a thread it has to register. The thread a library loads on is the one GC_init registered, or one its host registered, so the load no longer starts the markers. test-compile-foreign-thread-gc: two more libraries, one whose exported class has a static constructor and one whose top-level code calls an export, both allocating into a global so the optimizer keeps the call. Both hung before this change. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 5, 2026
Closed
ASDAlexander77
added a commit
that referenced
this pull request
Oct 6, 2026
…d; coroutine frames uncollectable (#524) * A gc shared library enables the collector's threads on the first call after its load Fixes #522, #523. #521 stopped __tslang_gc_enter from enabling threads for a thread the collector already knew, because on the thread loading the library that ran inside DllMain and deadlocked. That left two holes: - #522: the library's own async pool registers its threads through hooks that only the library's GC_enable_threads sets (each module has its own copy of the scheduler). When only registered threads called in, they were never set, and a pool thread that collected aborted ("Collecting from unknown thread"). - #523: for a host that is not a tslang program, the first unregistered thread enabled threads. GC_allow_register_threads makes the allocator take its lock only after it started the markers, so it raced a registered thread that was already allocating. Now the first call that may enables threads, whichever thread makes it, before it goes on. That is any call except one inside a DllMain (the thread holds the PEB's LoaderLock) while the markers have not started. Such a call goes on as it is, and the loading thread's next call, after the load, enables them. No other thread can call in before the load finishes. test-compile-foreign-thread-gc: - The host checks that the markers run after the loading thread's calls whenever they run at the end. A collector with no parallel marking never starts them, and then there is nothing to check. - An async-pool library works on its own pool and collects there. Both failed before this change. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Never enable the collector's threads under the loader lock, even with the markers counted GC_get_parallel counts the markers before they are waited for, so a call from inside another DLL's DllMain could pass the check and wait on the thread starting them, which waits for this thread's loader lock. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * -mm=gc: an async function's frame is uncollectable, freed when its coroutine finishes The coroutine lowering allocates a frame with aligned_alloc and frees it with free; GCPass renamed the pair to GC_memalign/GC_free, so the frame was collectable. A frame waiting to run is held only by the runtime's pool queue or a token's awaiter list, ordinary heap the collector does not scan, so a collection freed it and the pool resumed whatever reused the block. On Linux, eight host threads awaiting at once in a gc shared library crashed every time (another call's result came back now and then); GC_DONT_GC made it go away. It needs collections while frames wait, which several awaiting threads make likely. GCPass now rewrites aligned_alloc(align, size) to GC_malloc_uncollectable(size): it lives until GC_free and is still scanned, so what a frame holds across a suspension stays alive too. The alignment is dropped: frames ask for 8, and the collector aligns every object to two words. The JIT runtime exports GC_malloc_uncollectable beside GC_memalign. check-x86-run.sh expects the i686 frame allocator as GC_malloc_uncollectable(i32) under -mm=gc. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Every JIT run against a
-mm=gcdefault library has hung at 0% CPU since #516, and the release workflow hung with it: run 37377707693 (v0.0-pre-alpha90) sat in "Test Compiler with Default Library" until it was cancelled.Cause
#516 calls
__tslang_gc_enterfirst in every exported function of a gc shared library, and its first call ranGC_enable_threads. An exported function can run while the library is being loaded. The default library's exportedNumber.static_constructoris a global constructor, and top-level code can also call an export. On Windows that happens insideDllMain, under the loader lock.GC_allow_register_threadsstarts the collector's parallel marker threads and waits for them to come up. A new thread can't start until the loader lock is released, soLoadLibrarynever returned.Stacks of the hung
tslang --emit=jit(ProcDump + WinDbg):GC_MARKERS=1(no marker threads) makes the same run finish. A plain C++DllMainthat callsGC_allow_register_threadshangs the same way.--emit=exelinks the static default library and runs noDllMain, so it was unaffected. The age of the bundled gc.dll is not involved: CI builds bdwgc with its defaultenable_parallel_mark=ON.Fix
__tslang_gc_enternow enables threads only for a thread it has to register. The thread a library loads on is the oneGC_initregistered, or one its host registered, so loading no longer starts the markers. For that thread this is what every release before #516 did. The program's own injectedGC_enable_threadsstill starts the markers, outside the loader lock.Test
test-compile-foreign-thread-gcgets two more libraries. In one, an exported class has a static constructor; in the other, top-level code calls an export. Both allocate into a global, so--optcan't fold the call away: a first attempt whose top level calledwork(1)was folded to a constant and never hung. Both libraries hung before this change. A regression costs 2 × 240 s before the test fails.Local (Windows, Release): full ctest 3895/3895. A default library built from DefaultLib main against this runtime passes test-jit-/test-compile-gc-defaultlib-collector, and
tests.ps1release compile 162/162 and release JIT 162/162. The debug pass, Linux and Android were not run locally.Known gaps (follow-up)
__tslang_gc_enter. Its pool threads would then run coroutines unregistered, with the allocator unlocked. That matches every release before A gc shared library registers the threads its host calls it on #516, but A gc shared library registers the threads its host calls it on #516 had covered it.GC_allow_register_threadsstarts them before it setsGC_need_to_lock, so a registered thread allocating at that moment races it. The window is once per process and narrow (A gc shared library registers the threads its host calls it on #516 had the same window when a foreign thread called first).Both want threads enabled from a registered thread once the load has finished (a flag set when loading ends, or loader-lock detection), and that needs its own design.
🤖 Generated with Claude Code