Problem
The resource manager loses large DAX devices on every reboot because auto-rescan only runs while the pool is empty.
On boot, the RM starts and scans before the kernel has finished bringing up the big CXL device. Small devices come up fast and join the pool; the auto-rescan loop (rescanIfEmpty, added in #48) then stops retrying because the pool is no longer empty — and the device that shows up a few seconds later never gets picked up.
Measured timeline from the 2026-07-13 reboot (turin-gb2):
00:18:38 boot
00:18:49 RM starts, initial scan fails for all devices (header read -6, nodes not ready)
00:18:59 auto-rescan: dax1.0 + dax2.0 join (466 GiB pool)
00:19:xx /dev/dax0.0 (3.84 TB gaia) node appears — ~10 s too late, never scanned again
Result: the 3.5 TiB device silently missing from the pool, and any allocation that needs it fails (Alloc failed with status -2 for explicit dax_path, -12 for sized requests that no longer fit). Reproduced on both the 07-13 and 07-14 reboots — it is deterministic whenever RM startup beats the device node.
Current workaround: kill -HUP <rm pid> triggers a live rescan (existing pools untouched, missing devices adopted). We have had to do this manually after every reboot.
Proposed fix
Drop the empty-pool condition: run the periodic rescan in the main loop unconditionally (same 10 s cadence is fine). rescanDevicesLocked() already skips pools that exist by id/path, so a no-op rescan is cheap — it is a directory scan plus a header read per new device only.
Alternatives considered: a udev rule that HUPs the RM on dax device add (more moving parts), or systemd device ordering (can't know which dax devices to wait for in general). The unconditional periodic rescan is a few lines in main.cpp and also covers devices hotplugged long after boot.
Problem
The resource manager loses large DAX devices on every reboot because auto-rescan only runs while the pool is empty.
On boot, the RM starts and scans before the kernel has finished bringing up the big CXL device. Small devices come up fast and join the pool; the auto-rescan loop (
rescanIfEmpty, added in #48) then stops retrying because the pool is no longer empty — and the device that shows up a few seconds later never gets picked up.Measured timeline from the 2026-07-13 reboot (turin-gb2):
Result: the 3.5 TiB device silently missing from the pool, and any allocation that needs it fails (
Alloc failed with status -2for explicit dax_path, -12 for sized requests that no longer fit). Reproduced on both the 07-13 and 07-14 reboots — it is deterministic whenever RM startup beats the device node.Current workaround:
kill -HUP <rm pid>triggers a live rescan (existing pools untouched, missing devices adopted). We have had to do this manually after every reboot.Proposed fix
Drop the empty-pool condition: run the periodic rescan in the main loop unconditionally (same 10 s cadence is fine).
rescanDevicesLocked()already skips pools that exist by id/path, so a no-op rescan is cheap — it is a directory scan plus a header read per new device only.Alternatives considered: a udev rule that HUPs the RM on dax device add (more moving parts), or systemd device ordering (can't know which dax devices to wait for in general). The unconditional periodic rescan is a few lines in
main.cppand also covers devices hotplugged long after boot.