linux: support protection keys for Chromium sandboxing - #522
Conversation
Backport Linuxulator pkey_alloc, pkey_free, and pkey_mprotect support for Chromium V8 sandboxing. Add the required XSAVE layout helpers and Linux-compatible PKRU process lifecycle handling. AI-Assisted-by: OpenAI Codex (GPT-5) Signed-off-by: Lucas Holt <luke@foolishgames.com>
AI-Assisted-by: OpenAI Codex (GPT-5) Signed-off-by: Lucas Holt <luke@foolishgames.com>
Reviewer's GuideThis backport enables Linux protection-key syscalls for Chromium/V8 sandboxing by combining common Linux ABI validation, amd64 native PKU and XSAVE state handling, per-process allocation lifecycle management, and unsupported-architecture stubs, with corresponding module and documentation updates. Sequence diagram for Linux protection-key allocation and memory taggingsequenceDiagram
participant App as Linux application
participant ABI as Linux pkey syscall ABI
participant Common as linux_pkey_*_common
participant PKU as amd64 PKU backend
participant VM as amd64_pkru_update
App->>ABI: pkey_alloc(flags, init_val)
ABI->>Common: linux_pkey_alloc_common(flags, init_val)
Common->>PKU: linux_pkey_alloc_machdep(td, init_val)
PKU-->>Common: allocated key
Common-->>App: pkey
App->>ABI: pkey_mprotect(addr, len, prot, pkey)
ABI->>Common: linux_pkey_mprotect_common(addr, len, prot, pkey)
Common->>PKU: linux_pkey_mprotect_machdep(td, addr, len, prot, pkey)
PKU->>VM: amd64_pkru_update(td, addr, len, pkey, flags, clear)
VM-->>App: result
State diagram for Linux protection-key lifecyclestateDiagram-v2
[*] --> LinuxProcess
LinuxProcess --> LinuxProcess: linux_pemuldata_init_md()
LinuxProcess --> ForkedProcess: fork inherits md_pkey_allocation_map
ForkedProcess --> ForkedProcess: pkey_alloc / pkey_free
LinuxProcess --> ExecedProcess: exec
ExecedProcess --> ExecedProcess: linux_pemuldata_exec_md()
ExecedProcess --> PKRUInitialized: linux_pkru_exec_init()
PKRUInitialized --> LinuxProcess: allocation map = LINUX_PKEY_INITIAL_MAP
File-Level Changes
Tips and commandsInteracting with Sourcery
Customizing Your ExperienceAccess your dashboard to:
Getting Help
|
There was a problem hiding this comment.
Hey - I've found 3 issues
Prompt for AI Agents
Please address the comments from this code review:
## Individual Comments
### Comment 1
<location path="sys/amd64/linux/linux_pkru.c" line_range="185-187" />
<code_context>
+
+ pem = pem_find(td->td_proc);
+ LINUX_PEM_XLOCK(pem);
+ if ((pem->pem_md.md_pkey_allocation_map & (1u << pkey)) == 0) {
+ LINUX_PEM_XUNLOCK(pem);
+ return (EINVAL);
+ }
+ pem->pem_md.md_pkey_allocation_map &= ~(1u << pkey);
</code_context>
<issue_to_address>
**issue (bug_risk):** `pkey_free(0)` succeeds and removes key 0 from the allocation map, even though Linux reserves key 0 and rejects attempts to free it with `EINVAL`. Subsequent `pkey_mprotect(..., 0)` calls then incorrectly fail because the default key is no longer marked allocated.
**Triggers:** When a Linux application calls `pkey_free(0)`.
**Suggested fix:** Reject pkey 0 explicitly before modifying the allocation map.
</issue_to_address>
### Comment 2
<location path="sys/amd64/linux/linux_pkru.c" line_range="217-226" />
<code_context>
+ }
+ LINUX_PEM_SUNLOCK(pem);
+
+ error = linux_mprotect_common(td, addr, len, prot);
+ if (error != 0 || len == 0)
+ return (error);
+
+ /*
+ * Tag the range; a pkey of 0 untags it. The tag is not
+ * persistent: it dies with the mapping, matching Linux VMA
+ * semantics.
+ */
+ return (amd64_pkru_update(td, addr, len, pkey, 0, pkey == 0));
+}
</code_context>
<issue_to_address>
**issue (broader_impact):** `linux_pkey_mprotect_machdep` applies the ordinary memory protections before assigning the protection-key tag. If `amd64_pkru_update` fails, the syscall returns an error while leaving the mapping's protections changed and its protection-key tag unchanged, producing a partial update.
**Triggers:** When `linux_mprotect_common` succeeds but `amd64_pkru_update` fails, such as from a pmap allocation failure.
**Suggested fix:** Arrange for validation and pkey tagging to succeed before committing the protection change, or roll back the ordinary protection change when the pkey update fails.
</issue_to_address>
### Comment 3
<location path="sys/amd64/linux/linux_emul_md.h" line_range="26-27" />
<code_context>
+#define LINUX_PKEY_INITIAL_MAP 0x1
+
+/*
+ * Initial PKRU at exec: access disabled for keys 1..15, key 0 open;
+ * the Linux init_pkru default.
+ */
+#define LINUX_PKRU_INIT 0x55555554
+
+struct thread;
</code_context>
<issue_to_address>
**nitpick:** The comments describe `0x55555554` as disabling access for keys 1 through 15, but that value sets the write-disable bit for those keys and leaves their access-disable bits clear. The in-code explanation therefore misstates the actual PKRU behavior and can mislead future changes to the sandbox initialization.
**Suggested fix:** Describe the value as disabling writes while allowing reads, or use the actual access-disable encoding if full access denial is intended.
```suggestion
* Initial PKRU at exec: writes disabled for keys 1..15, reads allowed;
* the Linux init_pkru default.
```
</issue_to_address>|
@claude review this |
|
Claude encountered an error after 2s —— View job I'll analyze this and get back to you. |
|
@claude try to review again |
| */ | ||
|
|
||
| int | ||
| linux_pkey_alloc_machdep(struct thread *td, uint64_t init_val) |
There was a problem hiding this comment.
High: linux64.ko on arm64 will have undefined symbols. This PR removes the DUMMY(pkey_mprotect/pkey_alloc/pkey_free) stubs from sys/compat/linux/linux_dummy.c (compiled into linux_common.ko), and sys/arm64/linux/linux_sysent.c references linux_pkey_mprotect, linux_pkey_alloc and linux_pkey_free (syscalls 288-290). The syscall entry points were only added to sys/amd64/linux/linux_machdep.c and sys/amd64/linux32/linux32_machdep.c; nothing on arm64 defines them, so the module fails to link/load. Upstream FreeBSD put the wrappers in the shared sys/compat/linux/linux_misc.c for this reason. Fix: add the three wrappers to sys/arm64/linux/linux_machdep.c (next to linux_mprotect), or move them to linux_misc.c.
There was a problem hiding this comment.
Addressed by moving the three pkey syscall wrappers into shared sys/compat/linux/linux_misc.c, matching the current FreeBSD layout. This gives arm64 linux64.ko definitions for linux_pkey_mprotect, linux_pkey_alloc, and linux_pkey_free after the dummy stubs are removed.
| */ | ||
|
|
||
| int | ||
| linux_pkey_alloc_machdep(struct thread *td, uint64_t init_val) |
There was a problem hiding this comment.
High: Same problem as arm64: sys/i386/linux/linux_sysent.c references linux_pkey_mprotect/linux_pkey_alloc/linux_pkey_free (syscalls 380-382), the DUMMY() definitions were removed from linux_dummy.c (compiled into linux.ko on i386), and no i386 file defines the wrappers. linux.ko on i386 will have unresolved symbols and fail to load. Add the wrappers to sys/i386/linux/linux_machdep.c next to linux_mprotect.
There was a problem hiding this comment.
Addressed by moving the three wrappers into shared sys/compat/linux/linux_misc.c and removing the amd64/amd64-linux32 duplicates. The shared source is built by native i386 linux.ko, so the generated i386 syscall entries now have definitions. A clean amd64 GENERIC build also confirmed that both linux.ko and linux64.ko export all three symbols.
| linux_pkey_supported(void) | ||
| { | ||
|
|
||
| return ((cpu_stdext_feature2 & CPUID_STDEXT2_OSPKE) != 0); |
There was a problem hiding this comment.
Medium: linux_pkey_supported() only checks CPUID_STDEXT2_OSPKE, which can be stale. identify_cpu_ext_features() re-reads leaf 7 in initializecpu() right after CR4.PKE is set, so OSPKE is latched as 1. Later, fpuinit_bsp1() (sys/amd64/amd64/fpu.c ~line 403) honours the hw.xsave_mask tunable: if PKRU saving was disabled it clears CR4.PKE and CPUID_STDEXT2_PKU, but leaves OSPKE set. On such a boot pkey_alloc() still reports support and calls linux_pkru_write(): the fast path executes wrpkru with CR4.PKE clear (#UD -> kernel panic), and the slow path calls xsave_area_offset(xsave_mask, XFEATURE_ENABLED_PKRU, ...) with PKRU absent from xsave_mask, which trips the KASSERT (or reads xsave_area_desc out of bounds on non-INVARIANTS kernels, since the array is sized by flsl(xsave_mask | xsave_mask_supervisor)). Suggest also requiring CPUID_STDEXT2_PKU and (xsave_mask & XFEATURE_ENABLED_PKRU) != 0 (i.e. the same condition pmap_pkru_check_uargs and the fpu code use), so the Linux side degrades to the no-PKU stubs instead.
There was a problem hiding this comment.
Addressed. linux_pkey_supported() now requires both CPUID_STDEXT2_PKU and CPUID_STDEXT2_OSPKE, plus XFEATURE_ENABLED_PKRU in xsave_mask. If the tunable disables PKRU state saving, the Linux pkey paths now degrade to their unsupported return values without executing WRPKRU or looking up a missing XSAVE component. A clean GENERIC kernel and module build completed successfully.
|
Claude finished @laffer1's task in 3m 43s —— View job Review of PR #522 (re-review)
Blocking: i386 and arm64 Linux modules will fail to build/load
But With the dummy stubs gone and no arch-specific replacement, It's worse than just the pkey syscalls on i386: Fix suggestions:
Other notesThe three issues Sourcery raised earlier (pkey 0 free semantics, protection-vs-tagging ordering in No malloc/free asymmetries, bounds-check gaps, or style(9) issues found in the amd64-specific code ( |
Define the pkey syscall wrappers in shared Linux compatibility code so arm64 and i386 modules resolve their syscall entries. Require active PKU, OSPKE, and PKRU XSAVE state before accessing PKRU. AI-Assisted-by: OpenAI Codex (GPT-5) Signed-off-by: Lucas Holt <luke@foolishgames.com>
|
Claude finished @laffer1's task in 2m 46s —— View job Reviewing PR #522 (re-review requested)
Re-review of the latest commit (
|
| } | ||
| LINUX_PEM_SUNLOCK(pem); | ||
|
|
||
| error = linux_mprotect_common(td, addr, len, prot); |
There was a problem hiding this comment.
[low] pkey_mprotect is not atomic here: linux_mprotect_common() changes the page protections first, and only afterwards does amd64_pkru_update() apply the key tag. If the tagging step fails (e.g. vm_map_check_boundary() rejecting a range that straddles a fixed-boundary entry, or pmap_pkru_set() returning an error), the syscall reports failure to the caller while the new prot has already been applied. On Linux do_mprotect_pkey() performs both under one mmap write lock and either both happen or neither. Callers that treat a failed pkey_mprotect() as "nothing changed" (Chromium checks the return value and falls back to plain mprotect semantics) can end up with a range whose protection changed but which carries the old key. Consider validating the range (vm_map_check_boundary on the same page-rounded bounds) before calling linux_mprotect_common(), or reverting the protection change when tagging fails.
There was a problem hiding this comment.
Confirmed. The fixed-boundary rejection can occur after linux_mprotect_common() has already changed protections. I have added a preflight check under the VM map read lock using the same vm_map_check_boundary() predicate before calling linux_mprotect_common(). amd64_pkru_update() retains its second locked validation to synchronize the actual tag update with concurrent map changes. This avoids the deterministic partial-update case without attempting an unsafe rollback across mappings that may have different prior protections. The remaining PKRU checks are invariant for the call, and pmap_pkru_set()/pmap_pkru_clear() retry allocation failures internally. An incremental amd64 GENERIC kernel and full module build passes, as do formatting and cppcheck. The fix is staged locally and will be pushed in the follow-up commit.
linux_emul.h declares linux_common_execve with struct image_args. Include sys/imgact.h so i386 and arm64 builds see the complete declaration and do not fail -Wvisibility. AI-Assisted-by: Codex GPT-5 Signed-off-by: Lucas Holt <luke@foolishgames.com>
|
Claude finished @laffer1's task in 1m 22s —— View job Review of PR #522 (re-review — latest commit
|
Check fixed-boundary VM entries before changing ordinary page protections so a predictable PKRU tagging rejection cannot leave a partial update. Keep the validation in amd64_pkru_update to cover concurrent map changes. AI-Assisted-by: Codex GPT-5 Signed-off-by: Lucas Holt <luke@foolishgames.com>
|
Claude finished @laffer1's task in 1m 2s —— View job Review of PR #522 (re-review — latest commit
|
| return (xsave_area_desc[idx].offset); | ||
| offs = sizeof(struct savefpu) + sizeof(struct xstate_hdr); | ||
| xstate_bv &= ~(XFEATURE_ENABLED_X87 | XFEATURE_ENABLED_SSE); | ||
| while ((i = ffs(xstate_bv) - 1) > 0 && i < idx) { |
There was a problem hiding this comment.
Low: xstate_bv is uint64_t but ffs() takes int, so any component bit >= 32 is silently dropped from the compact-offset walk. Not reachable today (the only caller passes compact=false, and no enabled component is above bit 31), but the loop is wrong as written: with a high component set in xstate_bv the loop terminates early and returns a too-small offset. ffsl() (or ffsll()) is the intended primitive.
| if (len != 0 && !linux_pkey_range_valid(td, addr, len)) | ||
| return (EINVAL); | ||
|
|
||
| error = linux_mprotect_common(td, addr, len, prot); |
There was a problem hiding this comment.
Low: the protection change and the key tagging are two separate operations with the map unlocked in between, so the syscall is not atomic the way Linux's is. Two consequences: (1) if amd64_pkru_update() fails after linux_mprotect_common() succeeded, the caller gets an error but the page protections were already changed; (2) vm_map_check_boundary() does not verify the range is mapped, so if another thread munmaps part of [addr, addr+len) between the two calls, pmap_pkru_set() still installs the rangeset entry over the hole and a later, unrelated mmap landing there is silently tagged with this pkey (Linux would give the new VMA key 0). Narrow race window, but worth a comment or ordering the tag before/under the same lock as the protect.
Summary
pkey_alloc,pkey_free, andpkey_mprotectValidation
make -j2 buildkernel KERNCONF=GENERIClinux_common.ko,linux64.ko, andlinux.kofpu.oandsys_machdep.owith-Werrorlinux_common.kogit diff --checkThe repository C precommit script was also run. cppcheck reported parser failures in existing IFUNC and macro constructs rather than diagnostics in the new code. Splint skipped the kernel-only C sources by design.
Runtime Brave validation requires installing this kernel and rebooting.
AI-Assisted-by: OpenAI Codex (GPT-5)
Obtained from: FreeBSD commits 7bcaff05223e, b9951017bab3, and bdb561843e86
Tested by: Lucas Holt luke@foolishgames.com
Summary by Sourcery
Add Linux protection-key support backed by native amd64 PKU to enable Chromium and Brave sandboxing.
New Features:
Enhancements:
Build:
Documentation: