You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
The packed GPU↔CXL copy path currently resolves its transfer kernels from the optional lmcache package at runtime, with a pure-torch per-layer fallback when they are unavailable (see #72 for the resolution logic). #72 made that resolution robust across upstream LMCache packaging changes, but the coupling remains: any future change on the external side lands on us first, and a direct connector should own its data plane end to end.
This issue tracks bringing the transfer kernel in-tree as a maru-native CUDA C++ kernel.
Goal
A maru-owned CUDA kernel serving the packed store (D2H) and packed load (H2D) paths, so maru_vllm runs kernel-speed transfers with no external kernel resolution. (The opt-in maru_enable_fused_load experiment's single-layer variant, default off, can follow later or be dropped.)
Design
Stride-based addressing, no format enum. Generic kernel libraries dispatch over an engine-format enum because they must serve arbitrary engines. maru already resolves the paged-cache axis order and strides at registration (kv_layout.py, fix(maru_vllm): detect the paged KV axis order instead of assuming it #69), so the kernel can take strides as plain arguments. This removes the format-translation layer entirely and stays correct when vLLM changes layouts — only the Python-side layout detection needs to track them.
Copy semantics identical to today. Slab side: contiguous pinned CXL slab (cudaHostRegistered mmap) accessed directly from the kernel over UVA. Paged side: per-layer base-pointer table + slot mapping, one launch per chunk-slab covering all layers × K/V.
Performance target: parity with the current kernel path. Vectorized (128-bit) accesses where alignment allows. Measured baselines: CXL→GPU ~20 GB/s (32k-token workloads); on the 7k-token cross-instance benchmark, cold 1.062s / warm 0.124s (kernel path) vs 1.374s / 0.135s (per-layer fallback).
Integration seam already exists._resolve_lmc_ops() (fix(maru_vllm): adapt LMCache kernel integration to relocated KV-type enums #72) hands call sites an ops namespace; the native kernel slots in behind the same shape. The per-layer fallback stays as the safety net, and the existing toggle pattern extends to backend selection for A/B during rollout.
Build.torch.utils.cpp_extension.CUDAExtension in setup.py, marked optional like the existing _cxl_flush extension: missing nvcc or a failed build degrades to the torch fallback instead of failing the install.
Rollout
Kernel + extension behind a backend flag, default off; A/B against the current kernel path with the existing benchmarks.
Flip the default after correctness (layout-roundtrip unit tests, p2p_example.sh output verification) and perf parity are shown.
Remove the lmcache resolution path (and retire maru_use_lmcache_kernels).
Validation
Unit: existing KV-layer roundtrip tests parametrized over the native kernel across all supported layouts (NHD/HND × legacy/0.23 axis orders).
E2E: examples/vllm/p2p_sharing/p2p_example.sh (cross-instance hit + output verification).
Perf: cross-instance TTFT benchmark vs the baselines above; CXL→GPU effective bandwidth vs 20 GB/s.
Notes
If any offset/launch logic is adapted from LMCache's kernels, keep the Apache-2.0 attribution.
Background
The packed GPU↔CXL copy path currently resolves its transfer kernels from the optional
lmcachepackage at runtime, with a pure-torch per-layer fallback when they are unavailable (see #72 for the resolution logic). #72 made that resolution robust across upstream LMCache packaging changes, but the coupling remains: any future change on the external side lands on us first, and a direct connector should own its data plane end to end.This issue tracks bringing the transfer kernel in-tree as a maru-native CUDA C++ kernel.
Goal
A maru-owned CUDA kernel serving the packed store (D2H) and packed load (H2D) paths, so
maru_vllmruns kernel-speed transfers with no external kernel resolution. (The opt-inmaru_enable_fused_loadexperiment's single-layer variant, default off, can follow later or be dropped.)Design
kv_layout.py, fix(maru_vllm): detect the paged KV axis order instead of assuming it #69), so the kernel can take strides as plain arguments. This removes the format-translation layer entirely and stays correct when vLLM changes layouts — only the Python-side layout detection needs to track them.cudaHostRegistered mmap) accessed directly from the kernel over UVA. Paged side: per-layer base-pointer table + slot mapping, one launch per chunk-slab covering all layers × K/V._resolve_lmc_ops()(fix(maru_vllm): adapt LMCache kernel integration to relocated KV-type enums #72) hands call sites anopsnamespace; the native kernel slots in behind the same shape. The per-layer fallback stays as the safety net, and the existing toggle pattern extends to backend selection for A/B during rollout.torch.utils.cpp_extension.CUDAExtensioninsetup.py, marked optional like the existing_cxl_flushextension: missing nvcc or a failed build degrades to the torch fallback instead of failing the install.Rollout
p2p_example.shoutput verification) and perf parity are shown.maru_use_lmcache_kernels).Validation
examples/vllm/p2p_sharing/p2p_example.sh(cross-instance hit + output verification).Notes