Skip to content

feat(maru_vllm): maru-native CUDA kernel for packed GPU↔CXL KV transfers #74

Description

@seohui-XCENA

Background

The packed GPU↔CXL copy path currently resolves its transfer kernels from the optional lmcache package at runtime, with a pure-torch per-layer fallback when they are unavailable (see #72 for the resolution logic). #72 made that resolution robust across upstream LMCache packaging changes, but the coupling remains: any future change on the external side lands on us first, and a direct connector should own its data plane end to end.

This issue tracks bringing the transfer kernel in-tree as a maru-native CUDA C++ kernel.

Goal

A maru-owned CUDA kernel serving the packed store (D2H) and packed load (H2D) paths, so maru_vllm runs kernel-speed transfers with no external kernel resolution. (The opt-in maru_enable_fused_load experiment's single-layer variant, default off, can follow later or be dropped.)

Design

  • Stride-based addressing, no format enum. Generic kernel libraries dispatch over an engine-format enum because they must serve arbitrary engines. maru already resolves the paged-cache axis order and strides at registration (kv_layout.py, fix(maru_vllm): detect the paged KV axis order instead of assuming it #69), so the kernel can take strides as plain arguments. This removes the format-translation layer entirely and stays correct when vLLM changes layouts — only the Python-side layout detection needs to track them.
  • Copy semantics identical to today. Slab side: contiguous pinned CXL slab (cudaHostRegistered mmap) accessed directly from the kernel over UVA. Paged side: per-layer base-pointer table + slot mapping, one launch per chunk-slab covering all layers × K/V.
  • Performance target: parity with the current kernel path. Vectorized (128-bit) accesses where alignment allows. Measured baselines: CXL→GPU ~20 GB/s (32k-token workloads); on the 7k-token cross-instance benchmark, cold 1.062s / warm 0.124s (kernel path) vs 1.374s / 0.135s (per-layer fallback).
  • Integration seam already exists. _resolve_lmc_ops() (fix(maru_vllm): adapt LMCache kernel integration to relocated KV-type enums #72) hands call sites an ops namespace; the native kernel slots in behind the same shape. The per-layer fallback stays as the safety net, and the existing toggle pattern extends to backend selection for A/B during rollout.
  • Build. torch.utils.cpp_extension.CUDAExtension in setup.py, marked optional like the existing _cxl_flush extension: missing nvcc or a failed build degrades to the torch fallback instead of failing the install.

Rollout

  1. Kernel + extension behind a backend flag, default off; A/B against the current kernel path with the existing benchmarks.
  2. Flip the default after correctness (layout-roundtrip unit tests, p2p_example.sh output verification) and perf parity are shown.
  3. Remove the lmcache resolution path (and retire maru_use_lmcache_kernels).

Validation

  • Unit: existing KV-layer roundtrip tests parametrized over the native kernel across all supported layouts (NHD/HND × legacy/0.23 axis orders).
  • E2E: examples/vllm/p2p_sharing/p2p_example.sh (cross-instance hit + output verification).
  • Perf: cross-instance TTFT benchmark vs the baselines above; CXL→GPU effective bandwidth vs 20 GB/s.

Notes

  • If any offset/launch logic is adapted from LMCache's kernels, keep the Apache-2.0 attribution.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions