Skip to content

perf(load): CXL→GPU load path does not scale with in-flight requests — TTFT ≈ N × 176 ms when decode is short #79

Description

@kihwan-XCENA

Background

When a request's KV cache is already in the CXL pool, vLLM skips prefill and loads that KV into GPU memory instead — 4.19 GB of it for a 32k-token prompt on Llama-3.1-8B. The time that load takes is what the client sees as time-to-first-token (TTFT), and the loads are served one at a time, so concurrent requests queue behind each other.

Setup. One node, two GPUs, meta-llama/Llama-3.1-8B-Instruct. One vLLM instance writes 16 documents of ~32k tokens into the CXL pool and a second reads them back, so every measured request is a full cache hit and its TTFT is almost entirely load time. Commands under Reproduction.

Observation

Two things were swept: how many requests are in flight at once (max_inflight, 1 or 8), and how many tokens each request generates (output_len, 1 through 512).

With one request at a time, TTFT does not move — 254–262 ms whether it generates 1 token or 512. Nothing queues, so output length has nothing to act on.

With 8 in flight, output length decides TTFT. The 32 measured requests split in two: the first 8 are released together onto an empty loader, and the remaining 24 enter one at a time as earlier ones finish.

arm output_len first 8 reqs remaining 24
vllm-maruconnector 1 939 ms 1,407 ms
vllm-maruconnector 50 1,035 ms 285 ms
lmcache-maru (MP) 1 936 ms 1,381 ms
lmcache-maru (MP) 50 1,265 ms 487 ms

The first 8 cost about the same either way. The whole difference is in the remaining 24: asking for 50 tokens gets a first token in 285 ms, asking for 1 token waits 1,407 ms — five times longer for less work.

Per-request timelines below, vllm-maruconnector first and lmcache-maru second. Each bar is one request: solid is TTFT, hatched is decode. The red dashed line is the max_inflight boundary.

vllm-maruconnector at max_inflight=8, output_len 1 vs 50. Left panel: every request keeps a full-length TTFT bar of about 1,407 ms. Right panel: from the 9th request on the TTFT segment collapses to about 285 ms and decode fills the rest of each bar. lmcache-maru (LMCache MP plus maru) on the same axes: the same collapse from the 9th request on, and a first wave whose completion order is more scrambled than the direct path's.

Output length is what hides this. With 50-token outputs a request decodes long enough that the loader clears before the next one needs it. With 1-token outputs there is no decode phase, so all 8 sit on the loader at once and each waits for seven others. A benchmark run at a single output length can miss the effect entirely.

Both connector paths land on the same number. 1,407 ms direct against 1,381 ms through LMCache MP — 2% apart, from two independently written load implementations. That points at a bottleneck shared underneath both rather than a mistake in either one.

Reproduction

Servers

MaruServer — metadata only, the KV itself moves over the mapped DAX slabs:

python3 -m maru_server --port 22000 --log-level INFO \
  --dax-path /dev/dax1.0 --dax-path /dev/dax2.0

Two vLLM instances on the same node, one GPU each. inst1 (CUDA_VISIBLE_DEVICES=0) stores, inst2 (CUDA_VISIBLE_DEVICES=1) queries; only maru_instance_id and the port differ between them.

CUDA_VISIBLE_DEVICES=1 vllm serve meta-llama/Llama-3.1-8B-Instruct \
  --port 20011 \
  --gpu-memory-utilization 0.85 \
  --no-enable-prefix-caching \
  --kv-transfer-config '{
    "kv_connector": "MaruKVConnector",
    "kv_connector_module_path": "maru_vllm",
    "kv_role": "kv_both",
    "kv_buffer_device": "cuda",
    "kv_buffer_size": 1000000000.0,
    "kv_connector_extra_config": {
      "maru_server_url": "tcp://localhost:22000",
      "maru_instance_id": "inst2",
      "maru_kv_chunk_tokens": 256,
      "maru_pool_size": "107374182400",
      "maru_async_load": true,
      "maru_async_store": true,
      "maru_load_admission_window": 1,
      "maru_overlap_load_with_compute": true,
      "maru_log_timing": true
    }}'

--no-enable-prefix-caching is load-bearing: without it vLLM's own prefix cache absorbs the repeats and the loads under measurement never happen.

The lmcache-maru rows are the same two instances running LMCacheMPConnector against an MP server backed by the same DAX devices, in place of MaruKVConnector.

Workload

16 documents, none of them a prefix of another (the leading id is what guarantees that):

prompt = f"{doc_id} " + " ".join(["hi"] * 32000)   # ~32k tokens -> 4.19 GB of KV

Three phases, in order:

  1. pre-warmup — 5 throwaway prompts, f"{i}xx " + " ".join(["hi"] * 1000), not recorded.
  2. warmup round on inst1 — 16 requests, one per document, populating the CXL pool.
  3. query round on inst2 — the same 16 documents tiled twice, 32 requests, every one a full hit. The charts above are this round only.

The dispatcher is a closed loop:

sem = asyncio.Semaphore(max_inflight)
tasks = [asyncio.create_task(one(prompt, sem)) for prompt in prompts]
await asyncio.gather(*tasks)

Every task exists up front, so exactly max_inflight requests are outstanding at all times and a new one is admitted the instant one finishes. At max_inflight=8 the first 8 arrive together and every request after them waits behind a full queue.

Knobs that produce the effect

knob values why it matters
max_inflight 1, 8 1 is the control — it never queues.
output_len 1, 50, 128, 256, 512 Time spent decoding is what holds a request off the loader. Short output removes that buffer.
ignore_eos must be true Without it the model stops at EOS, every point in the sweep measures the same decode, and the effect never appears at all.

Requests go through the completions API with temperature left at its default.

For readers inside the org: the harness configs are cfg/p2p/long_doc_decode_len_{1,50,128,256,512}.yaml, run with ./run.sh <config>, and tools/plot_outputlen_compare.py --backend <arm> redraws the charts above.

Notes

  • Not a correctness bug. Every request succeeds and returns the right output; this is about how latency scales.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions