Background
When a request's KV cache is already in the CXL pool, vLLM skips prefill and loads that KV into GPU memory instead — 4.19 GB of it for a 32k-token prompt on Llama-3.1-8B. The time that load takes is what the client sees as time-to-first-token (TTFT), and the loads are served one at a time, so concurrent requests queue behind each other.
Setup. One node, two GPUs, meta-llama/Llama-3.1-8B-Instruct. One vLLM instance writes 16 documents of ~32k tokens into the CXL pool and a second reads them back, so every measured request is a full cache hit and its TTFT is almost entirely load time. Commands under Reproduction.
Observation
Two things were swept: how many requests are in flight at once (max_inflight, 1 or 8), and how many tokens each request generates (output_len, 1 through 512).
With one request at a time, TTFT does not move — 254–262 ms whether it generates 1 token or 512. Nothing queues, so output length has nothing to act on.
With 8 in flight, output length decides TTFT. The 32 measured requests split in two: the first 8 are released together onto an empty loader, and the remaining 24 enter one at a time as earlier ones finish.
| arm |
output_len |
first 8 reqs |
remaining 24 |
vllm-maruconnector |
1 |
939 ms |
1,407 ms |
vllm-maruconnector |
50 |
1,035 ms |
285 ms |
lmcache-maru (MP) |
1 |
936 ms |
1,381 ms |
lmcache-maru (MP) |
50 |
1,265 ms |
487 ms |
The first 8 cost about the same either way. The whole difference is in the remaining 24: asking for 50 tokens gets a first token in 285 ms, asking for 1 token waits 1,407 ms — five times longer for less work.
Per-request timelines below, vllm-maruconnector first and lmcache-maru second. Each bar is one request: solid is TTFT, hatched is decode. The red dashed line is the max_inflight boundary.
Output length is what hides this. With 50-token outputs a request decodes long enough that the loader clears before the next one needs it. With 1-token outputs there is no decode phase, so all 8 sit on the loader at once and each waits for seven others. A benchmark run at a single output length can miss the effect entirely.
Both connector paths land on the same number. 1,407 ms direct against 1,381 ms through LMCache MP — 2% apart, from two independently written load implementations. That points at a bottleneck shared underneath both rather than a mistake in either one.
Reproduction
Servers
MaruServer — metadata only, the KV itself moves over the mapped DAX slabs:
python3 -m maru_server --port 22000 --log-level INFO \
--dax-path /dev/dax1.0 --dax-path /dev/dax2.0
Two vLLM instances on the same node, one GPU each. inst1 (CUDA_VISIBLE_DEVICES=0) stores, inst2 (CUDA_VISIBLE_DEVICES=1) queries; only maru_instance_id and the port differ between them.
CUDA_VISIBLE_DEVICES=1 vllm serve meta-llama/Llama-3.1-8B-Instruct \
--port 20011 \
--gpu-memory-utilization 0.85 \
--no-enable-prefix-caching \
--kv-transfer-config '{
"kv_connector": "MaruKVConnector",
"kv_connector_module_path": "maru_vllm",
"kv_role": "kv_both",
"kv_buffer_device": "cuda",
"kv_buffer_size": 1000000000.0,
"kv_connector_extra_config": {
"maru_server_url": "tcp://localhost:22000",
"maru_instance_id": "inst2",
"maru_kv_chunk_tokens": 256,
"maru_pool_size": "107374182400",
"maru_async_load": true,
"maru_async_store": true,
"maru_load_admission_window": 1,
"maru_overlap_load_with_compute": true,
"maru_log_timing": true
}}'
--no-enable-prefix-caching is load-bearing: without it vLLM's own prefix cache absorbs the repeats and the loads under measurement never happen.
The lmcache-maru rows are the same two instances running LMCacheMPConnector against an MP server backed by the same DAX devices, in place of MaruKVConnector.
Workload
16 documents, none of them a prefix of another (the leading id is what guarantees that):
prompt = f"{doc_id} " + " ".join(["hi"] * 32000) # ~32k tokens -> 4.19 GB of KV
Three phases, in order:
- pre-warmup — 5 throwaway prompts,
f"{i}xx " + " ".join(["hi"] * 1000), not recorded.
- warmup round on
inst1 — 16 requests, one per document, populating the CXL pool.
- query round on
inst2 — the same 16 documents tiled twice, 32 requests, every one a full hit. The charts above are this round only.
The dispatcher is a closed loop:
sem = asyncio.Semaphore(max_inflight)
tasks = [asyncio.create_task(one(prompt, sem)) for prompt in prompts]
await asyncio.gather(*tasks)
Every task exists up front, so exactly max_inflight requests are outstanding at all times and a new one is admitted the instant one finishes. At max_inflight=8 the first 8 arrive together and every request after them waits behind a full queue.
Knobs that produce the effect
| knob |
values |
why it matters |
max_inflight |
1, 8 |
1 is the control — it never queues. |
output_len |
1, 50, 128, 256, 512 |
Time spent decoding is what holds a request off the loader. Short output removes that buffer. |
ignore_eos |
must be true |
Without it the model stops at EOS, every point in the sweep measures the same decode, and the effect never appears at all. |
Requests go through the completions API with temperature left at its default.
For readers inside the org: the harness configs are cfg/p2p/long_doc_decode_len_{1,50,128,256,512}.yaml, run with ./run.sh <config>, and tools/plot_outputlen_compare.py --backend <arm> redraws the charts above.
Notes
- Not a correctness bug. Every request succeeds and returns the right output; this is about how latency scales.
Background
When a request's KV cache is already in the CXL pool, vLLM skips prefill and loads that KV into GPU memory instead — 4.19 GB of it for a 32k-token prompt on Llama-3.1-8B. The time that load takes is what the client sees as time-to-first-token (TTFT), and the loads are served one at a time, so concurrent requests queue behind each other.
Setup. One node, two GPUs,
meta-llama/Llama-3.1-8B-Instruct. One vLLM instance writes 16 documents of ~32k tokens into the CXL pool and a second reads them back, so every measured request is a full cache hit and its TTFT is almost entirely load time. Commands under Reproduction.Observation
Two things were swept: how many requests are in flight at once (
max_inflight, 1 or 8), and how many tokens each request generates (output_len, 1 through 512).With one request at a time, TTFT does not move — 254–262 ms whether it generates 1 token or 512. Nothing queues, so output length has nothing to act on.
With 8 in flight, output length decides TTFT. The 32 measured requests split in two: the first 8 are released together onto an empty loader, and the remaining 24 enter one at a time as earlier ones finish.
output_lenvllm-maruconnectorvllm-maruconnectorlmcache-maru(MP)lmcache-maru(MP)The first 8 cost about the same either way. The whole difference is in the remaining 24: asking for 50 tokens gets a first token in 285 ms, asking for 1 token waits 1,407 ms — five times longer for less work.
Per-request timelines below,
vllm-maruconnectorfirst andlmcache-marusecond. Each bar is one request: solid is TTFT, hatched is decode. The red dashed line is themax_inflightboundary.Output length is what hides this. With 50-token outputs a request decodes long enough that the loader clears before the next one needs it. With 1-token outputs there is no decode phase, so all 8 sit on the loader at once and each waits for seven others. A benchmark run at a single output length can miss the effect entirely.
Both connector paths land on the same number. 1,407 ms direct against 1,381 ms through LMCache MP — 2% apart, from two independently written load implementations. That points at a bottleneck shared underneath both rather than a mistake in either one.
Reproduction
Servers
MaruServer — metadata only, the KV itself moves over the mapped DAX slabs:
Two vLLM instances on the same node, one GPU each.
inst1(CUDA_VISIBLE_DEVICES=0) stores,inst2(CUDA_VISIBLE_DEVICES=1) queries; onlymaru_instance_idand the port differ between them.--no-enable-prefix-cachingis load-bearing: without it vLLM's own prefix cache absorbs the repeats and the loads under measurement never happen.The
lmcache-marurows are the same two instances runningLMCacheMPConnectoragainst an MP server backed by the same DAX devices, in place ofMaruKVConnector.Workload
16 documents, none of them a prefix of another (the leading id is what guarantees that):
Three phases, in order:
f"{i}xx " + " ".join(["hi"] * 1000), not recorded.inst1— 16 requests, one per document, populating the CXL pool.inst2— the same 16 documents tiled twice, 32 requests, every one a full hit. The charts above are this round only.The dispatcher is a closed loop:
Every task exists up front, so exactly
max_inflightrequests are outstanding at all times and a new one is admitted the instant one finishes. Atmax_inflight=8the first 8 arrive together and every request after them waits behind a full queue.Knobs that produce the effect
max_inflight1is the control — it never queues.output_lenignore_eosRequests go through the completions API with
temperatureleft at its default.For readers inside the org: the harness configs are
cfg/p2p/long_doc_decode_len_{1,50,128,256,512}.yaml, run with./run.sh <config>, andtools/plot_outputlen_compare.py --backend <arm>redraws the charts above.Notes