Repository navigation
Conversation
…c-profiler onto it, and instrument the front-alignment path `feat/batch-admission-ragged` branched before PR #78 merged, so it carries no profiler at all. The question it now has to answer — is the +155.6 ms per forward at K=32 GPU-busy, host-blocked-on-GPU, or host-busy-while-the-GPU-idles — is precisely the one a wall-clock timer cannot answer and the one #78's three channels exist for. Merging master in was tried first and rejected: it compiles to 7 errors because master's scalar `XsRollingCache::trim_tail_to` and this branch's per-row `base`/`tokens` are the same design decision made twice, and resolving that is a change to someone else's workstream, not instrumentation. So the crate is taken verbatim (`arc-profiler/`, unmodified) and the spans are re-placed by hand at this branch's own call sites. What is instrumented, and why each choice: * `normal.rs` — `attach_device`/`maybe_selftest`/`set_meta`. Without `attach_device` every device span resolves to nothing, and a device column of zeros reads exactly like "the GPU was idle". * `engine/mod.rs` — `step` root, `decode` and `prompt` as separate subtrees (a prefill step is orders of magnitude larger; averaging them hides the decode cost), `scheduler.lock`/`scheduler.schedule`, `pipeline.lock`, `pipeline.step`. * `kv_cache/mod.rs` — `clone_in_cache`/`clone_out_cache`, `clone_in.alloc`/`clone_in.slice_set` (device), and the two that discriminate the candidate mechanisms: `cache.front_align` (ONE device span over the whole front-alignment region) and `kv.front_pad_grow` (a HOST span on `front_pad_single`'s padding branch only, so `calls` is literally how many row x slot x half buffers were re-materialised this step). * `deepseek4.rs` — `model_forward`, one device span over the whole forward, as the denominator every host-side cost has to be read against. `front_pad_grow` is a host span on purpose: it runs up to `rows*43*2` times per step, and two `cudaEventRecord`s per call would be a device-side cost of the same order as the thing being measured. The region's stream time is taken once, at `front_align_batch`. Off-state cost is unchanged at ~2.9 ns per call site; nothing here runs unless `ARC_PROFILE=1`. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…t CUDA device
`device::timer_for` refused to attach a CUDA event timer whenever
`cu_stream()` returned NULL:
let stream = dev.cuda_stream().cu_stream() as *mut c_void;
if !stream.is_null() { return Box::new(CudaTimer::new(stream)); }
A NULL `CUstream` is **not** "no stream". It is CUDA's legacy default stream,
and cudarc says so in as many words at `CudaContext::default_stream` — *"the
default stream for this context (the null ptr stream)"* — constructed with
`cu_stream: std::ptr::null_mut()`. A candle device that was never handed an
explicit stream returns exactly that, which is the ordinary configuration.
So the guard rejected the common case, fell through to `NullTimer`, and every
device span in the run recorded nothing. Measured on an H200 tonight, on a V4
profile built `--features "cuda flash-attn"` against a device the report itself
names `Cuda(CudaDevice(DeviceId(1)))`:
note: device selftest: NOT RUN (no CUDA event timer attached)
unresolved_device_spans: 10200
device_ns: 0.0 ms sync_ns: 0.0 ms
⇒ **the three-channel separation PR #78 exists for has never produced a number
on this fork.** wall was real; device and sync were structurally unmeasurable.
The report was honest about it ("unmeasured, not zero", per D18) — which is why
this was findable at all — but it attributed the failure to a missing timer
rather than to a guard that refused a valid stream.
`cudaEventRecord(ev, 0)` is valid and records on the default stream, so the
handle needs no validation. The runtime does, so it is now asked rather than
assumed: probe once with a real `record()`, keep the timer if the event was
created and recorded, and only then fall back to `NullTimer`.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Code Metrics Report━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Language Files Lines Code Comments Blanks ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ C Header 5 305 210 52 43 CSS 2 1181 1036 34 111 CUDA 72 24328 17592 4018 2718 Dockerfile 1 39 22 8 9 JavaScript 16 3546 2676 482 388 Jinja2 7 694 656 5 33 JSON 74 4600 4597 0 3 Makefile 1 6 5 0 1 Metal Shading Lan| 33 12224 9431 1142 1651 PowerShell 1 300 227 30 43 Python 143 14830 12217 797 1816 Shell 23 5756 3964 1412 380 Plain Text 4 3801 0 2479 1322 TOML 33 1485 1292 43 150 YAML 3 25 23 2 0 ───────────────────────────────────────────────────────────────────────────────── HTML 4 2687 2604 43 40 |- CSS 2 543 479 37 27 |- JavaScript 1 1233 1215 12 6 (Total) 4463 4298 92 73 ───────────────────────────────────────────────────────────────────────────────── Jupyter Notebooks 4 122 83 23 16 |- Markdown 1 60 30 22 8 |- Python 1 122 113 1 8 (Total) 304 226 46 32 ───────────────────────────────────────────────────────────────────────────────── Markdown 189 37703 0 28931 8772 |- BASH 70 1618 1189 313 116 |- C 2 12 12 0 0 |- CUDA 2 84 56 16 12 |- JSON 18 708 708 0 0 |- PowerShell 1 1 1 0 0 |- Python 23 1008 787 113 108 |- Rust 65 2048 1713 77 258 |- TOML 6 207 164 0 43 |- YAML 4 38 33 5 0 (Total) 43427 4663 29455 9309 ───────────────────────────────────────────────────────────────────────────────── Rust 662 307804 266560 13893 27351 |- Markdown 477 22853 471 19669 2713 (Total) 330657 267031 33562 30064 ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Total 1277 451971 330166 73659 48146 ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ |
…regated Two device spans in `forward_4d` — `mla_attn` around `self.attn.forward` and `moe` around `self.moe_or_mlp.forward` — aggregated across all layers. Enough to answer the question actually being asked (what fraction of a decode step grows with B); deliberately NOT the fourteen MLA sub-spans master carries, which this branch predates and which would each fire 43x per step. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
ArcGate: HELD — parked above #104, and the reason is now specific rather than open-ended. Full analysis on #104. Short form: #104 and #122 (landed as master
#104's Two things that are not true, both checked rather than assumed:
This is scope, not a verdict. Nothing here is ranked down or closeable; the blocker is now one A/B measurement — does #104's design buy anything #122's does not, priced against its per-step rebuild. Also note: this PR currently shows 1 check, the comment bot — zero CI lanes have ever run on it, because it predates the stacked-PR trigger fix that is now on master. Any push starts the full 16 lanes. Nothing deleted; branch untouched. |
|
Already landed on master (superseded, in a strictly richer form) — this PR should be closed. Checked the actual content, not the branch names:
Rebasing onto The only thing left on this branch that is not on master is its inherited base commit |
Closing: superseded — this content is already on master, by another route.This PR was one of five based on another PR's branch ( On inspection it does not need a rebase — it needs closing.
The second commit's content is on master too.
The remaining Reopen if any of the above is wrong — nothing here is deleted. |
feat/batch-admission-raggedbranched before PR #78 merged, so it carries noprofiler at all — and the question it now has to answer is exactly the one a
wall-clock timer cannot: the measured +155.6 ms per forward at K=32 (and
+0.0 ms at K=1,
mean_batchidentical across arms) is either GPU-busy,host-blocked-on-GPU, or host-busy-while-the-GPU-idles. Those three look
identical from outside and have completely different fixes. Separating them is
the entire reason PR #78 exists.
Why not merge master
Tried first, rejected.
git merge origin/mastertextually auto-merges butcompiles to 7 errors: master's scalar
XsRollingCache::trim_tail_toandthis branch's per-row
base/tokensare the same design decision made twice,and reconciling them is a change to someone else's workstream, not
instrumentation. So
arc-profiler/is taken verbatim from master and thespans are re-placed by hand at this branch's own call sites.
What is instrumented, and why each choice
normal.rsattach_device/maybe_selftest/set_metaattach_deviceevery device span resolves to nothing, and a device column of zeros reads exactly like "the GPU was idle"engine/mod.rssteproot,decode,prompt,scheduler.lock,scheduler.schedule,pipeline.lock,pipeline.stepkv_cache/mod.rsclone_in_cache,clone_out_cache,clone_in.alloc,clone_in.slice_setkv_cache/mod.rscache.front_align(device),kv.front_pad_grow(host)deepseek4.rsmodel_forward(device)kv.front_pad_growwraps onlyfront_pad_single's padding branch, so itscallsis literally how many row × slot × half buffers were re-materialisedthis step — the count that separates "thousands of tiny allocations" from "one
large copy" without any timing argument.
It is a host span on purpose: it runs up to
rows × 43 × 2times per step,and two
cudaEventRecords per call would be a device-side cost of the sameorder as the thing being measured. The region's stream time is taken once,
at
front_align_batch, sodevice_nsthere is the GPU time of all the paddingand
wall − sync − deviceis the host time spent issuing it.Off-state cost is unchanged at ~2.9 ns per call site; nothing here runs unless
ARC_PROFILE=1.Provenance
Measured with
arc-tools-style gating: the package that produces the binary isthe package built (
-p mistralrs-cli→mistralrs), the path comes fromcargo's own artifact stream rather than
target/release/<guess>, the binary isasserted to carry the built SHA, and the running server must report that same
git revisionbefore any cell is driven.Measured on an H200 (2026-08-17),
qtip2bUQFF, MTP depth 3,--max-seqs 256The instrument first, because a profiler reporting from dark instrumentation
is this project's worst recurring fault. Every cell below carries
device selftest: ... ratio=24.6x–44.2x — PASS: CUDA events are measuring execution, not launches,unresolved_device_spans: 0,violations: [],misnested_spans: 0. Cost of running it: 58.266 tok/s profiled vs 58.348unprofiled at K=32, i.e. −0.14%.
Where a decode step goes (uniform lengths, admission OFF)
model_forward(device)This regime is GPU-bound, not host-bound. Fitting the two rows gives
step ≈ 165 ms + 17.3 ms × B, which predicts the independently-measured b=12cell (372 vs 355 ms) and b=25 cell (596 vs 629 ms). The marginal cost of one
more sequence is 17.3 ms of GPU stream time; the host adds 0.7 ms.
What the two candidate mechanisms actually cost (mixed lengths, K=32)
clone_in_cache(mechanism B)slice_setcache.front_align(mechanism A)front_pad_grow@ 6.9 µsBoth are real, both engage exactly as predicted —
front_pad_growrisesfrom 133 calls/step on uniform lengths to 1,313 on mixed, so
callstracksraggedness — and both are GPU stream time, not host allocation stalls.
Together they are 37.6 ms against a 629 ms step (6%), ~97% of it device.
What ragged admission is worth here
On a uniform workload at temperature 0 every row decodes the identical
trajectory, so nothing is ragged, the grant costs nothing and buys nothing —
the two arms produce bit-identical MTP telemetry. On mixed lengths the
refused arm runs 12 of 32 admitted sequences per step, which is the
bucket-shattering law measured on hardware.
Prefill is the larger term, and it is not a K=128 phenomenon
12,160 prompt tokens in 43.2 s is ~281 tok/s of prefill. There is no chunked
prefill, so admitting K prompts blocks every running sequence for
K × prompt_len ÷ ~300seconds.What actually grows with B (uniform lengths, one bucket, admission OFF)
Same binary, same instrument, same step count, one run:
model_forwarddevmla_attndevmoedevMarginal cost of one more sequence, fitted over B=1→32: 19.4 ms/step, split
mla_attn11.0 ms — 57%, device time;moe4.2 ms — 21%, device time;Extrapolated fixed cost at B→0 is ~94 ms, so ~87% of a B=32 decode step is
work proportional to B, and 84% of that is on the GPU. The step-rate gain
from B=1 to B=8 is 3.1×, and B=1 to B=32 is 5.1×.
clone_in_cachedoes not appear in the uniform decode path at all: with astable cohort the engine issues
CacheInstruction::Nothing, so the 86allocations per step are never made. It costs 28.3 ms/step (4% of a B=32 step)
only when the cohort changes, or when ragged admission forces re-assembly every
step. That is what the fix for it is worth — measured, not assumed.