Skip to content

perf(arckv): make candle's caching allocator reachable, on for decode, and BOUNDED - #177

Merged
heydryft merged 18 commits into
masterfrom
perf/alloc-cache-reachable-on-decode
Aug 20, 2026
Merged

heydryft merged 18 commits into
masterfrom
perf/alloc-cache-reachable-on-decode

Conversation

@heydryft

Copy link
Copy Markdown
Contributor

Errors in the brief this had to correct first

Two premises I was given did not survive checking, and both change what "just turn it on" means:

  1. "with the correctness tests it already has" — there are none. The candle fork (aeonmindai/candle rev 88d86a2, which is what Cargo.lock pins) has zero tests for the allocator: cuda_backend/device.rs and cuda_backend/mod.rs are the only files that mention it and neither carries a #[cfg(test)]. Every arc-side exercise is #[cfg(feature = "cuda")] and unrunnable in CI. The only executable one is the capture_probe example, which has no assertions — it prints a max-abs-diff for a human to read.
  2. The measured pairs are not in the tree. host calls 25,622 -> 737, layout H2D 2,361.6 -> 6.0, driver frees exactly 0, bit-identical 52/52 — none of the "after" numbers exist on any of the 602 refs. 25,622 appears once, as a pre-arena baseline in COMPETITIVE_TEARDOWN.md:201-215. The nearest 52 is KERNEL_RULES.md:977-984, and it is a cautionary record, not a success: an arena that reported accounting OK, 4,976 cache hits/step and bit-identical output over 52 steps while silently bypassing itself for every buffer under 128 bytes. The only tell was a driver_frees that was supposed to be impossible.

So the arena is not a measured, tested thing waiting for a default flip. Consequently this PR flips what is safe to flip and says exactly what a GPU run must falsify.

The defect that is real

if probe && seq_len == 1 && std::env::var_os("ARC_CANDLE_ALLOC_CACHE").is_some()

The first conjunct tied a general-purpose caching allocator to the V4 capture probe. Recycling a decode step's frees has nothing to do with capture; it is worth having on its own. Three stacked default-off gates meant the ~11k allocations per token were a disabled feature, not a missing one.

What lands

A pure policy, alloc_cache_action(seq_len, enabled, killed):

forward action
decode (seq_len == 1) Enable
prefill (seq_len != 1) DrainAndDisable
already in that state Leave

ARC_CANDLE_ALLOC_CACHE=0 is the kill switch. Any other value, and unset, leave the policy in force — so the =1 the ops scripts pass (arcgraph_heap_probe.sh:191) still means what it always meant.

The capture probe keeps its own gate for the graph-mode positions. Only the allocator moved out from under it.

Why prefill must drain rather than coast

Not incidental, and it is the part I would have got wrong without reading the fork. The cache is keyed on exact byte count — free: HashMap<usize, Vec<CUdeviceptr>>, looked up with free.get_mut(&bytes). No bucketing, no smallest-fit, no splitting. It also has no capacity bound and no eviction. Leaving it on across a prefill parks that prefill's large one-shot buffers for the process lifetime under a key nothing will ever request again — which is the OOM-on-a-tight-model failure ARC_NO_DEDICATED_DECODE already exists to work around.

Why decode is the right place for it

A decode step allocates the same shapes 43 times over, once per layer, and the shapes that do not depend on KV length — hidden states, MLP and expert intermediates — repeat step after step. Exact-size keying is a good fit for exactly that traffic and a bad fit for everything else.

The number a GPU run should falsify

Predicted, not measured — I have no GPU. Buffers whose size tracks KV length (the causal mask most obviously) change size every decode step, so each step files one more never-reused entry. That is O(context^2) bytes worst case and this policy does not bound it: a [1,1,1,kv] BF16 mask summed over 4k tokens is ~16 MB, and over 128k it is not 16 MB.

The fixed-capacity graph-mode path (deepseek4.rs:4344-4353, which swaps the growing causal mask for graph_mode_length_mask at cfg_full.sliding_window) removes that growth entirely — but it is reached only under ARC_V4_CAPTURE_PROBE. So the allocator's steady-state safety is downstream of the shape-invariance work another agent owns, and the measurement that settles it is the long-context high-water mark. Until then, ARC_CANDLE_ALLOC_CACHE=0.

Allocations removed per token

Unmeasured. I will not put a number here. The mechanism removes a cuMemFreeAsync/cuMemAllocAsync pair for every decode allocation whose exact byte size recurs, against a baseline of 11,436 allocations per token. What fraction of those recur is precisely what has never been measured on a build where the cache was actually on — which is the point of this PR.

Evidence

Six host-runnable tests for the policy — the first tests this subsystem has anywhere. cargo test -p mistralrs-core --lib alloc_cache_policy: 6 passed; 0 failed.

🤖 Generated with Claude Code

https://claude.ai/code/session_01SpVNMpb13HkUXqSqbN1o9H

@github-actions

github-actions Bot commented Aug 19, 2026 •

Copy link
Copy Markdown
Code Metrics Report
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
 Language              Files        Lines         Code     Comments       Blanks
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
 C Header                  5          305          210           52           43
 CSS                       2         1181         1036           34          111
 CUDA                     74        25734        18232         4703         2799
 Dockerfile                1           39           22            8            9
 JavaScript               16         3546         2676          482          388
 Jinja2                    7          694          656            5           33
 JSON                     74         4600         4597            0            3
 Makefile                  1            6            5            0            1
 Metal Shading Lan|       33        12224         9431         1142         1651
 PowerShell                1          300          227           30           43
 Python                  145        15139        12482          811         1846
 Shell                    40         9948         6637         2655          656
 Plain Text                4         3801            0         2479         1322
 TOML                     33         1498         1294           54          150
 YAML                      3           25           23            2            0
─────────────────────────────────────────────────────────────────────────────────
 HTML                      4         2687         2604           43           40
 |- CSS                    2          543          479           37           27
 |- JavaScript             1         1233         1215           12            6
 (Total)                             4463         4298           92           73
─────────────────────────────────────────────────────────────────────────────────
 Jupyter Notebooks         4          122           83           23           16
 |- Markdown               1           60           30           22            8
 |- Python                 1          122          113            1            8
 (Total)                              304          226           46           32
─────────────────────────────────────────────────────────────────────────────────
 Markdown                204        45330            0        35183        10147
 |- BASH                  72         1655         1203          331          121
 |- C                      3           17           17            0            0
 |- CUDA                   2           84           56           16           12
 |- JSON                  18          708          708            0            0
 |- PowerShell             1            1            1            0            0
 |- Python                23         1008          787          113          108
 |- Rust                  66         2051         1716           77          258
 |- TOML                   6          207          164            0           43
 |- YAML                   5           41           36            5            0
 (Total)                            51102         4688        35725        10689
─────────────────────────────────────────────────────────────────────────────────
 Rust                    673       331228       285444        16819        28965
 |- Markdown             491        29538          471        25464         3603
 (Total)                           360766       285915        42283        32568
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
 Total                  1324       495675       352655        90603        52417
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

@heydryft

Copy link
Copy Markdown
Contributor Author

MEASURED on an H200 — the win is real, and so is the leak this PR predicted

Box: H200 143,771 MiB, exclusive. Binary 2df0eeb2ae59… = this PR's commit
cherry-picked onto graph @ 0875caa (master + the fused E4M3 kernel), so the
baseline below is the same tree minus this one commit.

Instrument: untraced bench for tok/s; nsys --trace=cuda for the counts, with
the step count pinned from the trace (the 25 s window holds 346 / 462 steps, so
neither leg is the "window caught no decode" void).

The counts — this is what the PR asked to be asserted

per decode step before after
cuMemAllocAsync 10,114.68 (7.450 ms host) 131.44 (0.384 ms host)
cuMemFreeAsync 10,114.69 (7.428 ms host) 0.00
both, host time 14.878 ms/step 0.384 ms/step
launches 7,935.9 7,924.1
kernel time 31.266 ms 31.477 ms
GPU-busy (kernel union) 43.3 % 58.2 %
blocking D2H 1.00 1.00

cuMemFreeAsync reaching exactly 0.00 is the load-bearing number: every free
is now recycled rather than returned to the driver. Kernel time is unchanged
(31.27 → 31.48 ms), which is the control — the GPU work is identical, only host
overhead moved.

Throughput

tok/s ms/token
baseline (0875caa) 19.3 51.85
this PR 20.8 48.03
this PR + ARC_CANDLE_ALLOC_CACHE=0 19.2 52.06

The kill-switch leg is the same binary with one env var, and it lands back on
baseline. So the +7.4 % is attributable to the allocator and to nothing else
in the rebuild.

25,622 → 737 does not reproduce and I did not try to make it — these are
freshly measured, and the "before" is measured too.

The number this PR asked a GPU run to falsify — it does not survive

The PR's own text: "each step files one more never-reused entry … that is
O(context^2) bytes worst case and this policy does not bound it."
Measured,
gen-len 2500, memory.used sampled at 2 Hz, same binary both legs:

cache OFF   t=70s 79,043 MiB ─ 79,043 ─ 79,043 ─ 79,075 ─ 79,107 ─ t=196s 79,139
            FLAT: +96 MiB over 128 s of decode

cache ON    t=84s 81,475 MiB ─ 83,011 ─ 84,579 ─ 86,147 ─ 89,315 ─ t=196s 101,493
            LINEAR, no plateau: +20,018 MiB over 112 s = +178.7 MiB/s

At 20.8 tok/s that is ~8.6 MiB per decoded token, unbounded. From the 101,493
MiB reached at 2,500 tokens there is 42,278 MiB of headroom left, so a single
sequence OOMs at roughly 7,400 tokens on a 143 GB H200.

The PR estimated ~16 MB at 4k tokens from the causal mask alone. The measurement
is ~21 GB at 2.5k, so the mask is not the story — many buffers are KV-length-keyed
and every one of them files a fresh never-reused byte-size key each step.

Note also that DrainAndDisable cannot rescue this in a decode-only workload:
the drain fires on prefill, and a long generation has exactly one of those, at
the start. In a mixed serving loop each prefill would drain — bounding the growth
but also discarding the decode buffers that produce the win.

Verdict

The mechanism is right and the 14.5 ms/step is real. Do not merge with the cache
on by default
— as measured it converts a 7.4 % speedup into an OOM at ~7.4k
tokens. Either bound the cache (cap + LRU eviction, or bucket the key instead of
exact-byte), or gate the default on the shape-invariant decode path that stops
KV-length-keyed allocations, and keep the unbounded form behind an opt-in until
then.

Method note, per KERNEL_RULES.md:977-984: the count falling was asserted, and
then the resource was checked anyway. The green count is what this PR wanted; the
memory curve is why "the count fell" is not by itself a merge signal.

@heydryft

Copy link
Copy Markdown
Contributor Author

Not merging yet — one blocking change needed. CI is fully green (18/18); this is not a CI objection.

🔴 This is the single largest blast radius in the current queue, and it is default-ON.

mistralrs-core/src/pipeline/normal.rs:183:

std::env::var("ARC_CANDLE_ALLOC_CACHE").is_ok_and(|v| v == "0")

That is a disable predicate. Unset ⇒ killed = false ⇒ want = seq_len == 1 ⇒ AllocCacheAction::Enable ⇒ cd.set_alloc_cache_enabled(true) on every decode step, with no flag set.

The previous condition was probe && seq_len == 1 && env::var_os(...).is_some() — three terms, all default-off. This PR drops probe and inverts the env var from opt-in to opt-out. That is two independent moves toward on-by-default in one change.

Why that matters here specifically: candle's caching allocator is exact-byte-keyed with no eviction and no capacity bound. This PR's own documentation concedes that the KV-length-tracking buffers (the causal mask) grow O(context^2) and are 'not bounded by this policy'. Turning an unbounded cache on by default for every decode is how a long-running server OOMs in production rather than in CI — and CI will never see it, because CI does not run long contexts.

Ask: invert the polarity — ARC_CANDLE_ALLOC_CACHE=1 to opt in, unset means off. Keep the seq_len == 1 narrowing. Then it lands.

Also note #182 (perf/alloc-cache-bounded, 'bound the caching allocator: +6.04 MiB/token leak -> flat') exists and is exactly the missing bound for this. These two should land together, bound first, or this one lands opt-in and #182 relaxes it later. Landing this one alone, default-ON, unbounded, is the worst of the three orderings.

Nothing else here is objectionable: no .cu touched, no format/ABI change, no dependency movement, no PagedAttention or prefix-cache contact.

heydryft and others added 18 commits August 20, 2026 18:25
…path

The V4 B=1 decode step pays 43 blocking `cuMemcpyDtoHAsync_v2` per token
(~109us each, 4.81 ms/token) inside `dsv4_kv_fp8::e4m3_codes_cpu`, which
round-trips every layer's scaled K block through the HOST because candle
has no CUDA F8E4M3 cast. Those copies are invisible to the obvious grep
(`*Synchronize*` = 0.0 calls/step) and they make CUDA graph capture
impossible: a graph cannot record a blocking D2H.

The measured trap: the existing sync-free path (`ARC_GPU_ACT_QUANT=1`,
`GpuApprox`) is SLOWER on H200 - interleaved A1 66.99 / B1 67.71 / A2
68.05 / B2 70.30 ms/token, with `kv_fp8_quant` 73.49 -> 134.18 us/call
(+83%). It swaps one blocking copy for ~19 extra elementwise launches per
layer, and this machine is bottlenecked on op count (9,131 launches/token,
median kernel 1.18us, memory controller 4% utilized), not kernel speed. So
`GpuApprox` is not the fix and is documented here as not being one.

This adds `mistralrs_quant::arc_kvquant`: one fused kernel per direction,
replacing ~11 candle ops on the quantize side and ~13 on the dequantize
side with a single launch each, and removing the D2H entirely. The byte
format is unchanged.

Bit-parity with the CPU path is the bar, and it is obtained by
construction rather than by hope:
  * the E4M3 rounding is a transcription of NVIDIA's
    `__nv_cvt_double_to_fp8(x, SATFINITE, E4M3)`, which is what the Rust
    `float8` crate ports and therefore what `F8E4M3::from_f32` computes;
  * `scale` reproduces `(amax / 448.0)?.affine(1.0, 1e-12)` exactly,
    including that candle lowers `Tensor / f64` to a MULTIPLY by the
    f32-rounded reciprocal;
  * mistralrs-quant compiles with `--use_fast_math`, so every float op is
    an explicit `__f*_rn` intrinsic (IEEE, unaffected by -prec-div/-ftz)
    and the amax reduction runs on absolute-value bit patterns as
    unsigned integers rather than through `fabsf`/`fmaxf`;
  * dequant indexes the SAME 256-entry `F8E4M3::from_bits(i).to_f32()`
    table the candle path fed to `index_select`.

D33: the kernel ships a deliberate mutant (RNE replaced by truncation,
everything else identical) reachable only from the parity test, so the
comparison is shown to fail on a wrong kernel. The test also asserts the
fused call counters moved - a parity check on this exact subsystem has
already passed vacuously by comparing an implementation to itself.

D14: the GPU parity test exits 2 (environment failure) rather than
passing when no CUDA device is present.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…z) + box paths

The emitted PTX showed nvcc 13.1 turning __fmul_rn/__fadd_rn/__fdiv_rn into
mul.rn.ftz.f32 / div.rn.ftz.f32 / add.rn.ftz.f32 under --use_fast_math, which
candle-kernels (no fast math) does not do. Replaced with inline PTX, which no
optimisation flag rewrites, so 'grep -c .ftz.f32' over the PTX is the audit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Runs on hardware with nvcc alone (no cargo build). Compares the kernel's
transcription against NVIDIA's software reference (what the Rust float8 crate
ports, hence what candle's CPU cast computes) and against the sm_90 hardware
path, over ALL 2^32 f32 bit patterns.

Measured on H200 / CUDA 13.1:
  visited               4,294,967,296 of 4,294,967,296
  mismatch vs NVIDIA sw 0
  mismatch sw vs hw     0
  negative control      123,731,850 inputs caught (2.88%)

The visited counter and the -DMUTANT=1 control exist because '0 mismatches'
from a sweep that ran zero iterations is indistinguishable from a pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The 43 blocking cuMemcpyDtoHAsync_v2 per V4 decode step are invisible to the
obvious instrument (*Synchronize* is 0.0 calls/step), so this counts the copies
themselves and pins the step count from the trace with two independent anchors
rather than assuming it.

Known-answer test on the recorded baseline trace (/root/budget-chain/nsys):
  D2H PER STEP      44.09   (BUDGET_V4_B1.md records 44)
  LAUNCHES PER STEP 9131.5  (records 9,131)
  1,792 B x 8,144 = 43.09/step; 517,120 B x 189 = 1.00/step -- the exact two
  DtoH sizes the budget names.
Instrument validated before being pointed at the new traces.

Drops the earlier measure_kv_fp8_fused.sh: it was written before the box paths
were known and its exclusivity check had a defect (it cleared the pid it was
meant to compare against). Replaced by the scripts actually run on the box.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Counts survive contention (nsys traces only this process), but V4 is ~79 GB of
the H200's 143 GB, so holding the bench lock is not enough — the previous
holder's process can still be resident. Gates on nvidia-smi free memory and
exits 2 on OOM or a missing report rather than reporting a partial trace.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…sync

Adds cuMemcpyDtoHAsync_v2 / cuLaunchKernel / cuMemAllocAsync / cuMemFreeAsync
per step, and prints ALL *Synchronize* alongside — it reads 0.00/step, which is
exactly why counting syncs on this workload finds nothing and concludes wrongly.

Known-answer re-test on the recorded baseline trace now reproduces every
headline in BUDGET_V4_B1.md:
  44.09/step cuMemcpyDtoHAsync_v2      (recorded 44)
  11436.33 + 11436.23 alloc/free/step  (recorded 11,436 each)
  2818.31/step cuMemcpyHtoDAsync_v2    (recorded 2,818)
  0.00/step ALL *Synchronize*          (recorded 0.0)
  9131.5 launches/step                 (recorded 9,131)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Measured on the box: a chain wrapping its whole pipeline in one flock held
/root/locks/bench.lock with 13-17 waiters queued while the GPU read
0 %, 0 MiB, 78.28 W. nsys report export and the per-step counting are pure CPU
on files already written, and must not hold the card.

The script now takes the lock itself, per leg, around the VRAM wait and the
traced run, and releases it the instant the bench exits — and says so in the
output so the release time is auditable. Do not wrap it in flock.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Lock covers exactly one leg (model load, which allocates, plus the timed run)
and is released before parsing. Interleaves A-B-A-B and prints the within-arm
drift beside the delta, because this box once drifted 3.3% monotonically —
more than the arm difference.

On a bench failure it dumps 'dmesg | grep -i xid' first: this box carries
~1,485 Xid GPU faults (ECC uncorrected 0, so not memory corruption), and a
process killed by one dies with no error line. Distinguishing the box's fault
from the code's is the difference between fixing a bug and chasing one that
does not exist.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…URED]

Box arc-v4-stack (H200), binary md5 e15259dc9ce935fa8782ba832cac1992 in a
PRIVATE target dir (/root/arc-wt/fp8-target, not the shared /root/arc-wt/target
that a neighbour's build overwrote tonight), 118 arc_kv_fp8_quantize symbol
hits, --features 'cuda flash-attn'. nsys, 20 s steady-state tail, step count
pinned from the trace by two independent anchors.

                          before (cpu)   after (fused)
  cuMemcpyDtoHAsync_v2      43.89/step       1.00/step
  ... of which 1,792 B      42.89/step   size ABSENT from the trace
  ... of which 517,120 B     1.00/step       1.00/step   (logits, not ours)
  kernel launches         9,081.8/step    8,404.5/step   (-677.3)
  cuLaunchKernel          7,892.4/step    7,119.0/step   (-773.5)
  cuMemAllocAsync        11,376.7/step   10,618.0/step   (-758.6)
  ALL *Synchronize*           0.00/step       0.00/step   (why grep lies here)

Per-call, which is the evidence that survives a contended box:
  before  cuMemcpyDtoHAsync_v2  48.56 us/call HOST time -> 2.131 ms/step blocking
  after   arc_kv_fp8_quantize_kernel    1.98 us/call, 42.93 calls/step
          arc_kv_fp8_dequantize_kernel  2.34 us/call, 42.93 calls/step
          fused total                             0.1853 ms/step device time

42.93 rather than 43.00 is one step straddling the window edge (99.84%); the
before arm's 1,792 B D2H shows the same 0.26% at 42.89/43. Engagement is
therefore proven per layer, not assumed.

NO end-to-end tok/s delta is reported. A naive A-B-A-B delta on this box is
biased by exactly one slot of drift, and four end-to-end numbers here turned
out to be pure environment in one night. The A/B driver is dropped rather than
shipped with a result it cannot support; the case rests on per-call cost and
launch counts, which do not depend on how long the step took.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… typos selecting GpuApprox

The name says "mode", which reads as an on/off for FP8 KV. It is not one, and
it never was: it selects WHICH ARITHMETIC produces the E4M3 code, and all three
variants quantize. There is no "off" — V4 is FP8-QAT, so the quantize/dequantize
round trip at `deepseek4.rs:1689` is the model's numerics, not an optimisation,
and it is correctly unconditional. The flag that gates FP8 KV *storage* is a
different variable, `ARC_V4_FP8_KV`.

Three defects, in descending severity:

1. A TYPO COULD SILENTLY SELECT NON-BIT-EXACT ARITHMETIC. The old table was
   `_ if ARC_GPU_ACT_QUANT is set => GpuApprox` followed by `_ => FusedDevice`.
   Both arms were reachable ONLY for an unset or MISSPELLED value. So on any box
   that still had `ARC_GPU_ACT_QUANT` exported from an earlier experiment,
   `ARC_KV_FP8_MODE=fusd` selected `GpuApprox` — the one variant this module
   documents as round-half-away-from-zero rather than round-half-to-even, and
   labels "NOT the fix and must not be shipped as one". Resolved once into a
   `OnceLock`, so it stuck for the process lifetime with no trace. An
   unparsable value now lands on the bit-exact default and SAYS SO, on both
   stderr and `tracing::error!` — never on `GpuApprox`.

2. The docs stated the opposite of the code. `PROFILING.md` annotated the
   `kv_fp8_quant` span "opt-in, ARC_V4_FP8_KV=1". That span opens at
   `deepseek4.rs:1688` and runs on every forward regardless of the flag; the
   adjacent `kv_fp8_dequant` line carried no such note, so the doc was
   internally inconsistent too. The annotation moves to `kv_cache_append`,
   which is where `ARC_V4_FP8_KV` actually acts.

3. `unwrap_or_default()` collapsed unset and empty-string, and nothing trimmed.

Renamed to `ARC_KV_FP8_IMPL`. The old spelling still works and prints a
deprecation on both channels — renaming it silently would convert every
operator's muscle memory into a fresh silent failure, which is the disease
being cured, not the cure.

The table is now a pure `parse_impl(Option<&str>, bool)` with a test that pins
every arm INCLUDING the typo-must-not-reach-GpuApprox regression. It runs on the
free CPU lane; no GPU is required to keep this honest.
…uant.cu

The ArcGate kernel-count tripwire caught a real regression introduced by
rebasing this branch onto current master, and it is worth recording how,
because the failure mode is subtle.

This branch originally carried `EXPECTED_KERNEL_COUNT 39 -> 40 — arc_kvquant.cu
was never counted`. Meanwhile master independently went 39 -> 40 for a
DIFFERENT kernel. On rebase, git compared patch texts, saw an identical
`-39 / +41`-shaped hunk already upstream, and dropped the commit as
"patch contents already upstream". The number was right; the REASON was not.
Net effect: the glob discovers 41 sources while the file still claims 40, so
`arc_kvquant.cu` would once again be the kernel that goes missing quietly —
which is the exact failure this file was created to make impossible.

Caught by `cuda_kernel_build_guard::expected_kernel_count_matches_disk` on the
free CPU lane, before any GPU time was spent. That is the tripwire earning its
keep, not a nuisance — do not "fix" a future occurrence by relaxing the guard.
`set_alloc_cache_enabled(true)` sat behind three stacked default-off gates:
`probe && seq_len == 1 && env("ARC_CANDLE_ALLOC_CACHE")`. The first conjunct
tied a general-purpose allocator to the V4 capture probe, so the only way to
recycle a decode step's frees was to also be capturing. The ~11k allocations
per token were a disabled feature, not a missing one.

Replaced with a pure policy, `alloc_cache_action(seq_len, enabled, killed)`:

* decode (`seq_len == 1`)  -> Enable
* prefill (`seq_len != 1`) -> DrainAndDisable
* already in that state    -> Leave

Prefill draining is not incidental. The cache is keyed on exact byte count
(`free: HashMap<usize, Vec<CUdeviceptr>>`, no bucketing, no smallest-fit) and
has no capacity bound and no eviction, so a prefill's large one-shot buffers
would be parked for the process lifetime under a key nothing asks for again.

`ARC_CANDLE_ALLOC_CACHE=0` is the kill switch. Any other value, and unset,
leave the policy in force, so the `=1` the ops scripts pass still means what
it always meant.

The capture probe keeps its own gate for the graph-mode positions; only the
allocator moved out from under it.

Six host-runnable tests for the policy. The allocator has no tests at all in
either repo — candle's `cuda_backend/{device,mod}.rs` carry no `#[cfg(test)]`
and every arc-side exercise is `#[cfg(feature = "cuda")]` — so the decision of
when it is on is now the part that CI can see.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpVNMpb13HkUXqSqbN1o9H
… prove it

Picks up candle b2a4dd80, which gives `AllocCache` a capacity and LRU eviction.
The allocator had neither, and arc only drains it on a prefill, so a single long
generation grew forever. Measured on an H200 over 2 600 tokens, `memory.used` at
2 Hz: **+6.04 MiB per decoded token with no plateau**, against +0.057 MiB/token
for the same run with the cache off — so the growth is the cache's and nothing
else's.

`ARC_ALLOC_CACHE_MAX_MB` sets the cap; unset leaves candle's 1 GiB default; `0`
restores the old unbounded behaviour for A/B. A typo deliberately does *not*
fall back to unbounded.

`ARC_ALLOC_CACHE_STATS=N` prints the allocator's counters every N decode steps:
allocations per step, **frees per step**, hit rate, bytes held against the cap.
Those are the numbers this has to be judged on. A green log is not evidence — an
earlier arena here reported "accounting OK" and bit-identical output over 52
steps while silently bypassing itself for every buffer under 128 bytes
(KERNEL_RULES.md:977-984). Allocations staying low *and* frees being non-zero is.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpVNMpb13HkUXqSqbN1o9H
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpVNMpb13HkUXqSqbN1o9H
Test-only change on the candle side; `git diff b2a4dd80..859c49c8` touches
nothing outside `#[cfg(test)] mod alloc_cache_tests`. The H200 numbers in the
branch description were measured at b2a4dd80 and stand.

Correction to that commit message: the allocator has **ten** tests, not eleven.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpVNMpb13HkUXqSqbN1o9H
Closes a free path this change introduced: leaving capture mode re-filed parked
buffers through the evicting put, which could hand a private-pool pointer back
to the driver at a moment arc-cuda-graph does not control. Decode is unaffected
— `set_capture_mode` is only called during a capture.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpVNMpb13HkUXqSqbN1o9H
Two sites in pipeline/normal.rs — a doc comment and a test name. No
behaviour change; the Typos job was the only red check on this PR.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`chore(deps): candle 89ab14ef` bumped all five candle crates in
`Cargo.toml` but left `Cargo.lock` pinning 88d86a2. That is the stale-lock
trap: a `--locked` CI build resolves the LOCK, so the lane would have
compiled the OLD, UNBOUNDED allocator while reporting green on a PR whose
entire subject is bounding it. `cargo metadata --locked` now succeeds; it
failed before this commit.

Regenerated with `cargo update -p candle-{core,nn,flash-attn,flash-attn-v3,metal-kernels}`.
The only non-candle churn is a `windows-core` 0.61.2/0.62.2 dedup that fell
out of the re-resolution; nothing arc builds on Linux or macOS reads it.
@heydryft
heydryft force-pushed the perf/alloc-cache-reachable-on-decode branch from 7d8760f to d71fec0 Compare August 20, 2026 17:48
@heydryft
heydryft changed the base branch from master to arckv-fp8-fused August 20, 2026 17:48
@heydryft heydryft changed the title perf(arckv): make candle's caching allocator reachable and on for decode perf(arckv): make candle's caching allocator reachable, on for decode, and BOUNDED Aug 20, 2026
@heydryft

Copy link
Copy Markdown
Contributor Author

Rebased onto current master, re-based on #175, and #182 folded in so this ships reachable and bounded as one change rather than a fast path plus a follow-up nobody merges.

The polarity was already right (ARC_CANDLE_ALLOC_CACHE=0 is a kill switch). The defect was that the cache was unbounded with no eviction: +6.04 MiB/token, forever. It is now bounded, for 0.15 ms/token.

Stale-lock defect fixed. #182 bumped all five candle crates in Cargo.toml to 89ab14ef but left Cargo.lock resolving 88d86a2; cargo metadata --locked failed on that branch. Since a --locked build resolves the LOCK, the lane would have compiled the OLD, UNBOUNDED allocator while reporting green on the PR whose entire subject is bounding it. Regenerated with cargo update -p candle-{core,nn,flash-attn,flash-attn-v3,metal-kernels}; the only non-candle churn is a windows-core 0.61.2/0.62.2 dedup that fell out of re-resolution and that nothing arc builds on Linux or macOS reads.

Candle ancestry verified before bumping: 88d86a2 -> b2a4dd80 -> 859c49c8 -> 89ab14ef is a clean linear chain. 89ab14ef is also the rev that stops eviction inside the capture window — freeing a private-pool pointer after cuMemPoolDestroy corrupts the glibc arena, which is the corrupted size vs. prev_size blocker on #181 — so it is the safer pin, not merely a newer one.

Reconciliation warning for whoever touches candle next. arcgraph/prewarm-profile (candle 1fa534d) is a SIBLING off 88d86a2, not an ancestor. It conflicts with 89ab14ef in four hunks of candle-core/src/cuda_backend/device.rs: the AllocCache struct body, set_alloc_cache_enabled, drain_alloc_cache_and_free, and cache_take. The substantive half is that prewarm assumes free's value type is a Vec while bounded changes it to a SizeClass, so three of its sites will not compile even after the markers are resolved, and alloc_cache_free_counts / prewarm_alloc_cache need the same adaptation. Reconcile — do not clobber.

New env vars, all fail-safe (a typo degrades to bounded, never to unbounded): ARC_ALLOC_CACHE_MAX_MB (unset -> candle's 1 GiB default; 0 -> explicitly unbounded), ARC_ALLOC_CACHE_STATS (off unless set; logs every N decode steps). The counter that proves the bound is real is free/step — the unbounded cache's was exactly zero, forever.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant