perf(ArcMoE): fuse the V4 router region — 18+9 launches become 2, and fix the sinkhorn local-memory spill - #185
Conversation
Code Metrics Report━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Language Files Lines Code Comments Blanks ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ C Header 5 305 210 52 43 CSS 2 1181 1036 34 111 CUDA 74 25734 18232 4703 2799 Dockerfile 1 39 22 8 9 JavaScript 16 3546 2676 482 388 Jinja2 7 694 656 5 33 JSON 74 4600 4597 0 3 Makefile 1 6 5 0 1 Metal Shading Lan| 33 12224 9431 1142 1651 PowerShell 1 300 227 30 43 Python 145 15139 12482 811 1846 Shell 40 9948 6637 2655 656 Plain Text 4 3801 0 2479 1322 TOML 33 1498 1294 54 150 YAML 3 25 23 2 0 ───────────────────────────────────────────────────────────────────────────────── HTML 4 2687 2604 43 40 |- CSS 2 543 479 37 27 |- JavaScript 1 1233 1215 12 6 (Total) 4463 4298 92 73 ───────────────────────────────────────────────────────────────────────────────── Jupyter Notebooks 4 122 83 23 16 |- Markdown 1 60 30 22 8 |- Python 1 122 113 1 8 (Total) 304 226 46 32 ───────────────────────────────────────────────────────────────────────────────── Markdown 204 45330 0 35183 10147 |- BASH 72 1655 1203 331 121 |- C 3 17 17 0 0 |- CUDA 2 84 56 16 12 |- JSON 18 708 708 0 0 |- PowerShell 1 1 1 0 0 |- Python 23 1008 787 113 108 |- Rust 66 2051 1716 77 258 |- TOML 6 207 164 0 43 |- YAML 5 41 36 5 0 (Total) 51102 4688 35725 10689 ───────────────────────────────────────────────────────────────────────────────── Rust 673 331228 285444 16819 28965 |- Markdown 491 29538 471 25464 3603 (Total) 360766 285915 42283 32568 ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Total 1324 495675 352655 90603 52417 ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ |
|
Housekeeping: I cancelled the queued CI runs on this PR and on #186-#190 immediately after opening them. Not a signal about the code — the repo's Actions queue was 27 runs deep and starving master's own post-merge confirmation, and opening six PRs at once is what pushed it there. That is my load to clear. These PRs all need the default-ON polarity change discussed above before they are mergeable anyway, so a CI result right now would not be actionable. Re-run with |
|
🚫 UNVERIFIED — NOT MERGEABLE. Do not read the check list on this PR as a pass. Every CI run on this PR is in The standing rule is that This PR has never had a real CI verdict. Treat it as red until it does. I am re-triggering CI staggered (not all six at once — simultaneous dispatch is what exhausted the runners). Once this PR shows a genuine To re-trigger by hand: |
|
Cause identified, and CI re-triggered — this one CAN reach a genuine verdict. The single That is This PR targets The UNVERIFIED note stands until this shows a genuine |
|
✅ UNVERIFIED note DISCHARGED — this PR now has a genuine CI verdict: 18/18 pass, including For the record, the full chain that got here: the original runs were cancelled by me for capacity triage, 🔴 The blast-radius objection stands and is untouched by this. CI cannot see a default flip. *ENABLED.get_or_init(|| !matches!(std::env::var("ARC_HC_FUSED").as_deref(), Ok("0")))— fused path ON when the variable is unset. This does not merge until that is inverted to opt-in ( Two separate gates, and only one is now satisfied: CI green ✅ · polarity ❌. |
…path
The V4 B=1 decode step pays 43 blocking `cuMemcpyDtoHAsync_v2` per token
(~109us each, 4.81 ms/token) inside `dsv4_kv_fp8::e4m3_codes_cpu`, which
round-trips every layer's scaled K block through the HOST because candle
has no CUDA F8E4M3 cast. Those copies are invisible to the obvious grep
(`*Synchronize*` = 0.0 calls/step) and they make CUDA graph capture
impossible: a graph cannot record a blocking D2H.
The measured trap: the existing sync-free path (`ARC_GPU_ACT_QUANT=1`,
`GpuApprox`) is SLOWER on H200 - interleaved A1 66.99 / B1 67.71 / A2
68.05 / B2 70.30 ms/token, with `kv_fp8_quant` 73.49 -> 134.18 us/call
(+83%). It swaps one blocking copy for ~19 extra elementwise launches per
layer, and this machine is bottlenecked on op count (9,131 launches/token,
median kernel 1.18us, memory controller 4% utilized), not kernel speed. So
`GpuApprox` is not the fix and is documented here as not being one.
This adds `mistralrs_quant::arc_kvquant`: one fused kernel per direction,
replacing ~11 candle ops on the quantize side and ~13 on the dequantize
side with a single launch each, and removing the D2H entirely. The byte
format is unchanged.
Bit-parity with the CPU path is the bar, and it is obtained by
construction rather than by hope:
* the E4M3 rounding is a transcription of NVIDIA's
`__nv_cvt_double_to_fp8(x, SATFINITE, E4M3)`, which is what the Rust
`float8` crate ports and therefore what `F8E4M3::from_f32` computes;
* `scale` reproduces `(amax / 448.0)?.affine(1.0, 1e-12)` exactly,
including that candle lowers `Tensor / f64` to a MULTIPLY by the
f32-rounded reciprocal;
* mistralrs-quant compiles with `--use_fast_math`, so every float op is
an explicit `__f*_rn` intrinsic (IEEE, unaffected by -prec-div/-ftz)
and the amax reduction runs on absolute-value bit patterns as
unsigned integers rather than through `fabsf`/`fmaxf`;
* dequant indexes the SAME 256-entry `F8E4M3::from_bits(i).to_f32()`
table the candle path fed to `index_select`.
D33: the kernel ships a deliberate mutant (RNE replaced by truncation,
everything else identical) reachable only from the parity test, so the
comparison is shown to fail on a wrong kernel. The test also asserts the
fused call counters moved - a parity check on this exact subsystem has
already passed vacuously by comparing an implementation to itself.
D14: the GPU parity test exits 2 (environment failure) rather than
passing when no CUDA device is present.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…z) + box paths The emitted PTX showed nvcc 13.1 turning __fmul_rn/__fadd_rn/__fdiv_rn into mul.rn.ftz.f32 / div.rn.ftz.f32 / add.rn.ftz.f32 under --use_fast_math, which candle-kernels (no fast math) does not do. Replaced with inline PTX, which no optimisation flag rewrites, so 'grep -c .ftz.f32' over the PTX is the audit. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Runs on hardware with nvcc alone (no cargo build). Compares the kernel's transcription against NVIDIA's software reference (what the Rust float8 crate ports, hence what candle's CPU cast computes) and against the sm_90 hardware path, over ALL 2^32 f32 bit patterns. Measured on H200 / CUDA 13.1: visited 4,294,967,296 of 4,294,967,296 mismatch vs NVIDIA sw 0 mismatch sw vs hw 0 negative control 123,731,850 inputs caught (2.88%) The visited counter and the -DMUTANT=1 control exist because '0 mismatches' from a sweep that ran zero iterations is indistinguishable from a pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The 43 blocking cuMemcpyDtoHAsync_v2 per V4 decode step are invisible to the obvious instrument (*Synchronize* is 0.0 calls/step), so this counts the copies themselves and pins the step count from the trace with two independent anchors rather than assuming it. Known-answer test on the recorded baseline trace (/root/budget-chain/nsys): D2H PER STEP 44.09 (BUDGET_V4_B1.md records 44) LAUNCHES PER STEP 9131.5 (records 9,131) 1,792 B x 8,144 = 43.09/step; 517,120 B x 189 = 1.00/step -- the exact two DtoH sizes the budget names. Instrument validated before being pointed at the new traces. Drops the earlier measure_kv_fp8_fused.sh: it was written before the box paths were known and its exclusivity check had a defect (it cleared the pid it was meant to compare against). Replaced by the scripts actually run on the box. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Counts survive contention (nsys traces only this process), but V4 is ~79 GB of the H200's 143 GB, so holding the bench lock is not enough — the previous holder's process can still be resident. Gates on nvidia-smi free memory and exits 2 on OOM or a missing report rather than reporting a partial trace. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…sync Adds cuMemcpyDtoHAsync_v2 / cuLaunchKernel / cuMemAllocAsync / cuMemFreeAsync per step, and prints ALL *Synchronize* alongside — it reads 0.00/step, which is exactly why counting syncs on this workload finds nothing and concludes wrongly. Known-answer re-test on the recorded baseline trace now reproduces every headline in BUDGET_V4_B1.md: 44.09/step cuMemcpyDtoHAsync_v2 (recorded 44) 11436.33 + 11436.23 alloc/free/step (recorded 11,436 each) 2818.31/step cuMemcpyHtoDAsync_v2 (recorded 2,818) 0.00/step ALL *Synchronize* (recorded 0.0) 9131.5 launches/step (recorded 9,131) Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Measured on the box: a chain wrapping its whole pipeline in one flock held /root/locks/bench.lock with 13-17 waiters queued while the GPU read 0 %, 0 MiB, 78.28 W. nsys report export and the per-step counting are pure CPU on files already written, and must not hold the card. The script now takes the lock itself, per leg, around the VRAM wait and the traced run, and releases it the instant the bench exits — and says so in the output so the release time is auditable. Do not wrap it in flock. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Lock covers exactly one leg (model load, which allocates, plus the timed run) and is released before parsing. Interleaves A-B-A-B and prints the within-arm drift beside the delta, because this box once drifted 3.3% monotonically — more than the arm difference. On a bench failure it dumps 'dmesg | grep -i xid' first: this box carries ~1,485 Xid GPU faults (ECC uncorrected 0, so not memory corruption), and a process killed by one dies with no error line. Distinguishing the box's fault from the code's is the difference between fixing a bug and chasing one that does not exist. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…URED]
Box arc-v4-stack (H200), binary md5 e15259dc9ce935fa8782ba832cac1992 in a
PRIVATE target dir (/root/arc-wt/fp8-target, not the shared /root/arc-wt/target
that a neighbour's build overwrote tonight), 118 arc_kv_fp8_quantize symbol
hits, --features 'cuda flash-attn'. nsys, 20 s steady-state tail, step count
pinned from the trace by two independent anchors.
before (cpu) after (fused)
cuMemcpyDtoHAsync_v2 43.89/step 1.00/step
... of which 1,792 B 42.89/step size ABSENT from the trace
... of which 517,120 B 1.00/step 1.00/step (logits, not ours)
kernel launches 9,081.8/step 8,404.5/step (-677.3)
cuLaunchKernel 7,892.4/step 7,119.0/step (-773.5)
cuMemAllocAsync 11,376.7/step 10,618.0/step (-758.6)
ALL *Synchronize* 0.00/step 0.00/step (why grep lies here)
Per-call, which is the evidence that survives a contended box:
before cuMemcpyDtoHAsync_v2 48.56 us/call HOST time -> 2.131 ms/step blocking
after arc_kv_fp8_quantize_kernel 1.98 us/call, 42.93 calls/step
arc_kv_fp8_dequantize_kernel 2.34 us/call, 42.93 calls/step
fused total 0.1853 ms/step device time
42.93 rather than 43.00 is one step straddling the window edge (99.84%); the
before arm's 1,792 B D2H shows the same 0.26% at 42.89/43. Engagement is
therefore proven per layer, not assumed.
NO end-to-end tok/s delta is reported. A naive A-B-A-B delta on this box is
biased by exactly one slot of drift, and four end-to-end numbers here turned
out to be pure environment in one night. The A/B driver is dropped rather than
shipped with a result it cannot support; the case rests on per-call cost and
launch counts, which do not depend on how long the step took.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… typos selecting GpuApprox The name says "mode", which reads as an on/off for FP8 KV. It is not one, and it never was: it selects WHICH ARITHMETIC produces the E4M3 code, and all three variants quantize. There is no "off" — V4 is FP8-QAT, so the quantize/dequantize round trip at `deepseek4.rs:1689` is the model's numerics, not an optimisation, and it is correctly unconditional. The flag that gates FP8 KV *storage* is a different variable, `ARC_V4_FP8_KV`. Three defects, in descending severity: 1. A TYPO COULD SILENTLY SELECT NON-BIT-EXACT ARITHMETIC. The old table was `_ if ARC_GPU_ACT_QUANT is set => GpuApprox` followed by `_ => FusedDevice`. Both arms were reachable ONLY for an unset or MISSPELLED value. So on any box that still had `ARC_GPU_ACT_QUANT` exported from an earlier experiment, `ARC_KV_FP8_MODE=fusd` selected `GpuApprox` — the one variant this module documents as round-half-away-from-zero rather than round-half-to-even, and labels "NOT the fix and must not be shipped as one". Resolved once into a `OnceLock`, so it stuck for the process lifetime with no trace. An unparsable value now lands on the bit-exact default and SAYS SO, on both stderr and `tracing::error!` — never on `GpuApprox`. 2. The docs stated the opposite of the code. `PROFILING.md` annotated the `kv_fp8_quant` span "opt-in, ARC_V4_FP8_KV=1". That span opens at `deepseek4.rs:1688` and runs on every forward regardless of the flag; the adjacent `kv_fp8_dequant` line carried no such note, so the doc was internally inconsistent too. The annotation moves to `kv_cache_append`, which is where `ARC_V4_FP8_KV` actually acts. 3. `unwrap_or_default()` collapsed unset and empty-string, and nothing trimmed. Renamed to `ARC_KV_FP8_IMPL`. The old spelling still works and prints a deprecation on both channels — renaming it silently would convert every operator's muscle memory into a fresh silent failure, which is the disease being cured, not the cure. The table is now a pure `parse_impl(Option<&str>, bool)` with a test that pins every arm INCLUDING the typo-must-not-reach-GpuApprox regression. It runs on the free CPU lane; no GPU is required to keep this honest.
…uant.cu The ArcGate kernel-count tripwire caught a real regression introduced by rebasing this branch onto current master, and it is worth recording how, because the failure mode is subtle. This branch originally carried `EXPECTED_KERNEL_COUNT 39 -> 40 — arc_kvquant.cu was never counted`. Meanwhile master independently went 39 -> 40 for a DIFFERENT kernel. On rebase, git compared patch texts, saw an identical `-39 / +41`-shaped hunk already upstream, and dropped the commit as "patch contents already upstream". The number was right; the REASON was not. Net effect: the glob discovers 41 sources while the file still claims 40, so `arc_kvquant.cu` would once again be the kernel that goes missing quietly — which is the exact failure this file was created to make impossible. Caught by `cuda_kernel_build_guard::expected_kernel_count_matches_disk` on the free CPU lane, before any GPU time was spent. That is the tripwire earning its keep, not a nuisance — do not "fix" a future occurrence by relaxing the guard.
`set_alloc_cache_enabled(true)` sat behind three stacked default-off gates:
`probe && seq_len == 1 && env("ARC_CANDLE_ALLOC_CACHE")`. The first conjunct
tied a general-purpose allocator to the V4 capture probe, so the only way to
recycle a decode step's frees was to also be capturing. The ~11k allocations
per token were a disabled feature, not a missing one.
Replaced with a pure policy, `alloc_cache_action(seq_len, enabled, killed)`:
* decode (`seq_len == 1`) -> Enable
* prefill (`seq_len != 1`) -> DrainAndDisable
* already in that state -> Leave
Prefill draining is not incidental. The cache is keyed on exact byte count
(`free: HashMap<usize, Vec<CUdeviceptr>>`, no bucketing, no smallest-fit) and
has no capacity bound and no eviction, so a prefill's large one-shot buffers
would be parked for the process lifetime under a key nothing asks for again.
`ARC_CANDLE_ALLOC_CACHE=0` is the kill switch. Any other value, and unset,
leave the policy in force, so the `=1` the ops scripts pass still means what
it always meant.
The capture probe keeps its own gate for the graph-mode positions; only the
allocator moved out from under it.
Six host-runnable tests for the policy. The allocator has no tests at all in
either repo — candle's `cuda_backend/{device,mod}.rs` carry no `#[cfg(test)]`
and every arc-side exercise is `#[cfg(feature = "cuda")]` — so the decision of
when it is on is now the part that CI can see.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpVNMpb13HkUXqSqbN1o9H
… prove it Picks up candle b2a4dd80, which gives `AllocCache` a capacity and LRU eviction. The allocator had neither, and arc only drains it on a prefill, so a single long generation grew forever. Measured on an H200 over 2 600 tokens, `memory.used` at 2 Hz: **+6.04 MiB per decoded token with no plateau**, against +0.057 MiB/token for the same run with the cache off — so the growth is the cache's and nothing else's. `ARC_ALLOC_CACHE_MAX_MB` sets the cap; unset leaves candle's 1 GiB default; `0` restores the old unbounded behaviour for A/B. A typo deliberately does *not* fall back to unbounded. `ARC_ALLOC_CACHE_STATS=N` prints the allocator's counters every N decode steps: allocations per step, **frees per step**, hit rate, bytes held against the cap. Those are the numbers this has to be judged on. A green log is not evidence — an earlier arena here reported "accounting OK" and bit-identical output over 52 steps while silently bypassing itself for every buffer under 128 bytes (KERNEL_RULES.md:977-984). Allocations staying low *and* frees being non-zero is. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SpVNMpb13HkUXqSqbN1o9H
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SpVNMpb13HkUXqSqbN1o9H
Test-only change on the candle side; `git diff b2a4dd80..859c49c8` touches nothing outside `#[cfg(test)] mod alloc_cache_tests`. The H200 numbers in the branch description were measured at b2a4dd80 and stand. Correction to that commit message: the allocator has **ten** tests, not eleven. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SpVNMpb13HkUXqSqbN1o9H
Closes a free path this change introduced: leaving capture mode re-filed parked buffers through the evicting put, which could hand a private-pool pointer back to the driver at a moment arc-cuda-graph does not control. Decode is unaffected — `set_capture_mode` is only called during a capture. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SpVNMpb13HkUXqSqbN1o9H
Two sites in pipeline/normal.rs — a doc comment and a test name. No behaviour change; the Typos job was the only red check on this PR. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`chore(deps): candle 89ab14ef` bumped all five candle crates in
`Cargo.toml` but left `Cargo.lock` pinning 88d86a2. That is the stale-lock
trap: a `--locked` CI build resolves the LOCK, so the lane would have
compiled the OLD, UNBOUNDED allocator while reporting green on a PR whose
entire subject is bounding it. `cargo metadata --locked` now succeeds; it
failed before this commit.
Regenerated with `cargo update -p candle-{core,nn,flash-attn,flash-attn-v3,metal-kernels}`.
The only non-candle churn is a `windows-core` 0.61.2/0.62.2 dedup that fell
out of the re-resolution; nothing arc builds on Linux or macOS reads it.
… fix the sinkhorn local-memory spill At b=1 the V4 decode step is launch-bound: ~7,924 cuLaunchKernel calls, of which the MoE router region is 3,259 (41.8%) across 86 contiguous spans — exactly two per layer (mhc_attn_pre, mhc_ffn_pre; 43 layers). Mean kernel in that region is 2.79 us and 89% are under 5 us, on [1,24] / [1,4,4] / [1,256] tensors. That is launch overhead wearing a kernel costume. Three changes, all bit-identical to what they replace: 1. cuda/hc_fused.cu `hc_pre_fused_f32` collapses hc_pre's 18-launch middle — a hand-decomposed RMS statistic in SEVEN launches (sqr, fast_sum, affine, affine, recip, sqrt, bmul) plus ELEVEN for the pre/post/comb scoring — into one kernel. `mixes` is never materialised. 2. cuda/hc_fused.cu `sqrt_softplus_f32` collapses MoeGate's sqrt(softplus(x)) from nine launches (zeros_like, bmaximum, uabs, uneg, uexp, affine, ulog, badd, usqrt) into one. 3. cuda/sinkhorn.cu is templated on `hc`. It took `hc` at runtime, so nvcc could not keep r[16]/buf[16]/col[16] in registers and demoted all three to local memory — `ptxas -v` reported 192 bytes stack frame. With hc=4 the kernel runs ONE block of FOUR threads, so nothing hides that latency: measured 30.0 us/call on a [1,4,4] tensor, 2.549 ms/step over 86 calls, 28% of the router's GPU time. Templating makes every index compile-time; ptxas now reports 0 bytes stack frame, 29 registers. Arithmetic unchanged. Bit-identity is by construction, not by tolerance: this region decides WHICH EXPERTS RUN, so a reassociated sum or a contracted FMA can change the emitted token. hc_fused.cu joins sinkhorn.cu in build.rs's dedicated no-fast-math, --fmad=false builder and carries the same #error guard, and it transcribes candle's ops exactly — including that `recipg` is `1.0 / a` with a *double* literal, that `mean_keepdim` is fast_sum followed by a separate affine, and that candle's FastReduce block_dim is min(1024, len).next_power_of_two(), which fixes the reduction tree's shape and therefore the f32 rounding. cuda/hc_fused.rs carries scalar replicas of both sides asserted bit-identical without a GPU, plus two anti-vacuity guards. The softplus one earned its place: its first version sampled [-30, -12.5, -1, 0.5, 3.25, 17] and passed against the naive log(1+exp(x)) form, because the two agree bit for bit until exp overflows near 88.7. It is now a three-way mutation probe. ARC_HC_FUSED=0 restores the eager chains so both can be A/B'd from one binary. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SpVNMpb13HkUXqSqbN1o9H
…phic wrap closure Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SpVNMpb13HkUXqSqbN1o9H
The sinkhorn guard asserted build.rs contained the literal `.exclude(&["sinkhorn.cu"])`, so adding a second bit-identity-critical kernel to that list failed it for the one reason that is not a regression — the kind of failure that gets 'fixed' by deleting the assertion. It now matches on the exclude list's CONTENTS, and hc_fused.rs carries the mirror guard: fast-math #error present, IEEE intrinsics present, __expf/__logf/ __fdividef/rsqrtf absent, and the file wired into the --fmad=false builder. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SpVNMpb13HkUXqSqbN1o9H
Measured on H200, git 6f4cd0d, bin f6ec4e85, toggled with ARC_HC_FUSED on one binary so the comparison has no second variable: ARC_HC_FUSED=0 7,494.3 launches/step 10,892.4 allocs/step 17.3 tok/s 57.70 ms/T ARC_HC_FUSED=1 5,691.0 launches/step 8,918.5 allocs/step 20.1 tok/s 49.66 ms/T The off leg reproduces origin/master (7,477.9 launches, 17.0 tok/s) to within 0.2%, so the switch is the only thing that moved and nothing else regressed. Router region 76 -> 38 kernels/layer, 0.207 -> 0.100 ms GPU. Sinkhorn 30.03 -> 9.12 us/call (2.583 -> 0.785 ms/step). Output bit-identical across the toggle. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SpVNMpb13HkUXqSqbN1o9H
1e81f56 to
ab5fbd5
Compare
Bottom rung of the session-8 decode ladder (15.0 -> 32.6 tok/s at b=1). Opening it because the work was never proposed anywhere and losing it would be the worst outcome of the session.
Fuses the V4 router region: 18+9 launches become 2, and fixes the sinkhorn local-memory spill. Adds
mistralrs-core/src/cuda/hc_fused.cuplus a build-wiring tripwire over it.Blast radius — read before merging
mistralrs-core/src/cuda/hc_fused.rs:31reads*ENABLED.get_or_init(|| !matches!(std::env::var("ARC_HC_FUSED").as_deref(), Ok("0")))— the fused path is ON when the variable is unset. House rule is that new behaviour lands default-OFF with the toggle named. Either flip the polarity to opt-in before merging, or merge deliberately and say so. I am not making that call unilaterally.
hc_fused.cu, a new kernel on the default V4 decode path.88d86a2— identical to master. Candle-neutral,Cargo.lockagrees withCargo.toml.--use_fast_mathon the new kernel.On the numbers
The 15.0 -> 32.6 ladder was measured on this branch's base, not on today's master. Treat the number as needing re-measurement, not as a merged fact. The mechanism (launch-count reduction) is auditable from the diff; the tok/s is not, until a box exists.
Ordering
Bottom of a five-branch stack: this ->
agent/mhc-fuse-post->agent/datamove-fuse->agent/f32-cast-fuse->agent/bookkeep-collapse. Merge base-first (D20). Note that the next rung already merges in #175 and #177, so those two land before it.🤖 Generated with Claude Code