Skip to content

perf(ArcAttention): port vLLM rms_norm_kernel — fuse the V4 per-head Q RMSNorm (7 launches -> 1) - #184

Open
heydryft wants to merge 2 commits into
release/openrouter-readyfrom
perf/port-vllm-qnorm
Open

heydryft wants to merge 2 commits into
release/openrouter-readyfrom
perf/port-vllm-qnorm

Conversation

@heydryft

Copy link
Copy Markdown
Contributor

Stacked on #183 — merge that first (D20).

Structural port of vLLM's csrc/layernorm_kernels.cu:108-181 rms_norm_kernel and SGLang's fused_add_rmsnorm.cuh:57: one block per row, the sum of squares reduced inside the kernel, the normalised row written back in the same launch, no intermediate tensors. Both Apache-2.0; notice in the header.

V4's per-head Q RMSNorm ran as seven candle launches per attention layer, 43x per decode token:

sqr | fast_sum | affine(1/n, 0) | affine(1, eps) | recip | sqrt | broadcast_mul

Measured on H200, DeepSeek-V4-Flash qtip2 UQFF, b=1 decode, 200 tokens

Same binary; ARC_NO_FUSED_QNORM=1 selects the old chain (SwiGLU fusion from #183 on in both legs, so this isolates the qnorm).

leg ms/token (3 runs) median
old chain 54.653 / 54.565 / 55.970 54.653
fused kernel 54.146 / 54.490 / 53.972 54.146

-0.5 ms/token, and the fused leg is faster in all three pairs. nsys, per decode token: -258.0 kernel launches (exactly 6 x 43 layers, i.e. 7 candle kernels replaced by 1), -86.0 H2D copies, -344.0 device allocations.

Cumulative with #183 vs. the unfused baseline: 58.146 -> 54.203 ms/token, -744.7 launches / -572.0 H2D / -1313.3 allocations per token.

Bit-identity

The arithmetic is candle's, not upstream's. Three things a naive rewrite gets wrong, all pinned here:

  1. Reduction order. candle's sum_keepdim runs fast_sum with block_dim = min(1024, n).next_power_of_two(), one element per thread, then a pairwise tree — not a sequential accumulation. This is the same trap that made the first sinkhorn kernel fail its H200 A/B.
  2. BF16 accumulator. candle sums the 512 squares in bf16; the kernel does too, deliberately, so the A/B is clean.
  3. bf16 constants. candle's affine down-converts 1/n and eps to bf16 on the host, so they are passed in as raw u16 bit patterns computed with the same from_f64.

On-GPU A/B vs the exact candle chain: bit-identical over 33,440 elements across the real [64, 512] shape and a non-power-of-two [7, 96] shape that exercises the identity padding, with a negative control (1 bf16 ULP -> reported difference). End-to-end, every model leg produced the same SHA-256 over 200 generated tokens as the baseline.

Surfaced, not shipped

Accumulating 512 squares with an 8-bit mantissa carries roughly sqrt(512) * 2^-8 ~ 9% relative error in the norm. That is a pre-existing quality defect in the code being replaced, not one this port introduces — and it is preserved here on purpose so this PR is a pure perf change. ARC_QNORM_F32_ACC=1 exposes a float accumulator so it can be measured in a perplexity A/B; it is off by default and is not bit-identical. Worth a separate change.

Nirupam Bhowmick and others added 2 commits August 20, 2026 01:05
…o one launch

Ports vLLM's csrc/activation_kernels.cu `silu_and_mul_clamp` (and SGLang's
deepseek_v4/silu_and_mul_masked_post_quant.cuh `silu_and_mul<kApplySwigluLimit>`),
both Apache-2.0, attributed in the file header.

Replaces the 8-launch candle chain (2 casts + 3 clamp binaries + silu + mul +
cast, plus five 1-element H2D copies for the scalar clamp operands) with a
single kernel, on both the routed-expert and shared-expert paths.

Compiled by the dedicated IEEE (no fast-math) builder so it stays bit-identical
to candle-kernels. Also replaces sinkhorn.cu's vacuous `#if
defined(__USE_FAST_MATH__)` #error — nvcc 12.4 defines no such macro in either
pass — with assert_ieee_kernel_flags() in build.rs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…SNorm

Structural port of vLLM csrc/layernorm_kernels.cu:108-181 rms_norm_kernel and
SGLang fused_add_rmsnorm.cuh:57 (both Apache-2.0, attributed in the header):
one launch, in-kernel reduction, no intermediate tensors.

Replaces seven candle launches per attention layer (43x per decode token):
sqr -> fast_sum -> affine(1/n,0) -> affine(1,eps) -> recip -> sqrt ->
broadcast_mul.

Arithmetic is candle's, not upstream's: the chain accumulates the 512-term sum
of squares in BF16 and the kernel reproduces that exactly, including
fast_sum's pairwise reduction order and the host-side bf16 down-conversion of
the 1/n and eps constants. ARC_QNORM_F32_ACC=1 exposes the float accumulator
for a future quality A/B but is off by default and not bit-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown
Code Metrics Report
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
 Language              Files        Lines         Code     Comments       Blanks
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
 C Header                  5          305          210           52           43
 CSS                       2         1181         1036           34          111
 CUDA                     73        25015        17952         4311         2752
 Dockerfile                1           39           22            8            9
 JavaScript               16         3546         2676          482          388
 Jinja2                    7          694          656            5           33
 JSON                     74         4600         4597            0            3
 Makefile                  1            6            5            0            1
 Metal Shading Lan|       33        12224         9431         1142         1651
 PowerShell                1          300          227           30           43
 Python                  145        15139        12482          811         1846
 Shell                    39         9777         6538         2596          643
 Plain Text                4         3801            0         2479         1322
 TOML                     33         1498         1294           54          150
 YAML                      3           25           23            2            0
─────────────────────────────────────────────────────────────────────────────────
 HTML                      4         2687         2604           43           40
 |- CSS                    2          543          479           37           27
 |- JavaScript             1         1233         1215           12            6
 (Total)                             4463         4298           92           73
─────────────────────────────────────────────────────────────────────────────────
 Jupyter Notebooks         4          122           83           23           16
 |- Markdown               1           60           30           22            8
 |- Python                 1          122          113            1            8
 (Total)                              304          226           46           32
─────────────────────────────────────────────────────────────────────────────────
 Markdown                203        44777            0        34735        10042
 |- BASH                  72         1654         1202          331          121
 |- C                      3           17           17            0            0
 |- CUDA                   2           84           56           16           12
 |- JSON                  18          708          708            0            0
 |- PowerShell             1            1            1            0            0
 |- Python                23         1008          787          113          108
 |- Rust                  66         2051         1716           77          258
 |- TOML                   6          207          164            0           43
 |- YAML                   5           41           36            5            0
 (Total)                            50548         4687        35277        10584
─────────────────────────────────────────────────────────────────────────────────
 Rust                    672       328294       283104        16426        28764
 |- Markdown             490        28532          471        24619         3442
 (Total)                           356826       283575        41045        32206
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
 Total                  1320       490291       349935        88466        51890
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

@heydryft

Copy link
Copy Markdown
Contributor Author

Retargeted at the integration branch

Base changed: master → release/openrouter-ready.

The queue is being restructured to the shape the owner asked for: one PR open against master (#194), with everything else merging into a single integration branch. Agents branch off release/openrouter-ready and PR back into it; #194 is the single gate from there to master.

Two things had to land on master first for this to be workable, and both have:

  1. The base-branch CI lane used to hard-fail any PR not targeting master (ci(ArcGate): let the base-branch lane permit the single integration branch — 7 PRs are red for addressing, not quality #216). Seven PRs were red for their addressing, not their quality. The lane now accepts master or release/openrouter-ready — an exact-match allowlist, nothing else — and its intent is intact: release/openrouter-ready reaches master only through Arc → OpenRouter-ready: the single integration PR (everything else merges into this branch) #194, whose own base is master, so nothing enters master without a full run whose base is master.
  2. release/openrouter-ready was re-cut from current master. It had diverged badly — merging it as-is would have reverted TCFRAG (perf(ArcQuant/ArcKernels): TCFRAG-2B — put the qtip2b trellis GEMV on the tensor cores #203) and re-broken the flag-polarity work (fix(ArcGate): read boolean ARC_* flags by value, so =0 means off #212). It is now master plus only the work it uniquely carried.

This PR was not closed and is not considered stale. An audit of the queue found the overwhelming majority of it to be real work that was never merged, not noise.

What you need to do: rebase onto release/openrouter-ready and resolve conflicts against it rather than against master. CI lanes will now actually run and report on this PR instead of failing at the base check.


📚 Stack order — this is the TOP

This PR used to be based on its parent PR's branch, which is exactly the pattern the base-branch CI lane was built to refuse: a base that can be rebased or abandoned underneath it, so a green here described a tree that might never exist.

Both are now retargeted at release/openrouter-ready and are siblings rather than a stack. That makes the merge order load-bearing and no longer enforced by the base ref:

Merge the bottom PR FIRST, then this one. Until the bottom lands, this PR's diff against the integration branch will also contain the bottom's commits.


🔴 needs-work — this one is red for a REAL reason

Retargeting fixes the addressing problem, not this PR's own problem. Labelled needs-work so it is not mistaken for "just needs a rebase" when the queue is triaged.

This PR has genuine test failures. The base change will make the lanes run against the integration branch, but they are expected to stay red until the failures are actually fixed. Do not merge it on the strength of a green base-branch check alone.

@heydryft heydryft added the needs-work Red for a real defect, not for addressing — needs code changes, not just a rebase label Aug 21, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

needs-work Red for a real defect, not for addressing — needs code changes, not just a rebase

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant