Conversation
Code Metrics Report━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Language Files Lines Code Comments Blanks ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ C Header 5 305 210 52 43 CSS 2 1181 1036 34 111 CUDA 74 25734 18232 4703 2799 Dockerfile 1 39 22 8 9 JavaScript 16 3546 2676 482 388 Jinja2 7 694 656 5 33 JSON 74 4600 4597 0 3 Makefile 1 6 5 0 1 Metal Shading Lan| 33 12224 9431 1142 1651 PowerShell 1 300 227 30 43 Python 145 15139 12482 811 1846 Shell 40 9948 6637 2655 656 Plain Text 4 3801 0 2479 1322 TOML 33 1498 1294 54 150 YAML 3 25 23 2 0 ───────────────────────────────────────────────────────────────────────────────── HTML 4 2687 2604 43 40 |- CSS 2 543 479 37 27 |- JavaScript 1 1233 1215 12 6 (Total) 4463 4298 92 73 ───────────────────────────────────────────────────────────────────────────────── Jupyter Notebooks 4 122 83 23 16 |- Markdown 1 60 30 22 8 |- Python 1 122 113 1 8 (Total) 304 226 46 32 ───────────────────────────────────────────────────────────────────────────────── Markdown 204 45330 0 35183 10147 |- BASH 72 1655 1203 331 121 |- C 3 17 17 0 0 |- CUDA 2 84 56 16 12 |- JSON 18 708 708 0 0 |- PowerShell 1 1 1 0 0 |- Python 23 1008 787 113 108 |- Rust 66 2051 1716 77 258 |- TOML 6 207 164 0 43 |- YAML 5 41 36 5 0 (Total) 51102 4688 35725 10689 ───────────────────────────────────────────────────────────────────────────────── Rust 673 331228 285444 16819 28965 |- Markdown 491 29538 471 25464 3603 (Total) 360766 285915 42283 32568 ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Total 1324 495675 352655 90603 52417 ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ |
|
CI fixed. Not merging — one blocking polarity change, plus a claim that needs hardware. What I fixed ( Credit where due: replacing the 🔴 Blocking: default-ON.
Also: the −3.44 ms/token was measured on this branch's base, not today's master. Re-measure before quoting. #184 stacks on this and inherits the same guard fix. |
…o one launch Ports vLLM's csrc/activation_kernels.cu `silu_and_mul_clamp` (and SGLang's deepseek_v4/silu_and_mul_masked_post_quant.cuh `silu_and_mul<kApplySwigluLimit>`), both Apache-2.0, attributed in the file header. Replaces the 8-launch candle chain (2 casts + 3 clamp binaries + silu + mul + cast, plus five 1-element H2D copies for the scalar clamp operands) with a single kernel, on both the routed-expert and shared-expert paths. Compiled by the dedicated IEEE (no fast-math) builder so it stays bit-identical to candle-kernels. Also replaces sinkhorn.cu's vacuous `#if defined(__USE_FAST_MATH__)` #error — nvcc 12.4 defines no such macro in either pass — with assert_ieee_kernel_flags() in build.rs. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The guard asserted the literal `.exclude(&["sinkhorn.cu"])`, which this branch replaced with a named `IEEE_SOURCES` list so swiglu_clamp.cu could join the same no-fast-math contract. Assert the invariant at both ends instead of the old spelling: the fast-math builder still excludes IEEE_SOURCES, and sinkhorn.cu is still in it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
791eba0 to
9fa561f
Compare
Retargeted at the integration branchBase changed: The queue is being restructured to the shape the owner asked for: one PR open against Two things had to land on
This PR was not closed and is not considered stale. An audit of the queue found the overwhelming majority of it to be real work that was never merged, not noise. What you need to do: rebase onto 📚 Stack order — this is the BOTTOMThis PR and its dependent were both retargeted at Merge this PR FIRST. Its dependent carries changes that assume this one is already in the branch, and merging them out of order will produce conflicts or silently drop this PR's changes. |
Ports vLLM's
csrc/activation_kernels.cusilu_and_mul_clamp(launcher at:298, layer atactivation.py:200) and SGLang'sdeepseek_v4/silu_and_mul_masked_post_quant.cuh:55-73silu_and_mul<kApplySwigluLimit = true>. Both Apache-2.0; notice and line cites are in the file header.V4 ran its SwiGLU clamp as a candle op chain on both the routed and shared expert paths:
plus a one-element host-to-device copy for every scalar clamp operand (candle's
binary_op_scalar!builds the operand on the CPU andto_devices it). This collapses all of it into one launch.Measured on H200, DeepSeek-V4-Flash qtip2 UQFF, b=1 decode, 200 tokens
Same binary both legs;
ARC_NO_FUSED_SWIGLU=1selects the old chain.-3.44 ms/token (17.21 -> 18.29 tok/s, +5.6%).
nsys, per decode token: -486.7 kernel launches, -486.0 H2D copies, -969.3 device allocations.
Bit-identity
4a3055e1...).The kernel is compiled by the dedicated IEEE (no fast-math) builder, because
--use_fast_mathprovably rewrites this expression (measured PTX, nvcc 12.4/sm_90): accurateexpf+div.rn.f32becomeex2.approx.ftz.f32+div.approx.ftz.f32.One deliberate numeric change, stated: the shared expert previously ran its silu through
mistralrs-quant's fast-mathfused_glu_f32, so it used the__expfapproximation. It now uses accurateexpf, making it bit-identical to the routed experts and to the upstream reference, which it was not before.Also fixes a vacuous guard
sinkhorn.cucarries#if defined(__USE_FAST_MATH__) #error .... That macro does not exist. Probed directly on this toolchain:__USE_FAST_MATH__,__CUDA_FAST_MATH__,__FAST_MATH__,__CUDACC_FAST_MATH__,__CUDA_PREC_DIV__and__CUDA_FTZ__are all undefined in both the host and device pass, with and without the flag — so that guard has never fired. Replaced withassert_ieee_kernel_flags()inbuild.rs, which asserts the flag sets where the regression is actually introduced. Proved red: injecting--use_fast_mathintoIEEE_ARGSpanics the build before nvcc runs.