feat(turboquant): CUDA kernels at head_dim 64/128/256/512, Hopper + Blackwell - #94
Conversation
Code Metrics Report━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Language Files Lines Code Comments Blanks ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ C Header 5 305 210 52 43 CSS 2 1181 1036 34 111 CUDA 81 29475 20161 6275 3039 Dockerfile 1 39 22 8 9 JavaScript 16 3546 2676 482 388 Jinja2 7 694 656 5 33 JSON 75 4896 4893 0 3 Makefile 1 6 5 0 1 Metal Shading Lan| 33 12224 9431 1142 1651 PowerShell 1 300 227 30 43 Python 147 15371 12677 824 1870 Shell 42 10304 6858 2769 677 Plain Text 4 3801 0 2479 1322 TOML 33 1498 1294 54 150 YAML 3 25 23 2 0 ───────────────────────────────────────────────────────────────────────────────── HTML 4 2687 2604 43 40 |- CSS 2 543 479 37 27 |- JavaScript 1 1233 1215 12 6 (Total) 4463 4298 92 73 ───────────────────────────────────────────────────────────────────────────────── Jupyter Notebooks 4 122 83 23 16 |- Markdown 1 60 30 22 8 |- Python 1 122 113 1 8 (Total) 304 226 46 32 ───────────────────────────────────────────────────────────────────────────────── Markdown 213 48932 0 38081 10851 |- BASH 72 1655 1203 331 121 |- C 4 19 19 0 0 |- CUDA 2 84 56 16 12 |- JSON 19 779 779 0 0 |- PowerShell 1 1 1 0 0 |- Python 23 1008 787 113 108 |- Rust 68 2063 1727 78 258 |- TOML 6 207 164 0 43 |- YAML 5 41 36 5 0 (Total) 54789 4772 38624 11393 ───────────────────────────────────────────────────────────────────────────────── Rust 681 340988 293252 18096 29640 |- Markdown 504 32562 471 28063 4028 (Total) 373550 293723 46159 33668 ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Total 1353 516771 363188 99077 54506 ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ |
Post-review fixes (5bfebd6)An independent CUDA review re-derived the indexing in Python and type-checked the kernel bodies against a shim. It found no compile errors and confirmed the head_dim generalization is exact — the vectorized gather reproduces the original scalar It found three real problems, now fixed:
Still-open items from that review, deliberately not fixed here
Verification without nvccNo GPU and no nvcc on the authoring box, so the CUDA remains compiled-in-theory only. What was actually run locally: the C preprocessor over both
|
Gate corrected (05282f0) — D18 exit codes + the sourced-preflight landmineThe gate exiting Exit codes. Nine conditions now exit Preflight lookup is now branch-independent: The if [ "$_ARC_PF_FAILED" = "1" ]; then
[ "$_arc_pf_sourced" = "1" ] || exit 1Sourced on a failing box, Reordered so the D16 proof is STEP 1 and needs no full workspace build — PR #98 already showed these kernels compile clean for sm_90 in 8m53s, so the native build is confirmation, not discovery, and is skippable. STEP 1 also now asserts on Verified locally against a stub reproducing the real preflight's sourced-exit structure — all four environment cases return 2, syntax clean. |
|
Blocking on a silent default change — the kernels are welcome, the default is not.
config.k_head_dim() == TURBOQUANT_HEAD_DIM // 128with mistralrs_quant::turboquant::cuda_supports_head_dim(config.k_head_dim()) // {64,128,256,512}So any paged-attention model at head_dim 64 or 256 that today falls back to The tell is in this PR's own fixture: the "expected to fall back" list changes from This is the wave43-BU shape exactly, which #98 cites by name as the reason it refuses to default itself on: FP8 KV shipped default-on and unmeasured, and every V4 request died. Every other PR in this batch is scrupulously opt-in; this is the one that isn't. Suggested fix, which keeps everything valuable here: leave default resolution at head_dim 128 and let the wider set require an explicit Ordering dependency: #97 fixes an inverted Hardware gate is queued on the H200 ( |
|
Narrowed the default — pushed as
Nothing about the kernel work changed. The six The unsupported-geometry message now separates two facts that were being collapsed: "no kernel exists" versus "a kernel exists but is unmeasured, so the default will not choose it for you." The second tells a user to opt in; the first tells them to stop. On the tests. The old fixture had the regression written into it — the expected-to-fall-back list moved Both pins were mutation-checked by reverting the const to the full set. Worth recording that the first revision of one of them passed under that mutation — it skipped widths that were in the default set, so widening the set emptied its loop and it asserted nothing. That is the same silent-success shape the PR is about, reproduced inside its own test. It now asserts the loop is non-empty before iterating, and both pins fail under mutation and pass restored. Local: 377/377 |
…V4 KV branch
One conflict, in the V4 attention KV-append arm: master added an
`arc_profiler::device_span("kv_cache_append")` exactly where this branch
added the TurboQuant storage branch.
Resolved by keeping both and giving the TurboQuant arm its own span,
`kv_cache_append_turbo`. Eviction is the single cost this path adds to
decode, and the ON/OFF decode comparison this branch has to clear is a
throughput claim — folding eviction into the dense span's number would
have hidden the one line item the measurement is meant to expose.
|
Merged master in and corrected the public record —
The first three are straightforward: V4 reports The fourth is the one worth reading twice, because it is not this PR's doing: pub const PAGED_CACHE_TYPE: PagedCacheType = PagedCacheType::TurboQuant; // :113
So the README told readers TurboQuant never runs unless asked, while the shipped default ran it on the most common geometry there is. Nothing here changes behaviour — docs plus one clap doc comment. |
|
Took the moat agent's cudaforge 0.1.5 emits its own Our loop applied the same Fixed by hoisting the Scope, deliberately partial. This is the dedupe half of #108's The silent-arch hole is covered in the meantime, just not by the build: One thing I could not settle: whether nvcc rejects or tolerates the duplicate flag. macOS cannot run the cuda build path, so I am not claiming it was fatal. Emitting it once is correct under either answer, which is why it was worth doing without waiting for that. |
Merge sequencing: #111 lands first, then this — and one sentence here must go#111 ( #111 goes first because this PR is blocked on its H200 gate anyway (the archive must contain The blocking item
That is false, and it needs to come out in this PR's conflict resolution — not be left to merge order to overwrite. There is one measured serving run, and I verified it against the object store rather than taking it secondhand:
One correction to the provenance as it was handed to me: What stays as-is
Net effect on the composite: default-on and bounded to 128 (this PR's structure) plus one real measured run, correctly scoped (#111's provenance). The sentence above is the only thing here pulling against that. The narrowing commit and its mutation-checked tests are unaffected by any of this; the hardware gate remains the only other open item. |
|
Cleared: the Sequence of what happened here, because the stale status is exactly the thing that gets quoted an hour later by someone who wasn't in the thread:
🔑 The lesson I'm recording on the PR rather than just in my own notes: cancelling CI does not merely discard information — it WRITES a false negative into the PR's permanent check record. I had told the coordinator "nothing informative was destroyed", which was true about the runs and false about the record. Now: master is green ( Judge this PR on the new run. If you are reading the old aggregate, it is describing my cancellation, not this code. |
|
Triage verdict: REBASE, not close. Still carries unique work. Evidence against current master: Conflicts are small: 2 files — Two things the rebase must handle explicitly, given what has been learned since this PR was opened:
Leaving open with the decision recorded. Not merging it blind. |
Retargeted at the integration branchBase changed: The queue is being restructured to the shape the owner asked for: one PR open against Two things had to land on
This PR was not closed and is not considered stale. An audit of the queue found the overwhelming majority of it to be real work that was never merged, not noise. What you need to do: rebase onto 📚 Stack order — this is the BOTTOMThis PR and its dependent were both retargeted at Merge this PR FIRST. Its dependent carries changes that assume this one is already in the branch, and merging them out of order will produce conflicts or silently drop this PR's changes. |
…/512 TurboQuant's paged kernels were hardcoded to head_dim 128 — `constexpr int HS = 128`, a single `SGN[128]` sign table, `rotate128`, and six `extern "C"` launchers each opening with `if (hs != 128) return;`. That early return is a silent no-op: the output buffer is left uninitialized, so a mismatched head dim produces garbage *fast* rather than failing. Kernels - `tq_attn`, `tq_attn_clocked`, `tq_cache_k`, `tq_cache_v` and the two clocked cache variants now take HS as a template parameter. NT stays 128 for every HS, so the warp-reduction topology is unchanged; wider heads give each thread DPT = HS/NT output dims instead of adding threads. At HS=128 the emitted work is the same as before. - The per-lane packed-K gather is now one vector load (`tq_load_klane`) instead of BPL scalar byte loads. The x=16 cache swizzle keeps each lane's run contiguous and naturally aligned, which `packed_k_bytes_are_swizzle_aligned` asserts. - Every `if (hs != 128) return;` is replaced by a dispatch over the instantiated widths. D16 dual-arch - `TQ_PREFETCH` software-pipelines the K gather on SM90/SM100/SM103 and runs flat below Hopper — same arithmetic, scheduled for the part it runs on. - Launches now opt in to the arch's real dynamic shared-memory budget. The logits array scales with context length and CUDA caps dynamic shared memory at 48 KB without an explicit opt-in, so contexts past ~12k tokens could not launch at all on any arch. - `ARC_CUDA_ARCHS=90,100,103` emits cubins for Hopper *and* Blackwell rather than only the build box's GPU. Tables - Codebooks and sign diagonals are dimension-dependent, so 512 needs its own. All 20 tables are emitted by `emit_turboquant_cuda_tables`, which calls the same `generate_signs`/`get_codebook` the Rust compressor uses — correct by construction, not by transcription. The generator reproduces the shipped d=128 table byte-for-byte. - `turboquant::cuda_tables` re-derives every table at test time and diffs the checked-in header, and pins the Rust head-dim list against both the kernel's dispatch switch and the FFI wrapper's list. Nothing previously asserted that the tree's four copies of the sign table agreed with the generator or with each other. Gates - `TURBOQUANT_HEAD_DIM = 128` exact-match becomes membership in `TURBOQUANT_CUDA_HEAD_DIMS`; K and V are compressed independently and no longer have to be equal widths. Not measured on a GPU: the kernels are written and gated but this box has no nvcc, so nothing here is a claim about hardware behaviour.
… kernels Proves the four head-dim instantiations compile for sm_90a AND sm_100a/sm_103a (D16), not just the build box's GPU, and that the shipping decode path did not regress. Deliberately contains no quality A/B: TurboQuant quality is settled by prior measurement and re-proving it would spend GPU budget on a known number.
Every other .cu in this crate had a rerun-if-changed line; this one did not. Emitting any rerun-if-changed disables cargo's watch-the-whole-package default, so edits to the largest CUDA file in the crate were relying entirely on cudaforge's object cache to trigger a rebuild.
…d-kernel parity Three problems from review of the head_dim generalization. 1. 32-bit overflow on cache byte offsets. `pb*kbs`, `pb*vbs`, `bi*cbs` and `bi*vbs` were all `int`. At head_dim 512 with nkvh=8 and block_size=32, `kbs` is 65536 B, so `pb*kbs` wraps once the K cache passes 2 GB — an entirely reachable allocation on an H200, and 4x sooner than at 128. Now computed in `long long` at all eight sites. 2. The shared-memory opt-in cached a process-wide `granted` flag per kernel instantiation. The requestable maximum differs by architecture, so one flag cannot be right across a mixed fleet, and a failed request left `granted` at 0 and fell through to a launch that then failed with no diagnostic. The cache is gone; the call is cheap and now always attempted. 3. `tq_attn_clocked` kept the scalar byte gather while `tq_attn` moved to a vector load plus the arch-gated pipeline. Identical results, but the phase-2 stamps would have described a kernel nobody launches — at head_dim 512 that is 8 scalar loads measured against the 1 vector load actually executed. The loop now mirrors `tq_attn`. Verified without nvcc by type-checking the kernel bodies against a CUDA shim (clang++ -Wall -Wextra -Wshadow, both __CUDA_ARCH__ branches): clean. Still not run on a GPU.
…ine guard
The gate exited 1 when gpu_box_preflight.sh was absent. That read as a code
failure for a couple of minutes when it was really 'this machine cannot
answer'. Per D18 rule 2: 0 pass, 1 genuine failure (the only signal to act
on), 2 environment could not answer.
Nine conditions now exit 2 — missing/failing/incomplete preflight, no repo, no
nvcc, unreachable remote, un-checkout-able branch, no archive to inspect, no
cuobjdump. Four exit 1, all genuine: nvcc compile failure, a native build
failure, a test failure, and the archive missing an arch it claims to carry.
The preflight is now looked up at ARC_PREFLIGHT, then /root/arc-tools, then
/usr/local/lib/arc, and only last inside the repo. It lives on
wave61/box-preflight-shared-prefix, so a repo-relative path vanishes the moment
a runner checks out any other branch — which is exactly how this gate came to
refuse to run.
Guarded the sourced-preflight landmine directly. gpu_box_preflight.sh ends on
[ "$_arc_pf_sourced" = "1" ] || exit 1
so when SOURCED on a *failing* box that test succeeds, short-circuits the
'|| exit 1', and becomes the last command — 'source pf || handler' returns 0
and the handler never fires. The gate ignores the return value entirely and
reads _ARC_PF_FAILED, treating an unset flag as 'did not reach a verdict'
rather than as a pass.
Also reordered so the sm_90a+sm_100a+sm_103a compile is STEP 1 and needs no
full workspace build, and it now asserts on cuobjdump output that both cubins
are actually in the archive — a build that merely succeeds is not the claim.
Exit paths verified locally against a stub reproducing the real preflight's
sourced-exit structure: all four environment cases return 2.
…hat is asked for Widening the acceptance gate from head_dim == 128 to every instantiated width was correct for what a user may REQUEST, but PagedCacheType:: TurboQuant carries #[default] (cache_engine.rs:22), so it also silently moved standard-layout models at head_dim 64/256/512 onto kernels that have never executed — and silently dropped prefix caching with them, since supports_prefix_cache() is false for every TurboQuant variant. Two regressions, neither visible at the call site, for users who never asked for TurboQuant at all. This repo has paid for that exact shape once: FP8 KV shipped default-on and unmeasured (wave43-BU) and every V4 request died. #98 cites that incident by name as its own reason not to default itself on; this takes the same choice. resolve_for_model already distinguished explicit from ambient-default in order to pick error-vs-fallback, so the fix is to let that same 'forced' flag also pick WHICH head-dim set applies: compiled -> widens what you may ask for (TURBOQUANT_CUDA_HEAD_DIMS) measured -> widens what you get unasked (TURBOQUANT_DEFAULT_HEAD_DIMS) TURBOQUANT_DEFAULT_HEAD_DIMS is [128] and moves when a hardware gate passes, not when a kernel compiles. Nothing about the kernels changes: the six 'if (hs != 128) return;' no-ops (D18 instance 3) and the 48 KB dynamic smem cap fix are untouched, and explicit --pa-cache-type turboquant still reaches every instantiated width including V4's 512. The unsupported-geometry message now separates 'no kernel exists' from 'a kernel exists but is unmeasured, so the default will not choose it for you' — the second tells the user to opt in rather than to give up. Tests: the old fixture asserted the regression (its expected-to-fall-back list moved [64,96,192,256] -> [96,192,320,1024] precisely because 64 and 256 stopped falling back). Replaced with an explicit-path test over every instantiated width, plus two pins on the default path. Both pins were mutation-checked by reverting the const to the full set: both fail, and the first revision of one passed vacuously because widening the const emptied its loop, so it now asserts the loop is non-empty before iterating.
This PR branched before #101 ("correct the public record") landed, so it had never seen that text. Merging master in makes four README claims false — three falsified by this PR's own kernels, one that was already wrong when it was written. Falsified by this PR: - L31 "the paged kernel exists at head_dim 128 only" - L31 "there is no kernel at head_dim 512, so DeepSeek V4 cannot use it" DeepSeek V4 reports KvCacheLayout::Standard at head_dim 512 (normal_loaders.rs), so the 512 instantiation genuinely reaches it. - L141 an explicit `--pa-cache-type turboquant` off-128 "is a hard error" It is now accepted at any instantiated width; the hard error moved out to the uninstantiated set. Already wrong before this PR: - L141 "TurboQuant is not the default anywhere" `defaults::PAGED_CACHE_TYPE` is `PagedCacheType::TurboQuant` and the CLI's `--pa-cache-type` has no clap default, so leaving it unset on CUDA gives a standard-layout head_dim-128 model TurboQuant KV with no flag — and silently drops prefix caching, which no TurboQuant variant supports. That was true on master; the narrowing in a3a5faa bounds it to 128 but does not remove it. Stating the opposite understated the risk to anyone reading the README to decide whether to opt out, so it is corrected in place rather than quietly reworded, and the Rust example's comment now says `Auto` opts *out* rather than implying TurboQuant is off until asked for. Also refreshes the `--pa-cache-type` help text, which still described the explicit path as 128-or-error. No behaviour change: documentation and one clap doc comment.
…n arch
`mistralrs-paged-attn/build.rs` appended one `-gencode` per entry in
`ARC_CUDA_ARCHS` on top of the one cudaforge derives itself, and applied
the same `a` suffix for cap >= 90 that cudaforge's `GpuArch::auto_suffix`
does. Verified against cudaforge 0.1.5 rather than assumed:
compute_cap.rs auto_suffix(90).to_gencode_arg()
-> "-gencode=arch=compute_90a,code=sm_90a"
builder.rs:401 let gencode_arg = gpu_arch.to_gencode_arg();
builder.rs:405 command.arg(&gencode_arg)...
builder.rs:412 for arg in &self.extra_args { command.arg(arg); }
cudaforge emits its own gencode first and then appends `extra_args`
verbatim, so on an H200 building `ARC_CUDA_ARCHS=90,100,103` nvcc received
`-gencode=arch=compute_90a,code=sm_90a` twice, byte-identical.
Fixed by hoisting the existing `get_compute_cap()` binding above the loop
(it is what decides which requested arches are genuinely extra) and
skipping the one cudaforge already covers. The later duplicate binding is
deleted; there is now one.
This is the dedupe half of #108's `arc_target::build::split_primary()`,
done locally because `arc-target` does not exist on this branch. The other
half — `verify_and_export()`, the `cuobjdump` gate that fails a build whose
archive is missing a requested arch — genuinely needs that crate and is
left to #108, which should replace this block with the two calls. Until it
lands, `arc-tools/wave64_v4_turboquant_kv_gate.sh` STEP 2 asserts the arch
list externally with `cuobjdump --list-elf`, so a missing arch is caught by
the gate even though it is not yet caught by the build.
Not verified here: whether nvcc rejects or tolerates the duplicate flag.
macOS cannot run the cuda build path. Emitting it once is correct either
way, which is why this is worth doing without waiting on that answer.
0c71381 to
be7764b
Compare
What was actually blocking TurboQuant
static_assert(HEAD_SIZE == 128)in the.cuhis dead code — that header's kernel has no callers. The live kernel isturbo_paged_attention.cu, and its limit was worse than an assert: sixextern "C"launchers each opened withif (hs != 128) return;. That is a silent no-op. The output buffer is left uninitialized, so a mismatched head dim produces garbage quickly — a broken run reports a better number, not a worse one.Kernels
tq_attn,tq_attn_clocked,tq_cache_k,tq_cache_vand both clocked cache variants now takeHSas a template parameter, instantiated at 64, 128, 256 and 512.NTstays 128 at everyHS, so the warp-reduction topology (4 warps,red[]) is unchanged. Wider heads give each threadDPT = HS/NToutput dims rather than adding threads. AtHS=128the emitted work is what it was.tq_load_klane) instead ofBPLscalar byte loads. Thex=16cache swizzle keeps each lane's run contiguous and naturally aligned;packed_k_bytes_are_swizzle_alignedasserts the property the load depends on rather than trusting it.if (hs != 128) return;is gone, replaced by a dispatch over the instantiated widths.D16 — Hopper and Blackwell
The file had zero
__CUDA_ARCH__guards, and nothing in the workspace passed an explicit-gencode:cudaforgederives one arch fromCUDA_COMPUTE_CAPor the localnvidia-smi, so a binary built on an H200 carries SM90 only.TQ_PREFETCHsoftware-pipelines the K gather on SM90/SM100/SM103 and runs flat below Hopper. Same arithmetic, scheduled for the part it runs on.ARC_CUDA_ARCHS=90,100,103emits cubins for both architectures. Left unset the build is unchanged, because each extra arch recompiles every.cuand that is not worth paying on a rented box that runs one arch.Tables — generated, not transcribed
Codebooks and sign diagonals are both functions of the dimension, so 512 needed its own; the centroids shrink as
sqrt(1/d). All 20 tables come fromemit_turboquant_cuda_tables, which calls the samegenerate_signs/get_codebookthe Rust compressor uses. The generator reproduces the shipped d=128 sign table byte-for-byte, which is the check that it is the same format.turboquant::cuda_tablesre-derives every table at test time and diffs the checked-in header, and pins the Rust head-dim list against both the kernel's dispatch switch and the FFI wrapper's list. Previously nothing asserted that the tree's four separate copies of the sign table agreed with the generator or with each other, despite a comment inturbo_wht.cuclaiming they must.Gates
TURBOQUANT_HEAD_DIM = 128exact-match becomes membership inTURBOQUANT_CUDA_HEAD_DIMS. K and V are compressed independently, so they no longer have to be equal widths either.What this does not do — the honest gap
This does not put DeepSeek-V4 on TurboQuant, and the kernel was never what stood in the way.
DeepSeekV4Loader::supports_paged_attentionreturnsOk(false)(normal_loaders.rs:3231). V4 uses the eagerNormalCacheexclusively, so the paged TurboQuant kernel is unreachable for it at any head dim. The two things that actually block it are unrelated to head size:cache_write_and_gatherreturns a varlen pack[1, H, sum(seqlen), D]thatdsv4_attentionwould read as one sequence, and thexshistory slots are per-sequence only throughNormalCacheManager::clone_in_cache, which the paged engine arm never issues.Separately, V4 stores no V at all — V is K bit-for-bit under the fused
wkv, and the V half of the slot is a 1-wide zero marker (V4_V_MARKER_WIDTH = 1). So the K+V shape TurboQuant compresses does not map onto V4's KV as-is; the integration point isV4CachedK, alongsideDenseand the FP8Packed, not the paged kernel.The eager path is also not a shortcut:
TurboQuantSingleCache::current_datareconstructs the entire compressed context on the host every decode step, andrequire_normal_kv_slotrejects a TurboQuant slot for V4 outright.Making V4 default-on means teaching
V4CachedKa TurboQuant variant. That is a KV storage change, and this repo already has the scar from shipping one unmeasured — wave43-BU turned FP8 KV on by default without a GPU run and every V4 request died on a dtype mismatch, which is whyARC_V4_FP8_KVis opt-in today. It should be its own change with a GPU run attached.Verification
cargo checkgreen;cargo test -p mistralrs-core --lib375 passed;cargo test -p mistralrs-quant turboquant::32 passed.arc-tools/wave63_turboquant_hd512_gate.shis the D15 script that compiles all four instantiations for sm_90a and sm_100a/sm_103a and checks the shipping decode path for regression. It contains no quality A/B by design.