Repository navigation
Conversation
…/512 TurboQuant's paged kernels were hardcoded to head_dim 128 — `constexpr int HS = 128`, a single `SGN[128]` sign table, `rotate128`, and six `extern "C"` launchers each opening with `if (hs != 128) return;`. That early return is a silent no-op: the output buffer is left uninitialized, so a mismatched head dim produces garbage *fast* rather than failing. Kernels - `tq_attn`, `tq_attn_clocked`, `tq_cache_k`, `tq_cache_v` and the two clocked cache variants now take HS as a template parameter. NT stays 128 for every HS, so the warp-reduction topology is unchanged; wider heads give each thread DPT = HS/NT output dims instead of adding threads. At HS=128 the emitted work is the same as before. - The per-lane packed-K gather is now one vector load (`tq_load_klane`) instead of BPL scalar byte loads. The x=16 cache swizzle keeps each lane's run contiguous and naturally aligned, which `packed_k_bytes_are_swizzle_aligned` asserts. - Every `if (hs != 128) return;` is replaced by a dispatch over the instantiated widths. D16 dual-arch - `TQ_PREFETCH` software-pipelines the K gather on SM90/SM100/SM103 and runs flat below Hopper — same arithmetic, scheduled for the part it runs on. - Launches now opt in to the arch's real dynamic shared-memory budget. The logits array scales with context length and CUDA caps dynamic shared memory at 48 KB without an explicit opt-in, so contexts past ~12k tokens could not launch at all on any arch. - `ARC_CUDA_ARCHS=90,100,103` emits cubins for Hopper *and* Blackwell rather than only the build box's GPU. Tables - Codebooks and sign diagonals are dimension-dependent, so 512 needs its own. All 20 tables are emitted by `emit_turboquant_cuda_tables`, which calls the same `generate_signs`/`get_codebook` the Rust compressor uses — correct by construction, not by transcription. The generator reproduces the shipped d=128 table byte-for-byte. - `turboquant::cuda_tables` re-derives every table at test time and diffs the checked-in header, and pins the Rust head-dim list against both the kernel's dispatch switch and the FFI wrapper's list. Nothing previously asserted that the tree's four copies of the sign table agreed with the generator or with each other. Gates - `TURBOQUANT_HEAD_DIM = 128` exact-match becomes membership in `TURBOQUANT_CUDA_HEAD_DIMS`; K and V are compressed independently and no longer have to be equal widths. Not measured on a GPU: the kernels are written and gated but this box has no nvcc, so nothing here is a claim about hardware behaviour.
… kernels Proves the four head-dim instantiations compile for sm_90a AND sm_100a/sm_103a (D16), not just the build box's GPU, and that the shipping decode path did not regress. Deliberately contains no quality A/B: TurboQuant quality is settled by prior measurement and re-proving it would spend GPU budget on a known number.
Every other .cu in this crate had a rerun-if-changed line; this one did not. Emitting any rerun-if-changed disables cargo's watch-the-whole-package default, so edits to the largest CUDA file in the crate were relying entirely on cudaforge's object cache to trigger a rebuild.
…d-kernel parity Three problems from review of the head_dim generalization. 1. 32-bit overflow on cache byte offsets. `pb*kbs`, `pb*vbs`, `bi*cbs` and `bi*vbs` were all `int`. At head_dim 512 with nkvh=8 and block_size=32, `kbs` is 65536 B, so `pb*kbs` wraps once the K cache passes 2 GB — an entirely reachable allocation on an H200, and 4x sooner than at 128. Now computed in `long long` at all eight sites. 2. The shared-memory opt-in cached a process-wide `granted` flag per kernel instantiation. The requestable maximum differs by architecture, so one flag cannot be right across a mixed fleet, and a failed request left `granted` at 0 and fell through to a launch that then failed with no diagnostic. The cache is gone; the call is cheap and now always attempted. 3. `tq_attn_clocked` kept the scalar byte gather while `tq_attn` moved to a vector load plus the arch-gated pipeline. Identical results, but the phase-2 stamps would have described a kernel nobody launches — at head_dim 512 that is 8 scalar loads measured against the 1 vector load actually executed. The loop now mirrors `tq_attn`. Verified without nvcc by type-checking the kernel bodies against a CUDA shim (clang++ -Wall -Wextra -Wshadow, both __CUDA_ARCH__ branches): clean. Still not run on a GPU.
V4's raw K cache and TurboQuant are the same two-region shape with the same boundary, so this builds them as one mechanism rather than two. `dsv4_attention` reads raw keys over exactly the trailing `t_q + window - 1` tokens; everything older is `-inf` on every query row, because V4's raw branch is a sliding window and distant context arrives through `compressed_kv` instead. So TurboQuant's `fp16_window` IS V4's sliding window and its compressed region IS everything older. The consequence is that nothing on the decode path is ever dequantized: `span` over the reachable window is a narrow of a tensor that was never compressed. * `KvCache::V4Turbo` / `V4TurboKCache`, modelled on `XsRollingCache` — a grown capacity buffer of packed records plus a bounded dense tail, reporting length in tokens and mapping truncation onto both time bases. Norms ride inside the code record because `BatchSrc` gives a slot exactly two tensors. * `V4CachedK::Turbo`, wired into the eager decode arm behind `ARC_V4_TURBOQUANT=1`. 260 B/token at V4-Flash's geometry against 1,026 dense and 590 for FP8 KV. * Eviction is chunked (`EVICT_CHUNK = 128`): `quantize_into` is a host function, so per-step eviction would be ~86 device syncs per token. * `ARC_V4_TURBOQUANT` + `ARC_V4_STANDARD_DENSE` is a hard error — that ablation makes ratio-0 layers read past the window, which is the assumption eviction rests on. * One `Once`-guarded log line on the first eviction: flag-on and flag-on-but-never-compressing are otherwise indistinguishable from outside the process. Also fixes two pre-existing silently-ignored kernel parameters: `qs` (query stride) is now honoured — `query[sidx*qs + ...]`, a no-op for the contiguous queries every current caller passes — with the head/dim strides the kernel still assumes now checked in the Rust wrapper; and `softcapping` is applied with the same arithmetic and at the same position as `pagedattention.cuh:276`. Default stays off. wave43-BU is why. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Code Metrics Report━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Language Files Lines Code Comments Blanks ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ C Header 5 305 210 52 43 CSS 2 1181 1036 34 111 CUDA 72 24328 17592 4018 2718 Dockerfile 1 39 22 8 9 JavaScript 16 3546 2676 482 388 Jinja2 7 694 656 5 33 JSON 74 4600 4597 0 3 Makefile 1 6 5 0 1 Metal Shading Lan| 33 12224 9431 1142 1651 PowerShell 1 300 227 30 43 Python 143 14830 12217 797 1816 Shell 22 5235 3623 1280 332 Plain Text 4 3801 0 2479 1322 TOML 33 1485 1292 43 150 YAML 3 25 23 2 0 ───────────────────────────────────────────────────────────────────────────────── HTML 4 2687 2604 43 40 |- CSS 2 543 479 37 27 |- JavaScript 1 1233 1215 12 6 (Total) 4463 4298 92 73 ───────────────────────────────────────────────────────────────────────────────── Jupyter Notebooks 4 122 83 23 16 |- Markdown 1 60 30 22 8 |- Python 1 122 113 1 8 (Total) 304 226 46 32 ───────────────────────────────────────────────────────────────────────────────── Markdown 187 37238 0 28565 8673 |- BASH 68 1610 1185 309 116 |- C 1 10 10 0 0 |- CUDA 2 84 56 16 12 |- JSON 18 708 708 0 0 |- PowerShell 1 1 1 0 0 |- Python 23 1008 787 113 108 |- Rust 65 2048 1713 77 258 |- TOML 6 207 164 0 43 |- YAML 4 38 33 5 0 (Total) 42952 4657 29085 9210 ───────────────────────────────────────────────────────────────────────────────── Rust 660 305610 264700 13744 27166 |- Markdown 475 22145 471 19045 2629 (Total) 327755 265171 32789 29795 ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Total 1272 448073 327959 72384 47730 ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ |
…ine guard
The gate exited 1 when gpu_box_preflight.sh was absent. That read as a code
failure for a couple of minutes when it was really 'this machine cannot
answer'. Per D18 rule 2: 0 pass, 1 genuine failure (the only signal to act
on), 2 environment could not answer.
Nine conditions now exit 2 — missing/failing/incomplete preflight, no repo, no
nvcc, unreachable remote, un-checkout-able branch, no archive to inspect, no
cuobjdump. Four exit 1, all genuine: nvcc compile failure, a native build
failure, a test failure, and the archive missing an arch it claims to carry.
The preflight is now looked up at ARC_PREFLIGHT, then /root/arc-tools, then
/usr/local/lib/arc, and only last inside the repo. It lives on
wave61/box-preflight-shared-prefix, so a repo-relative path vanishes the moment
a runner checks out any other branch — which is exactly how this gate came to
refuse to run.
Guarded the sourced-preflight landmine directly. gpu_box_preflight.sh ends on
[ "$_arc_pf_sourced" = "1" ] || exit 1
so when SOURCED on a *failing* box that test succeeds, short-circuits the
'|| exit 1', and becomes the last command — 'source pf || handler' returns 0
and the handler never fires. The gate ignores the return value entirely and
reads _ARC_PF_FAILED, treating an unset flag as 'did not reach a verdict'
rather than as a pass.
Also reordered so the sm_90a+sm_100a+sm_103a compile is STEP 1 and needs no
full workspace build, and it now asserts on cuobjdump output that both cubins
are actually in the archive — a build that merely succeeds is not the claim.
Exit paths verified locally against a stub reproducing the real preflight's
sourced-exit structure: all four environment cases return 2.
… faults The wave64 gate passed `--features "cuda flash-attn"` to `-p mistralrs-paged-attn`, which declares only `cuda`/`metal`. STEP 2 died before it ran, so D16 dual-arch was never actually tested. STEP 1 — the first compile of the qs/softcapping kernel fix — passed on an H200. * STEP 2 now passes `--features cuda` only. * A feature-sanity check runs BEFORE the 9-minute build, reading the manifests themselves, so this class of bug costs seconds rather than a gate run. * STEP 2 now ASSERTS the cubin arches instead of printing them. v1 would have reported a missing sm_100a and still said PASS, which is the same absence-read-as-success shape the gate exists to catch. * STEP 3 reuses STEP 1's feature set so it does not force a second full non-cuda rebuild. * Exit codes (D18 rule 2): 1 means the code under test failed, 2 means environment/harness — missing toolkit, model, repo, or a feature flag this script got wrong. A missing UQFF now exits 2 rather than 0, so a partial run can never read as a pass. Both new logic blocks were checked against the real manifests and against synthetic arch lists before committing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
🟢 First hardware result: the never-before-compiled CUDA compilesGate ran on a live H200. That was the step that mattered — my own framing was "STEP 1 failing = the .cu edit is broken". It didn't fail. The 🔴 STEP 2 died on a bug in my gate script, not in the code
STEPS 3–7 remain outstanding, especially STEP 6 (engagement) — "flag on" and "flag on but silently no-opping" are otherwise indistinguishable from outside the process. Follow-ups (issues are disabled on this repo, so: here +
|
…hat is asked for Widening the acceptance gate from head_dim == 128 to every instantiated width was correct for what a user may REQUEST, but PagedCacheType:: TurboQuant carries #[default] (cache_engine.rs:22), so it also silently moved standard-layout models at head_dim 64/256/512 onto kernels that have never executed — and silently dropped prefix caching with them, since supports_prefix_cache() is false for every TurboQuant variant. Two regressions, neither visible at the call site, for users who never asked for TurboQuant at all. This repo has paid for that exact shape once: FP8 KV shipped default-on and unmeasured (wave43-BU) and every V4 request died. #98 cites that incident by name as its own reason not to default itself on; this takes the same choice. resolve_for_model already distinguished explicit from ambient-default in order to pick error-vs-fallback, so the fix is to let that same 'forced' flag also pick WHICH head-dim set applies: compiled -> widens what you may ask for (TURBOQUANT_CUDA_HEAD_DIMS) measured -> widens what you get unasked (TURBOQUANT_DEFAULT_HEAD_DIMS) TURBOQUANT_DEFAULT_HEAD_DIMS is [128] and moves when a hardware gate passes, not when a kernel compiles. Nothing about the kernels changes: the six 'if (hs != 128) return;' no-ops (D18 instance 3) and the 48 KB dynamic smem cap fix are untouched, and explicit --pa-cache-type turboquant still reaches every instantiated width including V4's 512. The unsupported-geometry message now separates 'no kernel exists' from 'a kernel exists but is unmeasured, so the default will not choose it for you' — the second tells the user to opt in rather than to give up. Tests: the old fixture asserted the regression (its expected-to-fall-back list moved [64,96,192,256] -> [96,192,320,1024] precisely because 64 and 256 stopped falling back). Replaced with an explicit-path test over every instantiated width, plus two pins on the default path. Both pins were mutation-checked by reverting the const to the full set: both fail, and the first revision of one passed vacuously because widening the const emptied its loop, so it now asserts the loop is non-empty before iterating.
This PR branched before #101 ("correct the public record") landed, so it had never seen that text. Merging master in makes four README claims false — three falsified by this PR's own kernels, one that was already wrong when it was written. Falsified by this PR: - L31 "the paged kernel exists at head_dim 128 only" - L31 "there is no kernel at head_dim 512, so DeepSeek V4 cannot use it" DeepSeek V4 reports KvCacheLayout::Standard at head_dim 512 (normal_loaders.rs), so the 512 instantiation genuinely reaches it. - L141 an explicit `--pa-cache-type turboquant` off-128 "is a hard error" It is now accepted at any instantiated width; the hard error moved out to the uninstantiated set. Already wrong before this PR: - L141 "TurboQuant is not the default anywhere" `defaults::PAGED_CACHE_TYPE` is `PagedCacheType::TurboQuant` and the CLI's `--pa-cache-type` has no clap default, so leaving it unset on CUDA gives a standard-layout head_dim-128 model TurboQuant KV with no flag — and silently drops prefix caching, which no TurboQuant variant supports. That was true on master; the narrowing in a3a5faa bounds it to 128 but does not remove it. Stating the opposite understated the risk to anyone reading the README to decide whether to opt out, so it is corrected in place rather than quietly reworded, and the Rust example's comment now says `Auto` opts *out* rather than implying TurboQuant is off until asked for. Also refreshes the `--pa-cache-type` help text, which still described the explicit path as 128-or-error. No behaviour change: documentation and one clap doc comment.
…V4 KV branch
One conflict, in the V4 attention KV-append arm: master added an
`arc_profiler::device_span("kv_cache_append")` exactly where this branch
added the TurboQuant storage branch.
Resolved by keeping both and giving the TurboQuant arm its own span,
`kv_cache_append_turbo`. Eviction is the single cost this path adds to
decode, and the ON/OFF decode comparison this branch has to clear is a
throughput claim — folding eviction into the dense span's number would
have hidden the one line item the measurement is meant to expose.
…n arch
`mistralrs-paged-attn/build.rs` appended one `-gencode` per entry in
`ARC_CUDA_ARCHS` on top of the one cudaforge derives itself, and applied
the same `a` suffix for cap >= 90 that cudaforge's `GpuArch::auto_suffix`
does. Verified against cudaforge 0.1.5 rather than assumed:
compute_cap.rs auto_suffix(90).to_gencode_arg()
-> "-gencode=arch=compute_90a,code=sm_90a"
builder.rs:401 let gencode_arg = gpu_arch.to_gencode_arg();
builder.rs:405 command.arg(&gencode_arg)...
builder.rs:412 for arg in &self.extra_args { command.arg(arg); }
cudaforge emits its own gencode first and then appends `extra_args`
verbatim, so on an H200 building `ARC_CUDA_ARCHS=90,100,103` nvcc received
`-gencode=arch=compute_90a,code=sm_90a` twice, byte-identical.
Fixed by hoisting the existing `get_compute_cap()` binding above the loop
(it is what decides which requested arches are genuinely extra) and
skipping the one cudaforge already covers. The later duplicate binding is
deleted; there is now one.
This is the dedupe half of #108's `arc_target::build::split_primary()`,
done locally because `arc-target` does not exist on this branch. The other
half — `verify_and_export()`, the `cuobjdump` gate that fails a build whose
archive is missing a requested arch — genuinely needs that crate and is
left to #108, which should replace this block with the two calls. Until it
lands, `arc-tools/wave64_v4_turboquant_kv_gate.sh` STEP 2 asserts the arch
list externally with `cuobjdump --list-elf`, so a missing arch is caught by
the gate even though it is not yet caught by the build.
Not verified here: whether nvcc rejects or tolerates the duplicate flag.
macOS cannot run the cuda build path. Emitting it once is correct either
way, which is why this is worth doing without waiting on that answer.
|
ArcGate: HELD on its parent — mechanically, not on merit. Base is Two things worth recording while it waits: It currently shows 1 check — the comment bot. Zero CI lanes have ever run on it. That is the stacked-PR trigger gap ( And the claim in the body is the interesting one, so it should not get lost in the queue: that Reviewers: that is the sentence to check, not the diff size. The tests pin it, but the reachability argument is what makes the whole thing worth having. Nothing deleted; branch untouched. |
|
Triage verdict: REBASE, not close — but only after #94, on which it stacks. Evidence against current master: Conflicts: 6 files, worst are Caveat for the rebase: this adds a KV storage variant. That is on-disk/in-memory format territory and the byte formats are the moat — the rebase must state whether Leaving open with the decision recorded. |
Retargeted at the integration branchBase changed: The queue is being restructured to the shape the owner asked for: one PR open against Two things had to land on
This PR was not closed and is not considered stale. An audit of the queue found the overwhelming majority of it to be real work that was never merged, not noise. What you need to do: rebase onto 📚 Stack order — this is the TOPThis PR used to be based on its parent PR's branch, which is exactly the pattern the Both are now retargeted at Merge the bottom PR FIRST, then this one. Until the bottom lands, this PR's diff against the integration branch will also contain the bottom's commits. |
…hat is asked for Widening the acceptance gate from head_dim == 128 to every instantiated width was correct for what a user may REQUEST, but PagedCacheType:: TurboQuant carries #[default] (cache_engine.rs:22), so it also silently moved standard-layout models at head_dim 64/256/512 onto kernels that have never executed — and silently dropped prefix caching with them, since supports_prefix_cache() is false for every TurboQuant variant. Two regressions, neither visible at the call site, for users who never asked for TurboQuant at all. This repo has paid for that exact shape once: FP8 KV shipped default-on and unmeasured (wave43-BU) and every V4 request died. #98 cites that incident by name as its own reason not to default itself on; this takes the same choice. resolve_for_model already distinguished explicit from ambient-default in order to pick error-vs-fallback, so the fix is to let that same 'forced' flag also pick WHICH head-dim set applies: compiled -> widens what you may ask for (TURBOQUANT_CUDA_HEAD_DIMS) measured -> widens what you get unasked (TURBOQUANT_DEFAULT_HEAD_DIMS) TURBOQUANT_DEFAULT_HEAD_DIMS is [128] and moves when a hardware gate passes, not when a kernel compiles. Nothing about the kernels changes: the six 'if (hs != 128) return;' no-ops (D18 instance 3) and the 48 KB dynamic smem cap fix are untouched, and explicit --pa-cache-type turboquant still reaches every instantiated width including V4's 512. The unsupported-geometry message now separates 'no kernel exists' from 'a kernel exists but is unmeasured, so the default will not choose it for you' — the second tells the user to opt in rather than to give up. Tests: the old fixture asserted the regression (its expected-to-fall-back list moved [64,96,192,256] -> [96,192,320,1024] precisely because 64 and 256 stopped falling back). Replaced with an explicit-path test over every instantiated width, plus two pins on the default path. Both pins were mutation-checked by reverting the const to the full set: both fail, and the first revision of one passed vacuously because widening the const emptied its loop, so it now asserts the loop is non-empty before iterating.
Rebase attempt onto
|
…lackwell (#94) * feat(turboquant): instantiate the CUDA kernels at head_dim 64/128/256/512 TurboQuant's paged kernels were hardcoded to head_dim 128 — `constexpr int HS = 128`, a single `SGN[128]` sign table, `rotate128`, and six `extern "C"` launchers each opening with `if (hs != 128) return;`. That early return is a silent no-op: the output buffer is left uninitialized, so a mismatched head dim produces garbage *fast* rather than failing. Kernels - `tq_attn`, `tq_attn_clocked`, `tq_cache_k`, `tq_cache_v` and the two clocked cache variants now take HS as a template parameter. NT stays 128 for every HS, so the warp-reduction topology is unchanged; wider heads give each thread DPT = HS/NT output dims instead of adding threads. At HS=128 the emitted work is the same as before. - The per-lane packed-K gather is now one vector load (`tq_load_klane`) instead of BPL scalar byte loads. The x=16 cache swizzle keeps each lane's run contiguous and naturally aligned, which `packed_k_bytes_are_swizzle_aligned` asserts. - Every `if (hs != 128) return;` is replaced by a dispatch over the instantiated widths. D16 dual-arch - `TQ_PREFETCH` software-pipelines the K gather on SM90/SM100/SM103 and runs flat below Hopper — same arithmetic, scheduled for the part it runs on. - Launches now opt in to the arch's real dynamic shared-memory budget. The logits array scales with context length and CUDA caps dynamic shared memory at 48 KB without an explicit opt-in, so contexts past ~12k tokens could not launch at all on any arch. - `ARC_CUDA_ARCHS=90,100,103` emits cubins for Hopper *and* Blackwell rather than only the build box's GPU. Tables - Codebooks and sign diagonals are dimension-dependent, so 512 needs its own. All 20 tables are emitted by `emit_turboquant_cuda_tables`, which calls the same `generate_signs`/`get_codebook` the Rust compressor uses — correct by construction, not by transcription. The generator reproduces the shipped d=128 table byte-for-byte. - `turboquant::cuda_tables` re-derives every table at test time and diffs the checked-in header, and pins the Rust head-dim list against both the kernel's dispatch switch and the FFI wrapper's list. Nothing previously asserted that the tree's four copies of the sign table agreed with the generator or with each other. Gates - `TURBOQUANT_HEAD_DIM = 128` exact-match becomes membership in `TURBOQUANT_CUDA_HEAD_DIMS`; K and V are compressed independently and no longer have to be equal widths. Not measured on a GPU: the kernels are written and gated but this box has no nvcc, so nothing here is a claim about hardware behaviour. * test(gpu): wave63 gate — dual-arch compile proof for the head_dim 512 kernels Proves the four head-dim instantiations compile for sm_90a AND sm_100a/sm_103a (D16), not just the build box's GPU, and that the shipping decode path did not regress. Deliberately contains no quality A/B: TurboQuant quality is settled by prior measurement and re-proving it would spend GPU budget on a known number. * fix(build): watch turbo_paged_attention.cu for changes Every other .cu in this crate had a rerun-if-changed line; this one did not. Emitting any rerun-if-changed disables cargo's watch-the-whole-package default, so edits to the largest CUDA file in the crate were relying entirely on cudaforge's object cache to trigger a rebuild. * fix(turboquant): 64-bit cache offsets, per-device smem opt-in, clocked-kernel parity Three problems from review of the head_dim generalization. 1. 32-bit overflow on cache byte offsets. `pb*kbs`, `pb*vbs`, `bi*cbs` and `bi*vbs` were all `int`. At head_dim 512 with nkvh=8 and block_size=32, `kbs` is 65536 B, so `pb*kbs` wraps once the K cache passes 2 GB — an entirely reachable allocation on an H200, and 4x sooner than at 128. Now computed in `long long` at all eight sites. 2. The shared-memory opt-in cached a process-wide `granted` flag per kernel instantiation. The requestable maximum differs by architecture, so one flag cannot be right across a mixed fleet, and a failed request left `granted` at 0 and fell through to a launch that then failed with no diagnostic. The cache is gone; the call is cheap and now always attempted. 3. `tq_attn_clocked` kept the scalar byte gather while `tq_attn` moved to a vector load plus the arch-gated pipeline. Identical results, but the phase-2 stamps would have described a kernel nobody launches — at head_dim 512 that is 8 scalar loads measured against the 1 vector load actually executed. The loop now mirrors `tq_attn`. Verified without nvcc by type-checking the kernel bodies against a CUDA shim (clang++ -Wall -Wextra -Wshadow, both __CUDA_ARCH__ branches): clean. Still not run on a GPU. * fix(gate): D18 exit codes, branch-independent preflight, source-landmine guard The gate exited 1 when gpu_box_preflight.sh was absent. That read as a code failure for a couple of minutes when it was really 'this machine cannot answer'. Per D18 rule 2: 0 pass, 1 genuine failure (the only signal to act on), 2 environment could not answer. Nine conditions now exit 2 — missing/failing/incomplete preflight, no repo, no nvcc, unreachable remote, un-checkout-able branch, no archive to inspect, no cuobjdump. Four exit 1, all genuine: nvcc compile failure, a native build failure, a test failure, and the archive missing an arch it claims to carry. The preflight is now looked up at ARC_PREFLIGHT, then /root/arc-tools, then /usr/local/lib/arc, and only last inside the repo. It lives on wave61/box-preflight-shared-prefix, so a repo-relative path vanishes the moment a runner checks out any other branch — which is exactly how this gate came to refuse to run. Guarded the sourced-preflight landmine directly. gpu_box_preflight.sh ends on [ "$_arc_pf_sourced" = "1" ] || exit 1 so when SOURCED on a *failing* box that test succeeds, short-circuits the '|| exit 1', and becomes the last command — 'source pf || handler' returns 0 and the handler never fires. The gate ignores the return value entirely and reads _ARC_PF_FAILED, treating an unset flag as 'did not reach a verdict' rather than as a pass. Also reordered so the sm_90a+sm_100a+sm_103a compile is STEP 1 and needs no full workspace build, and it now asserts on cuobjdump output that both cubins are actually in the archive — a build that merely succeeds is not the claim. Exit paths verified locally against a stub reproducing the real preflight's sourced-exit structure: all four environment cases return 2. * fix(paged): keep the TurboQuant DEFAULT at head_dim 128; widen only what is asked for Widening the acceptance gate from head_dim == 128 to every instantiated width was correct for what a user may REQUEST, but PagedCacheType:: TurboQuant carries #[default] (cache_engine.rs:22), so it also silently moved standard-layout models at head_dim 64/256/512 onto kernels that have never executed — and silently dropped prefix caching with them, since supports_prefix_cache() is false for every TurboQuant variant. Two regressions, neither visible at the call site, for users who never asked for TurboQuant at all. This repo has paid for that exact shape once: FP8 KV shipped default-on and unmeasured (wave43-BU) and every V4 request died. #98 cites that incident by name as its own reason not to default itself on; this takes the same choice. resolve_for_model already distinguished explicit from ambient-default in order to pick error-vs-fallback, so the fix is to let that same 'forced' flag also pick WHICH head-dim set applies: compiled -> widens what you may ask for (TURBOQUANT_CUDA_HEAD_DIMS) measured -> widens what you get unasked (TURBOQUANT_DEFAULT_HEAD_DIMS) TURBOQUANT_DEFAULT_HEAD_DIMS is [128] and moves when a hardware gate passes, not when a kernel compiles. Nothing about the kernels changes: the six 'if (hs != 128) return;' no-ops (D18 instance 3) and the 48 KB dynamic smem cap fix are untouched, and explicit --pa-cache-type turboquant still reaches every instantiated width including V4's 512. The unsupported-geometry message now separates 'no kernel exists' from 'a kernel exists but is unmeasured, so the default will not choose it for you' — the second tells the user to opt in rather than to give up. Tests: the old fixture asserted the regression (its expected-to-fall-back list moved [64,96,192,256] -> [96,192,320,1024] precisely because 64 and 256 stopped falling back). Replaced with an explicit-path test over every instantiated width, plus two pins on the default path. Both pins were mutation-checked by reverting the const to the full set: both fail, and the first revision of one passed vacuously because widening the const emptied its loop, so it now asserts the loop is non-empty before iterating. * docs: correct the public record for the widened TurboQuant kernels This PR branched before #101 ("correct the public record") landed, so it had never seen that text. Merging master in makes four README claims false — three falsified by this PR's own kernels, one that was already wrong when it was written. Falsified by this PR: - L31 "the paged kernel exists at head_dim 128 only" - L31 "there is no kernel at head_dim 512, so DeepSeek V4 cannot use it" DeepSeek V4 reports KvCacheLayout::Standard at head_dim 512 (normal_loaders.rs), so the 512 instantiation genuinely reaches it. - L141 an explicit `--pa-cache-type turboquant` off-128 "is a hard error" It is now accepted at any instantiated width; the hard error moved out to the uninstantiated set. Already wrong before this PR: - L141 "TurboQuant is not the default anywhere" `defaults::PAGED_CACHE_TYPE` is `PagedCacheType::TurboQuant` and the CLI's `--pa-cache-type` has no clap default, so leaving it unset on CUDA gives a standard-layout head_dim-128 model TurboQuant KV with no flag — and silently drops prefix caching, which no TurboQuant variant supports. That was true on master; the narrowing in a3a5faa bounds it to 128 but does not remove it. Stating the opposite understated the risk to anyone reading the README to decide whether to opt out, so it is corrected in place rather than quietly reworded, and the Rust example's comment now says `Auto` opts *out* rather than implying TurboQuant is off until asked for. Also refreshes the `--pa-cache-type` help text, which still described the explicit path as 128-or-error. No behaviour change: documentation and one clap doc comment. * fix(build): stop emitting a duplicate -gencode for the build box's own arch `mistralrs-paged-attn/build.rs` appended one `-gencode` per entry in `ARC_CUDA_ARCHS` on top of the one cudaforge derives itself, and applied the same `a` suffix for cap >= 90 that cudaforge's `GpuArch::auto_suffix` does. Verified against cudaforge 0.1.5 rather than assumed: compute_cap.rs auto_suffix(90).to_gencode_arg() -> "-gencode=arch=compute_90a,code=sm_90a" builder.rs:401 let gencode_arg = gpu_arch.to_gencode_arg(); builder.rs:405 command.arg(&gencode_arg)... builder.rs:412 for arg in &self.extra_args { command.arg(arg); } cudaforge emits its own gencode first and then appends `extra_args` verbatim, so on an H200 building `ARC_CUDA_ARCHS=90,100,103` nvcc received `-gencode=arch=compute_90a,code=sm_90a` twice, byte-identical. Fixed by hoisting the existing `get_compute_cap()` binding above the loop (it is what decides which requested arches are genuinely extra) and skipping the one cudaforge already covers. The later duplicate binding is deleted; there is now one. This is the dedupe half of #108's `arc_target::build::split_primary()`, done locally because `arc-target` does not exist on this branch. The other half — `verify_and_export()`, the `cuobjdump` gate that fails a build whose archive is missing a requested arch — genuinely needs that crate and is left to #108, which should replace this block with the two calls. Until it lands, `arc-tools/wave64_v4_turboquant_kv_gate.sh` STEP 2 asserts the arch list externally with `cuobjdump --list-elf`, so a missing arch is caught by the gate even though it is not yet caught by the build. Not verified here: whether nvcc rejects or tolerates the duplicate flag. macOS cannot run the cuda build path. Emitting it once is correct either way, which is why this is worth doing without waiting on that answer.
Stacked on #94 (
feat/turboquant-hd512-v4). Review #94 first.#94 finished the kernel side and deliberately did not ship the integration. This is that integration — and the head_dim was never what stood in the way.
🔑 V4 and TurboQuant are the same two-region shape
dsv4_attentionreads the raw key cache over exactly one span: the trailingt_q + window - 1tokens (raw_keep_span). Every earlier raw column is-infon every query row, because V4's raw branch is a sliding window, not dense causal attention. Distant context never arrives through the raw cache at all — it comes fromcompressed_kv, which the compressor builds from its own rolling history (XsRollingCache) and which never touches those bytes.fp16_window— recent tokens kept uncompressedSame boundary, so this builds them as one mechanism. The payoff is the thing that makes it worth shipping:
That is the opposite of
arc_turbo::TurboQuantSingleCache, whosecurrent_datareconstructs every compressed token on the host on every call — the documented reason that path is opt-in. A dequantizing path exists here too (spanmust stay total, and a deep rollback can ask for one), but the model never takes it, and a test pins that.What landed
KvCache::V4Turbo/V4TurboKCache, modelled directly onXsRollingCache: a grown capacity buffer of packed records plus a bounded dense tail, reporting length in tokens and mapping truncation onto both time bases. Norms ride inside the code record rather than in a third buffer, becauseBatchSrcgives a cache slot exactly two tensors and a third would not survive clone-in/clone-out.V4CachedK::Turbois wired into the eager decode arm behindARC_V4_TURBOQUANT=1.Bytes per token per layer at V4-Flash's geometry (
head_dim=512, MQA, BF16):ARC_V4_FP8_KV=1)The dense tail stays dense and is bounded by
window + max(t_q-1, margin) + EVICT_CHUNK— independent of context length.Eviction is chunked, on purpose
TurboQuantLayout::quantize_intois a host function over&[f32]. Compressing every decode step would be one device→host→device round trip per layer per token — ~86 syncs a token across V4's 43 layers, which is CLAUDE.md pitfall #5 and would cost exactly the decode speed this is supposed to be free of.EVICT_CHUNK = 128amortises it to ~0.34 syncs/token.Two refusals rather than fallbacks (D18)
ARC_V4_TURBOQUANT=1+ARC_V4_STANDARD_DENSE=1is a hard error. That ablation makes ratio-0 layers attend densely over the whole cache instead of through the window, so the reachability argument stops holding and the compressed region would be silently missing from those layers' softmax.append_kv_turborefuses a slot it did not build, naming both variants. Silently falling back to the dense path would make "flag on" and "flag doing nothing" identical.For the same reason the codec emits one
Once-guarded line on its first eviction. Without it, flag-on and flag-on-but-never-compressing are indistinguishable from outside the process — and that distinction is the only thing the GPU gate can actually check.Also: two silently-ignored kernel parameters
Pre-existing, and the same failure class as the six
if (hs != 128) return;no-ops #94 removed — accepted, plumbed through the FFI, never read.qs(query stride) is now honoured:query[sidx*qs + hidx*HS + i]. For a contiguous queryqs == nh*HS, so this is the same address every current caller already computed. The kernel still assumes the head and dim axes are packed, so the Rust wrapper now checks that and bails naming the actual strides.softcappingis applied with the same arithmetic and at the same position aspagedattention.cuh:276— after thescalemultiply, before both the logit store and the running max (capping only one would leave the softmax renormalising against a maximum no stored logit can reach).softcapping == 1.0stays free via a grid-uniform hoisted branch.📊 Measured — CPU unit tests (D14), not hardware
cargo test -p mistralrs-core --lib→ 386 passed, 0 failed (was 377) ·--test synthetic_load_smoke13 ·--test kv_sharing_workloads7 ·-p mistralrs-quant turboquant32 · workspacecargo checkgreen · scoped clippy lane exit 0 · rustfmt drift like-for-like (v4_turbo 0, kv_cache/mod 0→0, deepseek4 30→30).The load-bearing test is
turbo_and_dense_cached_spans_agree_bit_exactly: across a real prefill + 400 decode steps with periodic rollbacks, the span the TurboQuant cache handsdsv4_attentionis bit-identical to the same range of the dense cache. Bit-identical, not close — a tolerance there would mean the codec had reached tokens the model can still see. Combined with #94's existingraw_prefix_is_equivalent_to_passing_the_whole_cache, that means the attention output is unchanged.Mutation runs — and the three that survived
span ignores the offsetsurvives it legitimately — the model only ever asks foroff == 0, so the mutation is unreachable from that path. Caught by the directspantests.EVICT_CHUNK(128) is ≫margin(16), so chunked eviction holdsbaseso far below the retention floor that no end-to-end fixture can make the margin term binding. The arithmetic is defence in depth for a smaller chunk or an eager eviction. So it was extracted as a pureretention_floor()and tested directly over the regime where it does bind — where all three of its terms now fail when mutated.Recording it because a mutation table nobody failed is a table nobody ran.
🔴 The honest gap
Zero GPU exposure. Neither the CUDA edit nor the cache has been on a box — the
qs/softcappingfix has never even been compiled, since nvcc does not exist on macOS and thecudamodule is cfg-gated socargo checknever parsed its Rust side either.arc-tools/wave64_v4_turboquant_kv_gate.shis written and self-contained; STEP 1 is that first compile, STEP 6 is the engagement check, and STEPS 4/5/7 are decode tok/s and peak KV bytes with the flag off vs on.supports_paged_attentionstill returnsOk(false)for V4 and is untouched: V4 is eager-NormalCacheonly, so the paged TurboQuant kernel remains unreachable for it at any head dim. That is a separate change, and flipping it blind is how wave43-BU went.Default stays off until the gate runs. Default-on is the destination; an unmeasured default-on is how we lose a day.
🤖 Generated with Claude Code