Skip to content

feat(turboquant): CUDA kernels at head_dim 64/128/256/512, Hopper + Blackwell - #94

Merged
heydryft merged 8 commits into
release/openrouter-readyfrom
feat/turboquant-hd512-v4
Aug 21, 2026
Merged

heydryft merged 8 commits into
release/openrouter-readyfrom
feat/turboquant-hd512-v4

Conversation

@heydryft

Copy link
Copy Markdown
Contributor

What was actually blocking TurboQuant

static_assert(HEAD_SIZE == 128) in the .cuh is dead code — that header's kernel has no callers. The live kernel is turbo_paged_attention.cu, and its limit was worse than an assert: six extern "C" launchers each opened with if (hs != 128) return;. That is a silent no-op. The output buffer is left uninitialized, so a mismatched head dim produces garbage quickly — a broken run reports a better number, not a worse one.

Kernels

tq_attn, tq_attn_clocked, tq_cache_k, tq_cache_v and both clocked cache variants now take HS as a template parameter, instantiated at 64, 128, 256 and 512.

  • NT stays 128 at every HS, so the warp-reduction topology (4 warps, red[]) is unchanged. Wider heads give each thread DPT = HS/NT output dims rather than adding threads. At HS=128 the emitted work is what it was.
  • The per-lane packed-K gather is now one vector load (tq_load_klane) instead of BPL scalar byte loads. The x=16 cache swizzle keeps each lane's run contiguous and naturally aligned; packed_k_bytes_are_swizzle_aligned asserts the property the load depends on rather than trusting it.
  • Every if (hs != 128) return; is gone, replaced by a dispatch over the instantiated widths.

D16 — Hopper and Blackwell

The file had zero __CUDA_ARCH__ guards, and nothing in the workspace passed an explicit -gencode: cudaforge derives one arch from CUDA_COMPUTE_CAP or the local nvidia-smi, so a binary built on an H200 carries SM90 only.

  • TQ_PREFETCH software-pipelines the K gather on SM90/SM100/SM103 and runs flat below Hopper. Same arithmetic, scheduled for the part it runs on.
  • Launches now opt in to the arch's real dynamic shared-memory budget (164 KB Ampere, 227 KB Hopper/Blackwell). This fixes a live bug: the logits array scales with context length, and without an explicit opt-in CUDA caps dynamic shared memory at 48 KB, so any context past ~12k tokens could not launch on any arch.
  • ARC_CUDA_ARCHS=90,100,103 emits cubins for both architectures. Left unset the build is unchanged, because each extra arch recompiles every .cu and that is not worth paying on a rented box that runs one arch.

Tables — generated, not transcribed

Codebooks and sign diagonals are both functions of the dimension, so 512 needed its own; the centroids shrink as sqrt(1/d). All 20 tables come from emit_turboquant_cuda_tables, which calls the same generate_signs/get_codebook the Rust compressor uses. The generator reproduces the shipped d=128 sign table byte-for-byte, which is the check that it is the same format.

turboquant::cuda_tables re-derives every table at test time and diffs the checked-in header, and pins the Rust head-dim list against both the kernel's dispatch switch and the FFI wrapper's list. Previously nothing asserted that the tree's four separate copies of the sign table agreed with the generator or with each other, despite a comment in turbo_wht.cu claiming they must.

Gates

TURBOQUANT_HEAD_DIM = 128 exact-match becomes membership in TURBOQUANT_CUDA_HEAD_DIMS. K and V are compressed independently, so they no longer have to be equal widths either.

What this does not do — the honest gap

This does not put DeepSeek-V4 on TurboQuant, and the kernel was never what stood in the way.

DeepSeekV4Loader::supports_paged_attention returns Ok(false) (normal_loaders.rs:3231). V4 uses the eager NormalCache exclusively, so the paged TurboQuant kernel is unreachable for it at any head dim. The two things that actually block it are unrelated to head size: cache_write_and_gather returns a varlen pack [1, H, sum(seqlen), D] that dsv4_attention would read as one sequence, and the xs history slots are per-sequence only through NormalCacheManager::clone_in_cache, which the paged engine arm never issues.

Separately, V4 stores no V at all — V is K bit-for-bit under the fused wkv, and the V half of the slot is a 1-wide zero marker (V4_V_MARKER_WIDTH = 1). So the K+V shape TurboQuant compresses does not map onto V4's KV as-is; the integration point is V4CachedK, alongside Dense and the FP8 Packed, not the paged kernel.

The eager path is also not a shortcut: TurboQuantSingleCache::current_data reconstructs the entire compressed context on the host every decode step, and require_normal_kv_slot rejects a TurboQuant slot for V4 outright.

Making V4 default-on means teaching V4CachedK a TurboQuant variant. That is a KV storage change, and this repo already has the scar from shipping one unmeasured — wave43-BU turned FP8 KV on by default without a GPU run and every V4 request died on a dtype mismatch, which is why ARC_V4_FP8_KV is opt-in today. It should be its own change with a GPU run attached.

Verification

  • cargo check green; cargo test -p mistralrs-core --lib 375 passed; cargo test -p mistralrs-quant turboquant:: 32 passed.
  • Scoped clippy lane green.
  • No GPU ran. There is no nvcc on the authoring box, so nothing here is a claim about hardware behaviour — the CUDA is written and gated but unproven. arc-tools/wave63_turboquant_hd512_gate.sh is the D15 script that compiles all four instantiations for sm_90a and sm_100a/sm_103a and checks the shipping decode path for regression. It contains no quality A/B by design.

@github-actions

github-actions Bot commented Aug 17, 2026 •

Copy link
Copy Markdown
Code Metrics Report
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
 Language              Files        Lines         Code     Comments       Blanks
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
 C Header                  5          305          210           52           43
 CSS                       2         1181         1036           34          111
 CUDA                     81        29475        20161         6275         3039
 Dockerfile                1           39           22            8            9
 JavaScript               16         3546         2676          482          388
 Jinja2                    7          694          656            5           33
 JSON                     75         4896         4893            0            3
 Makefile                  1            6            5            0            1
 Metal Shading Lan|       33        12224         9431         1142         1651
 PowerShell                1          300          227           30           43
 Python                  147        15371        12677          824         1870
 Shell                    42        10304         6858         2769          677
 Plain Text                4         3801            0         2479         1322
 TOML                     33         1498         1294           54          150
 YAML                      3           25           23            2            0
─────────────────────────────────────────────────────────────────────────────────
 HTML                      4         2687         2604           43           40
 |- CSS                    2          543          479           37           27
 |- JavaScript             1         1233         1215           12            6
 (Total)                             4463         4298           92           73
─────────────────────────────────────────────────────────────────────────────────
 Jupyter Notebooks         4          122           83           23           16
 |- Markdown               1           60           30           22            8
 |- Python                 1          122          113            1            8
 (Total)                              304          226           46           32
─────────────────────────────────────────────────────────────────────────────────
 Markdown                213        48932            0        38081        10851
 |- BASH                  72         1655         1203          331          121
 |- C                      4           19           19            0            0
 |- CUDA                   2           84           56           16           12
 |- JSON                  19          779          779            0            0
 |- PowerShell             1            1            1            0            0
 |- Python                23         1008          787          113          108
 |- Rust                  68         2063         1727           78          258
 |- TOML                   6          207          164            0           43
 |- YAML                   5           41           36            5            0
 (Total)                            54789         4772        38624        11393
─────────────────────────────────────────────────────────────────────────────────
 Rust                    681       340988       293252        18096        29640
 |- Markdown             504        32562          471        28063         4028
 (Total)                           373550       293723        46159        33668
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
 Total                  1353       516771       363188        99077        54506
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

@heydryft

Copy link
Copy Markdown
Contributor Author

Post-review fixes (5bfebd6)

An independent CUDA review re-derived the indexing in Python and type-checked the kernel bodies against a shim. It found no compile errors and confirmed the head_dim generalization is exact — the vectorized gather reproduces the original scalar koff formula byte-for-byte for every byte every lane owns at all four widths, coverage is exactly once with no overlap, every vector-load base is naturally aligned, rotation is involutory to 5.5e-16, and no __syncthreads() is divergent at HS < NT or HS > NT.

It found three real problems, now fixed:

  1. 32-bit overflow on cache byte offsets. pb*kbs, pb*vbs, bi*cbs, bi*vbs were int. At head_dim 512 with nkvh=8/block_size=32, kbs is 65536 B, so pb*kbs wraps once the K cache passes 2 GB — reachable on an H200, and 4× sooner than at 128. Now long long at all eight sites. This one predates this PR but the wider head makes it 4× closer.
  2. Shared-memory opt-in cached a process-wide flag. The requestable maximum differs by architecture, so one granted flag cannot be right across a mixed fleet; worse, a failed request left it at 0 and fell through to a launch that failed with no diagnostic. Cache removed.
  3. tq_attn_clocked kept the scalar gather while tq_attn moved to a vector load. Identical results, but the phase-2 stamps would describe a kernel nobody launches — at 512 that is 8 scalar loads measured against 1 vector load executed. The loop now mirrors tq_attn.

Still-open items from that review, deliberately not fixed here

  • qs (query stride) is accepted, plumbed through the FFI, and ignored — both kernels hardcode query[sidx*nh*HS + hidx*HS + i], correct only while qs == nh*hs. A non-contiguous query (fused QKV slice) would read garbage silently. Pre-existing; deserves its own change that either uses it or deletes it from the ABI.
  • softcapping is likewise accepted and never applied. Any model needing logit softcapping gets silently wrong attention through this path. Pre-existing.
  • Instantiation count is now 4 HS × 3 block sizes × 4 (QT,OutT) = 48 attention kernels plus 24 clocked and 32 cache, roughly 4× the previous cubin and compile time. Widths 64 and 256 may have no consumer yet — worth a call on whether to trim the set.

Verification without nvcc

No GPU and no nvcc on the authoring box, so the CUDA remains compiled-in-theory only. What was actually run locally: the C preprocessor over both __CUDA_ARCH__ branches (confirms all four nested variadic macros expand, and that each of the four widths gets an identical instantiation count — 15 attn, 2 cache, 6 clocked — so no width is missing a dispatch path), and clang++ -std=c++17 -fsyntax-only -Wall -Wextra -Wshadow over the kernel bodies against a CUDA shim with <<<>>> stripped, clean on both branches.

arc-tools/wave63_turboquant_hd512_gate.sh remains the thing that actually proves it.

@heydryft

Copy link
Copy Markdown
Contributor Author

Gate corrected (05282f0) — D18 exit codes + the sourced-preflight landmine

The gate exiting 1 on a missing gpu_box_preflight.sh was my bug, not the runner's. A missing harness file is environment cannot answer. Fixed per D18 rule 2.

Exit codes. Nine conditions now exit 2 — missing / failing / incomplete preflight, no repo, no nvcc, unreachable remote, un-checkout-able branch, no archive to inspect, no cuobjdump. Exactly four exit 1, all genuine: an nvcc compile failure, a native build failure, a test failure, and the archive missing an arch it claims to carry. A run that ends 2 says nothing about PR #94's code.

Preflight lookup is now branch-independent: ARC_PREFLIGHT → /root/arc-tools/ → /usr/local/lib/arc/ → repo-relative last. The repo-relative path is the one that vanishes the moment a runner checks out another branch, which is exactly how this gate came to refuse to run.

The source landmine is guarded directly. I confirmed it in the source — arc-tools/gpu_box_preflight.sh:276-281:

if [ "$_ARC_PF_FAILED" = "1" ]; then
    [ "$_arc_pf_sourced" = "1" ] || exit 1

Sourced on a failing box, [ "1" = "1" ] succeeds, short-circuits the || exit 1, and becomes the last command — so source pf || handler returns 0 and the handler never fires. Reproduced it against a stub, then fixed it the way the sibling gate did: ignore the return value entirely, read _ARC_PF_FAILED, and treat an unset flag as "did not reach a verdict" rather than as a pass. (This gate execs rather than sources, so it was never hit by it directly — but a gate that only works because of how it happens to invoke the thing is one refactor from being wrong.)

Reordered so the D16 proof is STEP 1 and needs no full workspace build — PR #98 already showed these kernels compile clean for sm_90 in 8m53s, so the native build is confirmation, not discovery, and is skippable. STEP 1 also now asserts on cuobjdump --list-elf that sm_90a and sm_100a are actually in libmistralrspagedattention.a. A build that merely succeeds is not the claim; the cubins being present is.

Verified locally against a stub reproducing the real preflight's sourced-exit structure — all four environment cases return 2, syntax clean.

@heydryft

Copy link
Copy Markdown
Contributor Author

Blocking on a silent default change — the kernels are welcome, the default is not.

PagedCacheType::TurboQuant carries #[default] (paged_attention/cache_engine.rs:22). This PR replaces the acceptance gate

config.k_head_dim() == TURBOQUANT_HEAD_DIM   // 128

with

mistralrs_quant::turboquant::cuda_supports_head_dim(config.k_head_dim())   // {64,128,256,512}

So any paged-attention model at head_dim 64 or 256 that today falls back to Auto will, after this merge, run TurboQuant KV by default — through kernels that have never executed on hardware. It also loses prefix caching silently, since supports_prefix_cache() is false for every TurboQuant variant.

The tell is in this PR's own fixture: the "expected to fall back" list changes from [64, 96, 192, 256] to [96, 192, 320, 1024]. 64 and 256 were removed because they stopped falling back — that is the default change, recorded in a test edit rather than in the description.

This is the wave43-BU shape exactly, which #98 cites by name as the reason it refuses to default itself on: FP8 KV shipped default-on and unmeasured, and every V4 request died. Every other PR in this batch is scrupulously opt-in; this is the one that isn't.

Suggested fix, which keeps everything valuable here: leave default resolution at head_dim 128 and let the wider set require an explicit --pa-cache-type turboquant. The six removed if (hs != 128) return; no-ops (D18 instance 3 — a no-op that left the output buffer uninitialised) and the 48 KB dynamic-smem cap fix are real wins and are not what I'm objecting to.

Ordering dependency: #97 fixes an inverted SdpaParams::sliding_window on V4 that is inert today only because flash_sinks_ok rejects head_dim 512. This PR arms it. If #94 lands without #97, the 512-wide kernels activate the inverted window silently — the fused path takes no mask that could disagree. #97 must land first or with it.

Hardware gate is queued on the H200 (arc-tools/wave63_turboquant_hd512_gate.sh); the one number that matters is whether the built archive contains both sm_90a and sm_100a cubins, which is exactly where the last run genuinely failed.

@heydryft

Copy link
Copy Markdown
Contributor Author

Narrowed the default — pushed as a3a5faab. The blocker above is resolved; the kernels are untouched.

resolve_for_model already distinguished explicit from ambient-default in order to choose error-vs-fallback, so the fix was to let that same forced flag also choose which head-dim set applies:

set meaning
compiled TURBOQUANT_CUDA_HEAD_DIMS = {64,128,256,512} what you may ask for
measured TURBOQUANT_DEFAULT_HEAD_DIMS = {128} what you get without asking

TURBOQUANT_DEFAULT_HEAD_DIMS moves when a hardware gate passes, not when a kernel compiles.

Nothing about the kernel work changed. The six if (hs != 128) return; no-ops (D18 instance 3 — a no-op leaving the output buffer uninitialised) and the 48 KB dynamic-smem cap fix are exactly as you wrote them, and an explicit --pa-cache-type turboquant still reaches every instantiated width, including V4's 512.

The unsupported-geometry message now separates two facts that were being collapsed: "no kernel exists" versus "a kernel exists but is unmeasured, so the default will not choose it for you." The second tells a user to opt in; the first tells them to stop.

On the tests. The old fixture had the regression written into it — the expected-to-fall-back list moved [64, 96, 192, 256] → [96, 192, 320, 1024] precisely because 64 and 256 had stopped falling back. That's replaced by an explicit-path test across every instantiated width, plus two pins on the default path.

Both pins were mutation-checked by reverting the const to the full set. Worth recording that the first revision of one of them passed under that mutation — it skipped widths that were in the default set, so widening the set emptied its loop and it asserted nothing. That is the same silent-success shape the PR is about, reproduced inside its own test. It now asserts the loop is non-empty before iterating, and both pins fail under mutation and pass restored.

Local: 377/377 mistralrs-core lib tests, scoped clippy lane clean, both touched files rustfmt-clean. The hardware gate is still queued — the sm_90a/sm_100a cubin inventory remains the open question, and it is unaffected by this change.

heydryft added a commit that referenced this pull request Aug 17, 2026
…V4 KV branch

One conflict, in the V4 attention KV-append arm: master added an
`arc_profiler::device_span("kv_cache_append")` exactly where this branch
added the TurboQuant storage branch.

Resolved by keeping both and giving the TurboQuant arm its own span,
`kv_cache_append_turbo`. Eviction is the single cost this path adds to
decode, and the ON/OFF decode comparison this branch has to clear is a
throughput claim — folding eviction into the dense span's number would
have hidden the one line item the measurement is meant to expose.
@heydryft

Copy link
Copy Markdown
Contributor Author

Merged master in and corrected the public record — 50c173bf. This PR branched before #101 landed, so it had never seen that text. Four claims in it are false; three are falsified by this PR's own kernels, and one was already wrong when written.

README claim status
L31 "the paged kernel exists at head_dim 128 only" falsified by this PR
L31 "no kernel at head_dim 512, so DeepSeek V4 cannot use it" falsified by this PR
L141 explicit --pa-cache-type turboquant off-128 "is a hard error" falsified by this PR
L141 "TurboQuant is not the default anywhere" already false on master

The first three are straightforward: V4 reports KvCacheLayout::Standard at head_dim 512 (normal_loaders.rs:3620), so the 512 instantiation genuinely reaches it, and the hard error has moved out to the uninstantiated set.

The fourth is the one worth reading twice, because it is not this PR's doing:

pub const PAGED_CACHE_TYPE: PagedCacheType = PagedCacheType::TurboQuant;  // :113

--pa-cache-type carries no clap default, so unset means None → defaults::PAGED_CACHE_TYPE → TurboQuant, with cache_type_explicit = false. On CUDA, with PagedAttention auto-enabled, a standard-layout head_dim-128 model gets TurboQuant KV with no flag — and silently loses prefix caching, which no TurboQuant variant supports. default_turboquant_still_taken_on_measured_head_dims asserts exactly this and passes.

So the README told readers TurboQuant never runs unless asked, while the shipped default ran it on the most common geometry there is. a3a5faab bounds that to 128; it does not remove it. I corrected the sentence in place rather than quietly rewording it, since understating the risk is what made it worth catching. The Rust example's PagedCacheType::Auto comment now reads as opting out, and the --pa-cache-type help text no longer describes the explicit path as 128-or-error.

Nothing here changes behaviour — docs plus one clap doc comment. cargo check green, cache_type_tests 11/11.

@heydryft

Copy link
Copy Markdown
Contributor Author

Took the moat agent's build.rs finding — 2b1253bc. Verified on this branch rather than taken on report, and the mechanism is worth stating precisely, because the duplicate is byte-identical rather than merely redundant.

cudaforge 0.1.5 emits its own -gencode per compile job and then appends extra_args verbatim:

compute_cap.rs   auto_suffix(90).to_gencode_arg() -> "-gencode=arch=compute_90a,code=sm_90a"
builder.rs:401   let gencode_arg = gpu_arch.to_gencode_arg();
builder.rs:405   command.arg(&gencode_arg)...
builder.rs:412   for arg in &self.extra_args { command.arg(arg); }

Our loop applied the same a-suffix rule for cap >= 90, so an H200 building ARC_CUDA_ARCHS=90,100,103 handed nvcc -gencode=arch=compute_90a,code=sm_90a twice.

Fixed by hoisting the get_compute_cap() binding above the loop — it is what decides which requested arches are genuinely extra — and skipping the one cudaforge already covers. The later duplicate binding is deleted, so there is now exactly one.

Scope, deliberately partial. This is the dedupe half of #108's arc_target::build::split_primary(), done locally because arc-target does not exist on this branch — nor do the cited mistralrs-quant/build.rs:100/120/205 call sites, which are #108's line numbers. The other half, verify_and_export(), genuinely needs that crate and is left to #108; it should replace this block with the two calls. I did not touch mistralrs-quant/build.rs.

The silent-arch hole is covered in the meantime, just not by the build: arc-tools/wave64_v4_turboquant_kv_gate.sh STEP 2 already asserts the archive's arch list with cuobjdump --list-elf and fails the gate on a missing one. So the failure that produced today's genuine gate result is caught by the gate today and by the build once #108 lands.

One thing I could not settle: whether nvcc rejects or tolerates the duplicate flag. macOS cannot run the cuda build path, so I am not claiming it was fatal. Emitting it once is correct under either answer, which is why it was worth doing without waiting for that.

@heydryft

Copy link
Copy Markdown
Contributor Author

Merge sequencing: #111 lands first, then this — and one sentence here must go

#111 (docs/turboquant-measured-record) overlaps this PR on five files: README.md, mistralrs-cli/src/args/paged_attn.rs, paged_attention/cache_engine.rs, paged_attention/mod.rs, turboquant/mod.rs. Confirmed resolution: this PR's head_dim structure, #111's provenance sentences. Different halves of the same false record; neither is complete alone.

#111 goes first because this PR is blocked on its H200 gate anyway (the archive must contain sm_100a cubins, not just sm_90a — precisely where the last run genuinely failed), and resolving against a master that already carries the provenance produces the right composite without anyone re-inserting text.

The blocking item

README.md:143 currently reads:

None of this has a measured serving run behind it at any width, 128 included.

That is false, and it needs to come out in this PR's conflict resolution — not be left to merge order to overwrite.

There is one measured serving run, and I verified it against the object store rather than taking it secondhand:

Commit 4eba13905, 2026-04-06 — "fix: revert to proven working eager state (55 tok/s, correct output)"
Model Qwen/Qwen3-32B — deploy/modal_b200.py line 5, docstring "Arc inference — Qwen3-32B with CUDA graphs on B200"
Scope b=1, one card, head_dim 128, Blackwell

One correction to the provenance as it was handed to me: deploy/modal_b200.py was not in the tree at 4eba13905 — there is no deploy/ directory at that commit. It was added by 404ee1aad, 2026-04-07 ("deploy: ncu/nsys/bench wiring + B200 v73"), the day after. That's an ordinary order of events and doesn't weaken the finding, but the citation should read "harness committed the next day at 404ee1aad, naming Qwen3-32B" rather than implying the file was present at the measured commit. This whole episode started with an audit that quoted a commit, called the model "unstated", and concluded "never measured" — so the correction's own citation should be unchippable.

What stays as-is

  • +46% remains unattributable. It compares Arc's whole dedicated decode path against Candle's and isolates nothing; there is no A/B against an unquantised cache.
  • 4.27× stays retracted. docs: name every Arc system, and correct the public record #101 was right.
  • head_dim 512 / V4 numbers stay design claims. No kernel has executed there — which is the same fact the narrowing in this PR now enforces in code.

Net effect on the composite: default-on and bounded to 128 (this PR's structure) plus one real measured run, correctly scoped (#111's provenance). The sentence above is the only thing here pulling against that.

The narrowing commit and its mutation-checked tests are unaffected by any of this; the hardware gate remains the only other open item.

@heydryft

Copy link
Copy Markdown
Contributor Author

Cleared: the CI complete: FAILURE above is a CANCELLATION ARTIFACT of mine, not a verdict on this PR. Fresh lanes are running now — read those, not the stale aggregate.

Sequence of what happened here, because the stale status is exactly the thing that gets quoted an hour later by someone who wasn't in the thread:

  1. Master's test target broke (fix(ArcSched): stop bucketing paged decode by sequence length (8.7x on spread traffic) #118 + perf(engine): kill the SequenceGroup spin-lock and the await-under-lock response dispatch — and the measurement that refutes their premise [MEASURED] #115, a semantic conflict between two textually disjoint PRs that each had a genuine 17/17 — GitHub tests each PR against the base as it stood, so neither run could see the other). Fixed by fix(build): master's test target is red — #118's fixture builds a tokio Mutex where #115 now requires std's #131, now merged.
  2. I absorbed master into this branch while master was red, so this PR inherited that failure. Proven causally rather than inferred: on this branch's then-head, cargo check -p mistralrs-core --tests exited 101; cherry-picking fix(build): master's test target is red — #118's fixture builds a tokio Mutex where #115 now requires std's #131's one-word commit on the same tree took it to exit 0, 0 errors.
  3. To free runners for fix(build): master's test target is red — #118's fixture builds a tokio Mutex where #115 now requires std's #131 — which was queued behind ~20 runs re-deriving the break it fixes — I cancelled this PR's in-flight runs. That was my call and it had a consequence I did not anticipate: ci-complete uses if: always() over needs: every lane, so a CANCELLED dependency makes the aggregate FAILURE. The red you see is that, not a defect.

🔑 The lesson I'm recording on the PR rather than just in my own notes: cancelling CI does not merely discard information — it WRITES a false negative into the PR's permanent check record. I had told the coordinator "nothing informative was destroyed", which was true about the runs and false about the record.

Now: master is green (cca5e5c6e, verified unpiped — cargo check --workspace --tests exit 0, cargo test -p mistralrs-core --lib 537 passed / 0 failed), this branch has absorbed it, and full lanes are re-running.

Judge this PR on the new run. If you are reading the old aggregate, it is describing my cancellation, not this code.

@heydryft

Copy link
Copy Markdown
Contributor Author

Triage verdict: REBASE, not close. Still carries unique work.

Evidence against current master: mistralrs-paged-attn/src/cuda/turbo_paged_attention.cu still has 4x if (hs != 128) return; — head_dim 64/256/512 are still silently unreachable — and mistralrs-quant/src/turboquant/cuda_tables.rs does not exist on master. None of this landed by another route.

Conflicts are small: 2 files — mistralrs-cli/src/args/paged_attn.rs and mistralrs-core/src/paged_attention/cache_engine.rs.

Two things the rebase must handle explicitly, given what has been learned since this PR was opened:

  1. A return on an unsupported head_dim is the silent-success failure mode (D18). Replace with a loud refusal, not a wider if.
  2. Do not let the rebase flip a default. TurboQuant is already the default cache type and that silently disables prefix caching (open defect). New head_dim support lands reachable-but-not-default.

Leaving open with the decision recorded. Not merging it blind.

@heydryft

Copy link
Copy Markdown
Contributor Author

Retargeted at the integration branch

Base changed: master → release/openrouter-ready.

The queue is being restructured to the shape the owner asked for: one PR open against master (#194), with everything else merging into a single integration branch. Agents branch off release/openrouter-ready and PR back into it; #194 is the single gate from there to master.

Two things had to land on master first for this to be workable, and both have:

  1. The base-branch CI lane used to hard-fail any PR not targeting master (ci(ArcGate): let the base-branch lane permit the single integration branch — 7 PRs are red for addressing, not quality #216). Seven PRs were red for their addressing, not their quality. The lane now accepts master or release/openrouter-ready — an exact-match allowlist, nothing else — and its intent is intact: release/openrouter-ready reaches master only through Arc → OpenRouter-ready: the single integration PR (everything else merges into this branch) #194, whose own base is master, so nothing enters master without a full run whose base is master.
  2. release/openrouter-ready was re-cut from current master. It had diverged badly — merging it as-is would have reverted TCFRAG (perf(ArcQuant/ArcKernels): TCFRAG-2B — put the qtip2b trellis GEMV on the tensor cores #203) and re-broken the flag-polarity work (fix(ArcGate): read boolean ARC_* flags by value, so =0 means off #212). It is now master plus only the work it uniquely carried.

This PR was not closed and is not considered stale. An audit of the queue found the overwhelming majority of it to be real work that was never merged, not noise.

What you need to do: rebase onto release/openrouter-ready and resolve conflicts against it rather than against master. CI lanes will now actually run and report on this PR instead of failing at the base check.


📚 Stack order — this is the BOTTOM

This PR and its dependent were both retargeted at release/openrouter-ready, so they are now siblings rather than a stack. That makes the merge order load-bearing and no longer enforced by the base ref:

Merge this PR FIRST. Its dependent carries changes that assume this one is already in the branch, and merging them out of order will produce conflicts or silently drop this PR's changes.

…/512

TurboQuant's paged kernels were hardcoded to head_dim 128 — `constexpr int
HS = 128`, a single `SGN[128]` sign table, `rotate128`, and six `extern "C"`
launchers each opening with `if (hs != 128) return;`. That early return is a
silent no-op: the output buffer is left uninitialized, so a mismatched head
dim produces garbage *fast* rather than failing.

Kernels
- `tq_attn`, `tq_attn_clocked`, `tq_cache_k`, `tq_cache_v` and the two clocked
  cache variants now take HS as a template parameter. NT stays 128 for every
  HS, so the warp-reduction topology is unchanged; wider heads give each
  thread DPT = HS/NT output dims instead of adding threads. At HS=128 the
  emitted work is the same as before.
- The per-lane packed-K gather is now one vector load (`tq_load_klane`)
  instead of BPL scalar byte loads. The x=16 cache swizzle keeps each lane's
  run contiguous and naturally aligned, which `packed_k_bytes_are_swizzle_aligned`
  asserts.
- Every `if (hs != 128) return;` is replaced by a dispatch over the
  instantiated widths.

D16 dual-arch
- `TQ_PREFETCH` software-pipelines the K gather on SM90/SM100/SM103 and runs
  flat below Hopper — same arithmetic, scheduled for the part it runs on.
- Launches now opt in to the arch's real dynamic shared-memory budget. The
  logits array scales with context length and CUDA caps dynamic shared memory
  at 48 KB without an explicit opt-in, so contexts past ~12k tokens could not
  launch at all on any arch.
- `ARC_CUDA_ARCHS=90,100,103` emits cubins for Hopper *and* Blackwell rather
  than only the build box's GPU.

Tables
- Codebooks and sign diagonals are dimension-dependent, so 512 needs its own.
  All 20 tables are emitted by `emit_turboquant_cuda_tables`, which calls the
  same `generate_signs`/`get_codebook` the Rust compressor uses — correct by
  construction, not by transcription. The generator reproduces the shipped
  d=128 table byte-for-byte.
- `turboquant::cuda_tables` re-derives every table at test time and diffs the
  checked-in header, and pins the Rust head-dim list against both the kernel's
  dispatch switch and the FFI wrapper's list. Nothing previously asserted that
  the tree's four copies of the sign table agreed with the generator or with
  each other.

Gates
- `TURBOQUANT_HEAD_DIM = 128` exact-match becomes membership in
  `TURBOQUANT_CUDA_HEAD_DIMS`; K and V are compressed independently and no
  longer have to be equal widths.

Not measured on a GPU: the kernels are written and gated but this box has no
nvcc, so nothing here is a claim about hardware behaviour.
… kernels

Proves the four head-dim instantiations compile for sm_90a AND sm_100a/sm_103a
(D16), not just the build box's GPU, and that the shipping decode path did not
regress. Deliberately contains no quality A/B: TurboQuant quality is settled by
prior measurement and re-proving it would spend GPU budget on a known number.
Every other .cu in this crate had a rerun-if-changed line; this one did not.
Emitting any rerun-if-changed disables cargo's watch-the-whole-package default,
so edits to the largest CUDA file in the crate were relying entirely on
cudaforge's object cache to trigger a rebuild.
…d-kernel parity

Three problems from review of the head_dim generalization.

1. 32-bit overflow on cache byte offsets. `pb*kbs`, `pb*vbs`, `bi*cbs` and
   `bi*vbs` were all `int`. At head_dim 512 with nkvh=8 and block_size=32,
   `kbs` is 65536 B, so `pb*kbs` wraps once the K cache passes 2 GB — an
   entirely reachable allocation on an H200, and 4x sooner than at 128. Now
   computed in `long long` at all eight sites.

2. The shared-memory opt-in cached a process-wide `granted` flag per kernel
   instantiation. The requestable maximum differs by architecture, so one flag
   cannot be right across a mixed fleet, and a failed request left `granted`
   at 0 and fell through to a launch that then failed with no diagnostic. The
   cache is gone; the call is cheap and now always attempted.

3. `tq_attn_clocked` kept the scalar byte gather while `tq_attn` moved to a
   vector load plus the arch-gated pipeline. Identical results, but the phase-2
   stamps would have described a kernel nobody launches — at head_dim 512 that
   is 8 scalar loads measured against the 1 vector load actually executed. The
   loop now mirrors `tq_attn`.

Verified without nvcc by type-checking the kernel bodies against a CUDA shim
(clang++ -Wall -Wextra -Wshadow, both __CUDA_ARCH__ branches): clean. Still not
run on a GPU.
…ine guard

The gate exited 1 when gpu_box_preflight.sh was absent. That read as a code
failure for a couple of minutes when it was really 'this machine cannot
answer'. Per D18 rule 2: 0 pass, 1 genuine failure (the only signal to act
on), 2 environment could not answer.

Nine conditions now exit 2 — missing/failing/incomplete preflight, no repo, no
nvcc, unreachable remote, un-checkout-able branch, no archive to inspect, no
cuobjdump. Four exit 1, all genuine: nvcc compile failure, a native build
failure, a test failure, and the archive missing an arch it claims to carry.

The preflight is now looked up at ARC_PREFLIGHT, then /root/arc-tools, then
/usr/local/lib/arc, and only last inside the repo. It lives on
wave61/box-preflight-shared-prefix, so a repo-relative path vanishes the moment
a runner checks out any other branch — which is exactly how this gate came to
refuse to run.

Guarded the sourced-preflight landmine directly. gpu_box_preflight.sh ends on

    [ "$_arc_pf_sourced" = "1" ] || exit 1

so when SOURCED on a *failing* box that test succeeds, short-circuits the
'|| exit 1', and becomes the last command — 'source pf || handler' returns 0
and the handler never fires. The gate ignores the return value entirely and
reads _ARC_PF_FAILED, treating an unset flag as 'did not reach a verdict'
rather than as a pass.

Also reordered so the sm_90a+sm_100a+sm_103a compile is STEP 1 and needs no
full workspace build, and it now asserts on cuobjdump output that both cubins
are actually in the archive — a build that merely succeeds is not the claim.

Exit paths verified locally against a stub reproducing the real preflight's
sourced-exit structure: all four environment cases return 2.
…hat is asked for

Widening the acceptance gate from head_dim == 128 to every instantiated
width was correct for what a user may REQUEST, but PagedCacheType::
TurboQuant carries #[default] (cache_engine.rs:22), so it also silently
moved standard-layout models at head_dim 64/256/512 onto kernels that have
never executed — and silently dropped prefix caching with them, since
supports_prefix_cache() is false for every TurboQuant variant. Two
regressions, neither visible at the call site, for users who never asked
for TurboQuant at all.

This repo has paid for that exact shape once: FP8 KV shipped default-on and
unmeasured (wave43-BU) and every V4 request died. #98 cites that incident
by name as its own reason not to default itself on; this takes the same
choice.

resolve_for_model already distinguished explicit from ambient-default in
order to pick error-vs-fallback, so the fix is to let that same 'forced'
flag also pick WHICH head-dim set applies:

  compiled  -> widens what you may ask for   (TURBOQUANT_CUDA_HEAD_DIMS)
  measured  -> widens what you get unasked   (TURBOQUANT_DEFAULT_HEAD_DIMS)

TURBOQUANT_DEFAULT_HEAD_DIMS is [128] and moves when a hardware gate
passes, not when a kernel compiles. Nothing about the kernels changes: the
six 'if (hs != 128) return;' no-ops (D18 instance 3) and the 48 KB dynamic
smem cap fix are untouched, and explicit --pa-cache-type turboquant still
reaches every instantiated width including V4's 512.

The unsupported-geometry message now separates 'no kernel exists' from 'a
kernel exists but is unmeasured, so the default will not choose it for
you' — the second tells the user to opt in rather than to give up.

Tests: the old fixture asserted the regression (its expected-to-fall-back
list moved [64,96,192,256] -> [96,192,320,1024] precisely because 64 and
256 stopped falling back). Replaced with an explicit-path test over every
instantiated width, plus two pins on the default path. Both pins were
mutation-checked by reverting the const to the full set: both fail, and the
first revision of one passed vacuously because widening the const emptied
its loop, so it now asserts the loop is non-empty before iterating.
This PR branched before #101 ("correct the public record") landed, so it
had never seen that text. Merging master in makes four README claims
false — three falsified by this PR's own kernels, one that was already
wrong when it was written.

Falsified by this PR:

  - L31  "the paged kernel exists at head_dim 128 only"
  - L31  "there is no kernel at head_dim 512, so DeepSeek V4 cannot use it"
         DeepSeek V4 reports KvCacheLayout::Standard at head_dim 512
         (normal_loaders.rs), so the 512 instantiation genuinely reaches it.
  - L141 an explicit `--pa-cache-type turboquant` off-128 "is a hard error"
         It is now accepted at any instantiated width; the hard error moved
         out to the uninstantiated set.

Already wrong before this PR:

  - L141 "TurboQuant is not the default anywhere"

    `defaults::PAGED_CACHE_TYPE` is `PagedCacheType::TurboQuant` and the
    CLI's `--pa-cache-type` has no clap default, so leaving it unset on
    CUDA gives a standard-layout head_dim-128 model TurboQuant KV with no
    flag — and silently drops prefix caching, which no TurboQuant variant
    supports. That was true on master; the narrowing in a3a5faa bounds it
    to 128 but does not remove it. Stating the opposite understated the
    risk to anyone reading the README to decide whether to opt out, so it
    is corrected in place rather than quietly reworded, and the Rust
    example's comment now says `Auto` opts *out* rather than implying
    TurboQuant is off until asked for.

Also refreshes the `--pa-cache-type` help text, which still described the
explicit path as 128-or-error.

No behaviour change: documentation and one clap doc comment.
…n arch

`mistralrs-paged-attn/build.rs` appended one `-gencode` per entry in
`ARC_CUDA_ARCHS` on top of the one cudaforge derives itself, and applied
the same `a` suffix for cap >= 90 that cudaforge's `GpuArch::auto_suffix`
does. Verified against cudaforge 0.1.5 rather than assumed:

  compute_cap.rs  auto_suffix(90).to_gencode_arg()
                    -> "-gencode=arch=compute_90a,code=sm_90a"
  builder.rs:401  let gencode_arg = gpu_arch.to_gencode_arg();
  builder.rs:405  command.arg(&gencode_arg)...
  builder.rs:412  for arg in &self.extra_args { command.arg(arg); }

cudaforge emits its own gencode first and then appends `extra_args`
verbatim, so on an H200 building `ARC_CUDA_ARCHS=90,100,103` nvcc received
`-gencode=arch=compute_90a,code=sm_90a` twice, byte-identical.

Fixed by hoisting the existing `get_compute_cap()` binding above the loop
(it is what decides which requested arches are genuinely extra) and
skipping the one cudaforge already covers. The later duplicate binding is
deleted; there is now one.

This is the dedupe half of #108's `arc_target::build::split_primary()`,
done locally because `arc-target` does not exist on this branch. The other
half — `verify_and_export()`, the `cuobjdump` gate that fails a build whose
archive is missing a requested arch — genuinely needs that crate and is
left to #108, which should replace this block with the two calls. Until it
lands, `arc-tools/wave64_v4_turboquant_kv_gate.sh` STEP 2 asserts the arch
list externally with `cuobjdump --list-elf`, so a missing arch is caught by
the gate even though it is not yet caught by the build.

Not verified here: whether nvcc rejects or tolerates the duplicate flag.
macOS cannot run the cuda build path. Emitting it once is correct either
way, which is why this is worth doing without waiting on that answer.
@heydryft
heydryft force-pushed the feat/turboquant-hd512-v4 branch from 0c71381 to be7764b Compare August 21, 2026 16:21
@heydryft
heydryft merged commit 73ff79d into release/openrouter-ready Aug 21, 2026
18 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant