Skip to content

feat(v4): TurboQuant KV storage for DeepSeek-V4 — a V4CachedK variant, + honour qs/softcapping - #98

Open
heydryft wants to merge 13 commits into
release/openrouter-readyfrom
feat/turboquant-v4-cachedk
Open

heydryft wants to merge 13 commits into
release/openrouter-readyfrom
feat/turboquant-v4-cachedk

Conversation

@heydryft

Copy link
Copy Markdown
Contributor

Stacked on #94 (feat/turboquant-hd512-v4). Review #94 first.

#94 finished the kernel side and deliberately did not ship the integration. This is that integration — and the head_dim was never what stood in the way.

🔑 V4 and TurboQuant are the same two-region shape

dsv4_attention reads the raw key cache over exactly one span: the trailing t_q + window - 1 tokens (raw_keep_span). Every earlier raw column is -inf on every query row, because V4's raw branch is a sliding window, not dense causal attention. Distant context never arrives through the raw cache at all — it comes from compressed_kv, which the compressor builds from its own rolling history (XsRollingCache) and which never touches those bytes.

TurboQuant V4
fp16_window — recent tokens kept uncompressed the raw sliding window
the packed region everything older, which only the compressor reads

Same boundary, so this builds them as one mechanism. The payoff is the thing that makes it worth shipping:

Nothing on the decode path is ever dequantized. span over the reachable window is a narrow of a tensor that was never compressed. The compressed region is only ever written.

That is the opposite of arc_turbo::TurboQuantSingleCache, whose current_data reconstructs every compressed token on the host on every call — the documented reason that path is opt-in. A dequantizing path exists here too (span must stay total, and a deep rollback can ask for one), but the model never takes it, and a test pins that.

What landed

KvCache::V4Turbo / V4TurboKCache, modelled directly on XsRollingCache: a grown capacity buffer of packed records plus a bounded dense tail, reporting length in tokens and mapping truncation onto both time bases. Norms ride inside the code record rather than in a third buffer, because BatchSrc gives a cache slot exactly two tensors and a third would not survive clone-in/clone-out.

V4CachedK::Turbo is wired into the eager decode arm behind ARC_V4_TURBOQUANT=1.

Bytes per token per layer at V4-Flash's geometry (head_dim=512, MQA, BF16):

layout bytes
dense BF16 (512 + 1-wide marker) 1,026
FP8 codes (ARC_V4_FP8_KV=1) 590
this (256 packed + one f32 norm) 260

The dense tail stays dense and is bounded by window + max(t_q-1, margin) + EVICT_CHUNK — independent of context length.

Eviction is chunked, on purpose

TurboQuantLayout::quantize_into is a host function over &[f32]. Compressing every decode step would be one device→host→device round trip per layer per token — ~86 syncs a token across V4's 43 layers, which is CLAUDE.md pitfall #5 and would cost exactly the decode speed this is supposed to be free of. EVICT_CHUNK = 128 amortises it to ~0.34 syncs/token.

Two refusals rather than fallbacks (D18)

  • ARC_V4_TURBOQUANT=1 + ARC_V4_STANDARD_DENSE=1 is a hard error. That ablation makes ratio-0 layers attend densely over the whole cache instead of through the window, so the reachability argument stops holding and the compressed region would be silently missing from those layers' softmax.
  • append_kv_turbo refuses a slot it did not build, naming both variants. Silently falling back to the dense path would make "flag on" and "flag doing nothing" identical.

For the same reason the codec emits one Once-guarded line on its first eviction. Without it, flag-on and flag-on-but-never-compressing are indistinguishable from outside the process — and that distinction is the only thing the GPU gate can actually check.

Also: two silently-ignored kernel parameters

Pre-existing, and the same failure class as the six if (hs != 128) return; no-ops #94 removed — accepted, plumbed through the FFI, never read.

  • qs (query stride) is now honoured: query[sidx*qs + hidx*HS + i]. For a contiguous query qs == nh*HS, so this is the same address every current caller already computed. The kernel still assumes the head and dim axes are packed, so the Rust wrapper now checks that and bails naming the actual strides.
  • softcapping is applied with the same arithmetic and at the same position as pagedattention.cuh:276 — after the scale multiply, before both the logit store and the running max (capping only one would leave the softmax renormalising against a maximum no stored logit can reach). softcapping == 1.0 stays free via a grid-uniform hoisted branch.

📊 Measured — CPU unit tests (D14), not hardware

cargo test -p mistralrs-core --lib → 386 passed, 0 failed (was 377) · --test synthetic_load_smoke 13 · --test kv_sharing_workloads 7 · -p mistralrs-quant turboquant 32 · workspace cargo check green · scoped clippy lane exit 0 · rustfmt drift like-for-like (v4_turbo 0, kv_cache/mod 0→0, deepseek4 30→30).

The load-bearing test is turbo_and_dense_cached_spans_agree_bit_exactly: across a real prefill + 400 decode steps with periodic rollbacks, the span the TurboQuant cache hands dsv4_attention is bit-identical to the same range of the dense cache. Bit-identical, not close — a tolerance there would mean the codec had reached tokens the model can still see. Combined with #94's existing raw_prefix_is_equivalent_to_passing_the_whole_cache, that means the attention output is unchanged.

Mutation runs — and the three that survived

over-evict: reach = window only ................ 2 FAILED
span ignores the dense-tail offset ............. 1 FAILED
norms dropped from the record .................. 2 FAILED
retention_floor: drop the margin term .......... 2 FAILED
retention_floor: drop the t_new term ........... 1 FAILED
retention_floor: drop the max() ................ 2 FAILED

⚠️ Three initially survived the integration fixture. Two were fixture gaps and one was not:

  • span ignores the offset survives it legitimately — the model only ever asks for off == 0, so the mutation is unreachable from that path. Caught by the direct span tests.
  • Dropping the rollback margin survived because nothing in the fixture ever rolled back. Fixed by making rollbacks part of the run.
  • It still survived after that, and the reason is the real finding: EVICT_CHUNK (128) is ≫ margin (16), so chunked eviction holds base so far below the retention floor that no end-to-end fixture can make the margin term binding. The arithmetic is defence in depth for a smaller chunk or an eager eviction. So it was extracted as a pure retention_floor() and tested directly over the regime where it does bind — where all three of its terms now fail when mutated.

Recording it because a mutation table nobody failed is a table nobody ran.

🔴 The honest gap

Zero GPU exposure. Neither the CUDA edit nor the cache has been on a box — the qs/softcapping fix has never even been compiled, since nvcc does not exist on macOS and the cuda module is cfg-gated so cargo check never parsed its Rust side either. arc-tools/wave64_v4_turboquant_kv_gate.sh is written and self-contained; STEP 1 is that first compile, STEP 6 is the engagement check, and STEPS 4/5/7 are decode tok/s and peak KV bytes with the flag off vs on.

supports_paged_attention still returns Ok(false) for V4 and is untouched: V4 is eager-NormalCache only, so the paged TurboQuant kernel remains unreachable for it at any head dim. That is a separate change, and flipping it blind is how wave43-BU went.

Default stays off until the gate runs. Default-on is the destination; an unmeasured default-on is how we lose a day.

🤖 Generated with Claude Code

heydryft and others added 5 commits August 17, 2026 12:55
…/512

TurboQuant's paged kernels were hardcoded to head_dim 128 — `constexpr int
HS = 128`, a single `SGN[128]` sign table, `rotate128`, and six `extern "C"`
launchers each opening with `if (hs != 128) return;`. That early return is a
silent no-op: the output buffer is left uninitialized, so a mismatched head
dim produces garbage *fast* rather than failing.

Kernels
- `tq_attn`, `tq_attn_clocked`, `tq_cache_k`, `tq_cache_v` and the two clocked
  cache variants now take HS as a template parameter. NT stays 128 for every
  HS, so the warp-reduction topology is unchanged; wider heads give each
  thread DPT = HS/NT output dims instead of adding threads. At HS=128 the
  emitted work is the same as before.
- The per-lane packed-K gather is now one vector load (`tq_load_klane`)
  instead of BPL scalar byte loads. The x=16 cache swizzle keeps each lane's
  run contiguous and naturally aligned, which `packed_k_bytes_are_swizzle_aligned`
  asserts.
- Every `if (hs != 128) return;` is replaced by a dispatch over the
  instantiated widths.

D16 dual-arch
- `TQ_PREFETCH` software-pipelines the K gather on SM90/SM100/SM103 and runs
  flat below Hopper — same arithmetic, scheduled for the part it runs on.
- Launches now opt in to the arch's real dynamic shared-memory budget. The
  logits array scales with context length and CUDA caps dynamic shared memory
  at 48 KB without an explicit opt-in, so contexts past ~12k tokens could not
  launch at all on any arch.
- `ARC_CUDA_ARCHS=90,100,103` emits cubins for Hopper *and* Blackwell rather
  than only the build box's GPU.

Tables
- Codebooks and sign diagonals are dimension-dependent, so 512 needs its own.
  All 20 tables are emitted by `emit_turboquant_cuda_tables`, which calls the
  same `generate_signs`/`get_codebook` the Rust compressor uses — correct by
  construction, not by transcription. The generator reproduces the shipped
  d=128 table byte-for-byte.
- `turboquant::cuda_tables` re-derives every table at test time and diffs the
  checked-in header, and pins the Rust head-dim list against both the kernel's
  dispatch switch and the FFI wrapper's list. Nothing previously asserted that
  the tree's four copies of the sign table agreed with the generator or with
  each other.

Gates
- `TURBOQUANT_HEAD_DIM = 128` exact-match becomes membership in
  `TURBOQUANT_CUDA_HEAD_DIMS`; K and V are compressed independently and no
  longer have to be equal widths.

Not measured on a GPU: the kernels are written and gated but this box has no
nvcc, so nothing here is a claim about hardware behaviour.
… kernels

Proves the four head-dim instantiations compile for sm_90a AND sm_100a/sm_103a
(D16), not just the build box's GPU, and that the shipping decode path did not
regress. Deliberately contains no quality A/B: TurboQuant quality is settled by
prior measurement and re-proving it would spend GPU budget on a known number.
Every other .cu in this crate had a rerun-if-changed line; this one did not.
Emitting any rerun-if-changed disables cargo's watch-the-whole-package default,
so edits to the largest CUDA file in the crate were relying entirely on
cudaforge's object cache to trigger a rebuild.
…d-kernel parity

Three problems from review of the head_dim generalization.

1. 32-bit overflow on cache byte offsets. `pb*kbs`, `pb*vbs`, `bi*cbs` and
   `bi*vbs` were all `int`. At head_dim 512 with nkvh=8 and block_size=32,
   `kbs` is 65536 B, so `pb*kbs` wraps once the K cache passes 2 GB — an
   entirely reachable allocation on an H200, and 4x sooner than at 128. Now
   computed in `long long` at all eight sites.

2. The shared-memory opt-in cached a process-wide `granted` flag per kernel
   instantiation. The requestable maximum differs by architecture, so one flag
   cannot be right across a mixed fleet, and a failed request left `granted`
   at 0 and fell through to a launch that then failed with no diagnostic. The
   cache is gone; the call is cheap and now always attempted.

3. `tq_attn_clocked` kept the scalar byte gather while `tq_attn` moved to a
   vector load plus the arch-gated pipeline. Identical results, but the phase-2
   stamps would have described a kernel nobody launches — at head_dim 512 that
   is 8 scalar loads measured against the 1 vector load actually executed. The
   loop now mirrors `tq_attn`.

Verified without nvcc by type-checking the kernel bodies against a CUDA shim
(clang++ -Wall -Wextra -Wshadow, both __CUDA_ARCH__ branches): clean. Still not
run on a GPU.
V4's raw K cache and TurboQuant are the same two-region shape with the same
boundary, so this builds them as one mechanism rather than two.

`dsv4_attention` reads raw keys over exactly the trailing `t_q + window - 1`
tokens; everything older is `-inf` on every query row, because V4's raw branch
is a sliding window and distant context arrives through `compressed_kv`
instead. So TurboQuant's `fp16_window` IS V4's sliding window and its
compressed region IS everything older. The consequence is that nothing on the
decode path is ever dequantized: `span` over the reachable window is a narrow
of a tensor that was never compressed.

* `KvCache::V4Turbo` / `V4TurboKCache`, modelled on `XsRollingCache` — a grown
  capacity buffer of packed records plus a bounded dense tail, reporting length
  in tokens and mapping truncation onto both time bases. Norms ride inside the
  code record because `BatchSrc` gives a slot exactly two tensors.
* `V4CachedK::Turbo`, wired into the eager decode arm behind
  `ARC_V4_TURBOQUANT=1`. 260 B/token at V4-Flash's geometry against 1,026
  dense and 590 for FP8 KV.
* Eviction is chunked (`EVICT_CHUNK = 128`): `quantize_into` is a host
  function, so per-step eviction would be ~86 device syncs per token.
* `ARC_V4_TURBOQUANT` + `ARC_V4_STANDARD_DENSE` is a hard error — that
  ablation makes ratio-0 layers read past the window, which is the assumption
  eviction rests on.
* One `Once`-guarded log line on the first eviction: flag-on and
  flag-on-but-never-compressing are otherwise indistinguishable from outside
  the process.

Also fixes two pre-existing silently-ignored kernel parameters: `qs` (query
stride) is now honoured — `query[sidx*qs + ...]`, a no-op for the contiguous
queries every current caller passes — with the head/dim strides the kernel
still assumes now checked in the Rust wrapper; and `softcapping` is applied
with the same arithmetic and at the same position as `pagedattention.cuh:276`.

Default stays off. wave43-BU is why.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@github-actions

github-actions Bot commented Aug 17, 2026 •

Copy link
Copy Markdown
Code Metrics Report
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
 Language              Files        Lines         Code     Comments       Blanks
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
 C Header                  5          305          210           52           43
 CSS                       2         1181         1036           34          111
 CUDA                     72        24328        17592         4018         2718
 Dockerfile                1           39           22            8            9
 JavaScript               16         3546         2676          482          388
 Jinja2                    7          694          656            5           33
 JSON                     74         4600         4597            0            3
 Makefile                  1            6            5            0            1
 Metal Shading Lan|       33        12224         9431         1142         1651
 PowerShell                1          300          227           30           43
 Python                  143        14830        12217          797         1816
 Shell                    22         5235         3623         1280          332
 Plain Text                4         3801            0         2479         1322
 TOML                     33         1485         1292           43          150
 YAML                      3           25           23            2            0
─────────────────────────────────────────────────────────────────────────────────
 HTML                      4         2687         2604           43           40
 |- CSS                    2          543          479           37           27
 |- JavaScript             1         1233         1215           12            6
 (Total)                             4463         4298           92           73
─────────────────────────────────────────────────────────────────────────────────
 Jupyter Notebooks         4          122           83           23           16
 |- Markdown               1           60           30           22            8
 |- Python                 1          122          113            1            8
 (Total)                              304          226           46           32
─────────────────────────────────────────────────────────────────────────────────
 Markdown                187        37238            0        28565         8673
 |- BASH                  68         1610         1185          309          116
 |- C                      1           10           10            0            0
 |- CUDA                   2           84           56           16           12
 |- JSON                  18          708          708            0            0
 |- PowerShell             1            1            1            0            0
 |- Python                23         1008          787          113          108
 |- Rust                  65         2048         1713           77          258
 |- TOML                   6          207          164            0           43
 |- YAML                   4           38           33            5            0
 (Total)                            42952         4657        29085         9210
─────────────────────────────────────────────────────────────────────────────────
 Rust                    660       305610       264700        13744        27166
 |- Markdown             475        22145          471        19045         2629
 (Total)                           327755       265171        32789        29795
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
 Total                  1272       448073       327959        72384        47730
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

…ine guard

The gate exited 1 when gpu_box_preflight.sh was absent. That read as a code
failure for a couple of minutes when it was really 'this machine cannot
answer'. Per D18 rule 2: 0 pass, 1 genuine failure (the only signal to act
on), 2 environment could not answer.

Nine conditions now exit 2 — missing/failing/incomplete preflight, no repo, no
nvcc, unreachable remote, un-checkout-able branch, no archive to inspect, no
cuobjdump. Four exit 1, all genuine: nvcc compile failure, a native build
failure, a test failure, and the archive missing an arch it claims to carry.

The preflight is now looked up at ARC_PREFLIGHT, then /root/arc-tools, then
/usr/local/lib/arc, and only last inside the repo. It lives on
wave61/box-preflight-shared-prefix, so a repo-relative path vanishes the moment
a runner checks out any other branch — which is exactly how this gate came to
refuse to run.

Guarded the sourced-preflight landmine directly. gpu_box_preflight.sh ends on

    [ "$_arc_pf_sourced" = "1" ] || exit 1

so when SOURCED on a *failing* box that test succeeds, short-circuits the
'|| exit 1', and becomes the last command — 'source pf || handler' returns 0
and the handler never fires. The gate ignores the return value entirely and
reads _ARC_PF_FAILED, treating an unset flag as 'did not reach a verdict'
rather than as a pass.

Also reordered so the sm_90a+sm_100a+sm_103a compile is STEP 1 and needs no
full workspace build, and it now asserts on cuobjdump output that both cubins
are actually in the archive — a build that merely succeeds is not the claim.

Exit paths verified locally against a stub reproducing the real preflight's
sourced-exit structure: all four environment cases return 2.
… faults

The wave64 gate passed `--features "cuda flash-attn"` to
`-p mistralrs-paged-attn`, which declares only `cuda`/`metal`. STEP 2 died
before it ran, so D16 dual-arch was never actually tested. STEP 1 — the first
compile of the qs/softcapping kernel fix — passed on an H200.

* STEP 2 now passes `--features cuda` only.
* A feature-sanity check runs BEFORE the 9-minute build, reading the manifests
  themselves, so this class of bug costs seconds rather than a gate run.
* STEP 2 now ASSERTS the cubin arches instead of printing them. v1 would have
  reported a missing sm_100a and still said PASS, which is the same
  absence-read-as-success shape the gate exists to catch.
* STEP 3 reuses STEP 1's feature set so it does not force a second full
  non-cuda rebuild.
* Exit codes (D18 rule 2): 1 means the code under test failed, 2 means
  environment/harness — missing toolkit, model, repo, or a feature flag this
  script got wrong. A missing UQFF now exits 2 rather than 0, so a partial run
  can never read as a pass.

Both new logic blocks were checked against the real manifests and against
synthetic arch lists before committing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@heydryft

Copy link
Copy Markdown
Contributor Author

🟢 First hardware result: the never-before-compiled CUDA compiles

Gate ran on a live H200.

[13:16:35] STEP 1 build sm_90 (native) — first compile of the qs/softcapping fix
[13:25:29] PASS build sm_90                     <- 8m53s, clean

That was the step that mattered — my own framing was "STEP 1 failing = the .cu edit is broken". It didn't fail. The qs / softcapping fix compiles for sm_90.

🔴 STEP 2 died on a bug in my gate script, not in the code

error: the package 'mistralrs-paged-attn' does not contain this feature: flash-attn

mistralrs-paged-attn declares only cuda / metal; flash-attn is a mistralrs-core feature. I passed the core feature set to a -p mistralrs-paged-attn build. So D16 dual-arch has not been tested at all — that step never ran. Fixed in afe56c5:

  • STEP 2 passes --features cuda only.
  • A feature-sanity check now runs before the 9-minute build, reading the manifests themselves, so this class of bug costs seconds instead of a gate run. Verified against the real manifests — it correctly reports flash-attn absent from mistralrs-paged-attn and present on mistralrs-core.
  • STEP 2 now asserts the cubin arches rather than printing them. v1 would have reported a missing sm_100a and still said PASS — the same absence-read-as-success shape this gate exists to catch. Membership logic tested against synthetic arch lists.
  • STEP 3 reuses STEP 1's feature set instead of forcing a second full non-cuda rebuild.
  • Exit-code discipline (D18 rule 2): 1 = the code under test failed; 2 = environment/harness (missing toolkit, model, repo, or a feature flag this script got wrong). A missing UQFF now exits 2 rather than 0, so a partial run can never read as a pass.

STEPS 3–7 remain outstanding, especially STEP 6 (engagement) — "flag on" and "flag on but silently no-opping" are otherwise indistinguishable from outside the process.

Follow-ups (issues are disabled on this repo, so: here + memory/mission/BACKLOG.md)

Both surfaced by this PR and deliberately not bundled:

  1. bool DO_CAP template param. Compiles the cap out of the default path so the no-cap instantiation is provably byte-identical to pre-change (cuobjdump --dump-ptx), rather than argued from "the branch is uniform". Bigger win: it is the natural point to make tq_attn_clocked's "must stay a mirror of tq_attn" invariant structural instead of comment-enforced — this PR had to hand-mirror the softcapping edit into both kernels for exactly that reason.

  2. Cap the raw KV store at the window. This PR proves the evicted region is never read, which makes "don't store it at all" available — SGLang's DSV4 pool already does this, and dsv4_attention.rs flags it as unblocked-but-separate. Beats 3.9x outright. Not strictly dominant, though: this PR keeps every token addressable, so prefix reuse and deep rollback still work; dropping makes those a hard refusal. Probably a third mode (dense / turboquant / window-only) rather than a replacement. retention_floor() here is already the "what may be dropped" predicate.

…hat is asked for

Widening the acceptance gate from head_dim == 128 to every instantiated
width was correct for what a user may REQUEST, but PagedCacheType::
TurboQuant carries #[default] (cache_engine.rs:22), so it also silently
moved standard-layout models at head_dim 64/256/512 onto kernels that have
never executed — and silently dropped prefix caching with them, since
supports_prefix_cache() is false for every TurboQuant variant. Two
regressions, neither visible at the call site, for users who never asked
for TurboQuant at all.

This repo has paid for that exact shape once: FP8 KV shipped default-on and
unmeasured (wave43-BU) and every V4 request died. #98 cites that incident
by name as its own reason not to default itself on; this takes the same
choice.

resolve_for_model already distinguished explicit from ambient-default in
order to pick error-vs-fallback, so the fix is to let that same 'forced'
flag also pick WHICH head-dim set applies:

  compiled  -> widens what you may ask for   (TURBOQUANT_CUDA_HEAD_DIMS)
  measured  -> widens what you get unasked   (TURBOQUANT_DEFAULT_HEAD_DIMS)

TURBOQUANT_DEFAULT_HEAD_DIMS is [128] and moves when a hardware gate
passes, not when a kernel compiles. Nothing about the kernels changes: the
six 'if (hs != 128) return;' no-ops (D18 instance 3) and the 48 KB dynamic
smem cap fix are untouched, and explicit --pa-cache-type turboquant still
reaches every instantiated width including V4's 512.

The unsupported-geometry message now separates 'no kernel exists' from 'a
kernel exists but is unmeasured, so the default will not choose it for
you' — the second tells the user to opt in rather than to give up.

Tests: the old fixture asserted the regression (its expected-to-fall-back
list moved [64,96,192,256] -> [96,192,320,1024] precisely because 64 and
256 stopped falling back). Replaced with an explicit-path test over every
instantiated width, plus two pins on the default path. Both pins were
mutation-checked by reverting the const to the full set: both fail, and the
first revision of one passed vacuously because widening the const emptied
its loop, so it now asserts the loop is non-empty before iterating.
This PR branched before #101 ("correct the public record") landed, so it
had never seen that text. Merging master in makes four README claims
false — three falsified by this PR's own kernels, one that was already
wrong when it was written.

Falsified by this PR:

  - L31  "the paged kernel exists at head_dim 128 only"
  - L31  "there is no kernel at head_dim 512, so DeepSeek V4 cannot use it"
         DeepSeek V4 reports KvCacheLayout::Standard at head_dim 512
         (normal_loaders.rs), so the 512 instantiation genuinely reaches it.
  - L141 an explicit `--pa-cache-type turboquant` off-128 "is a hard error"
         It is now accepted at any instantiated width; the hard error moved
         out to the uninstantiated set.

Already wrong before this PR:

  - L141 "TurboQuant is not the default anywhere"

    `defaults::PAGED_CACHE_TYPE` is `PagedCacheType::TurboQuant` and the
    CLI's `--pa-cache-type` has no clap default, so leaving it unset on
    CUDA gives a standard-layout head_dim-128 model TurboQuant KV with no
    flag — and silently drops prefix caching, which no TurboQuant variant
    supports. That was true on master; the narrowing in a3a5faa bounds it
    to 128 but does not remove it. Stating the opposite understated the
    risk to anyone reading the README to decide whether to opt out, so it
    is corrected in place rather than quietly reworded, and the Rust
    example's comment now says `Auto` opts *out* rather than implying
    TurboQuant is off until asked for.

Also refreshes the `--pa-cache-type` help text, which still described the
explicit path as 128-or-error.

No behaviour change: documentation and one clap doc comment.
…V4 KV branch

One conflict, in the V4 attention KV-append arm: master added an
`arc_profiler::device_span("kv_cache_append")` exactly where this branch
added the TurboQuant storage branch.

Resolved by keeping both and giving the TurboQuant arm its own span,
`kv_cache_append_turbo`. Eviction is the single cost this path adds to
decode, and the ON/OFF decode comparison this branch has to clear is a
throughput claim — folding eviction into the dense span's number would
have hidden the one line item the measurement is meant to expose.
…n arch

`mistralrs-paged-attn/build.rs` appended one `-gencode` per entry in
`ARC_CUDA_ARCHS` on top of the one cudaforge derives itself, and applied
the same `a` suffix for cap >= 90 that cudaforge's `GpuArch::auto_suffix`
does. Verified against cudaforge 0.1.5 rather than assumed:

  compute_cap.rs  auto_suffix(90).to_gencode_arg()
                    -> "-gencode=arch=compute_90a,code=sm_90a"
  builder.rs:401  let gencode_arg = gpu_arch.to_gencode_arg();
  builder.rs:405  command.arg(&gencode_arg)...
  builder.rs:412  for arg in &self.extra_args { command.arg(arg); }

cudaforge emits its own gencode first and then appends `extra_args`
verbatim, so on an H200 building `ARC_CUDA_ARCHS=90,100,103` nvcc received
`-gencode=arch=compute_90a,code=sm_90a` twice, byte-identical.

Fixed by hoisting the existing `get_compute_cap()` binding above the loop
(it is what decides which requested arches are genuinely extra) and
skipping the one cudaforge already covers. The later duplicate binding is
deleted; there is now one.

This is the dedupe half of #108's `arc_target::build::split_primary()`,
done locally because `arc-target` does not exist on this branch. The other
half — `verify_and_export()`, the `cuobjdump` gate that fails a build whose
archive is missing a requested arch — genuinely needs that crate and is
left to #108, which should replace this block with the two calls. Until it
lands, `arc-tools/wave64_v4_turboquant_kv_gate.sh` STEP 2 asserts the arch
list externally with `cuobjdump --list-elf`, so a missing arch is caught by
the gate even though it is not yet caught by the build.

Not verified here: whether nvcc rejects or tolerates the duplicate flag.
macOS cannot run the cuda build path. Emitting it once is correct either
way, which is why this is worth doing without waiting on that answer.
@heydryft

Copy link
Copy Markdown
Contributor Author

ArcGate: HELD on its parent — mechanically, not on merit.

Base is feat/turboquant-hd512-v4 (#94). I have absorbed master into #94 and it is running its full 16 lanes now; once it has a real CI complete and merges, this PR gets retargeted to master before anything is touched, and then needs its own absorb + lanes.

Two things worth recording while it waits:

It currently shows 1 check — the comment bot. Zero CI lanes have ever run on it. That is the stacked-PR trigger gap (pull_request: branches: [master] matched no PR whose base was another branch); the fix is now on master, so any push here starts all 16. Nothing about this PR's state should be read as "green" today.

And the claim in the body is the interesting one, so it should not get lost in the queue: that dsv4_attention reads the raw key cache over exactly one span — the trailing t_q + window - 1 tokens — because V4's raw branch is a sliding window, not dense causal attention, so every earlier raw column is -inf on every query row. If that holds, TurboQuant's fp16_window/packed split and V4's raw/compressed split are the same boundary, and the decode path never dequantizes anything — the compressed region is only ever written. That is a much stronger property than "TurboQuant works on V4", and it is the opposite of arc_turbo::TurboQuantSingleCache, which reconstructs every compressed token on the host on every call.

Reviewers: that is the sentence to check, not the diff size. The tests pin it, but the reachability argument is what makes the whole thing worth having.

Nothing deleted; branch untouched.

@heydryft

Copy link
Copy Markdown
Contributor Author

Triage verdict: REBASE, not close — but only after #94, on which it stacks.

Evidence against current master: mistralrs-core/src/kv_cache/v4_turbo.rs and the KvCache::V4Turbo variant do not exist on master, and the design this depends on is alive — V4CachedK still has 13 references in mistralrs-core/src/models/deepseek4.rs.

Conflicts: 6 files, worst are kv_cache/mod.rs, models/deepseek4.rs, lib.rs. Retarget to master once #94 lands.

Caveat for the rebase: this adds a KV storage variant. That is on-disk/in-memory format territory and the byte formats are the moat — the rebase must state whether V4Turbo becomes selectable-only or default, and it must be selectable-only. Also note the live ABI defect it sits near: cb_mult == 0 means 'gather the table' in qtip_gemv.cu:266 and qtip_grouped_gemm_lut.cu but 'compute Box-Muller' in qtip_gather_gemv.cu — do not add a third meaning.

Leaving open with the decision recorded.

@heydryft

Copy link
Copy Markdown
Contributor Author

Retargeted at the integration branch

Base changed: master → release/openrouter-ready.

The queue is being restructured to the shape the owner asked for: one PR open against master (#194), with everything else merging into a single integration branch. Agents branch off release/openrouter-ready and PR back into it; #194 is the single gate from there to master.

Two things had to land on master first for this to be workable, and both have:

  1. The base-branch CI lane used to hard-fail any PR not targeting master (ci(ArcGate): let the base-branch lane permit the single integration branch — 7 PRs are red for addressing, not quality #216). Seven PRs were red for their addressing, not their quality. The lane now accepts master or release/openrouter-ready — an exact-match allowlist, nothing else — and its intent is intact: release/openrouter-ready reaches master only through Arc → OpenRouter-ready: the single integration PR (everything else merges into this branch) #194, whose own base is master, so nothing enters master without a full run whose base is master.
  2. release/openrouter-ready was re-cut from current master. It had diverged badly — merging it as-is would have reverted TCFRAG (perf(ArcQuant/ArcKernels): TCFRAG-2B — put the qtip2b trellis GEMV on the tensor cores #203) and re-broken the flag-polarity work (fix(ArcGate): read boolean ARC_* flags by value, so =0 means off #212). It is now master plus only the work it uniquely carried.

This PR was not closed and is not considered stale. An audit of the queue found the overwhelming majority of it to be real work that was never merged, not noise.

What you need to do: rebase onto release/openrouter-ready and resolve conflicts against it rather than against master. CI lanes will now actually run and report on this PR instead of failing at the base check.


📚 Stack order — this is the TOP

This PR used to be based on its parent PR's branch, which is exactly the pattern the base-branch CI lane was built to refuse: a base that can be rebased or abandoned underneath it, so a green here described a tree that might never exist.

Both are now retargeted at release/openrouter-ready and are siblings rather than a stack. That makes the merge order load-bearing and no longer enforced by the base ref:

Merge the bottom PR FIRST, then this one. Until the bottom lands, this PR's diff against the integration branch will also contain the bottom's commits.

heydryft added a commit that referenced this pull request Aug 21, 2026
…hat is asked for

Widening the acceptance gate from head_dim == 128 to every instantiated
width was correct for what a user may REQUEST, but PagedCacheType::
TurboQuant carries #[default] (cache_engine.rs:22), so it also silently
moved standard-layout models at head_dim 64/256/512 onto kernels that have
never executed — and silently dropped prefix caching with them, since
supports_prefix_cache() is false for every TurboQuant variant. Two
regressions, neither visible at the call site, for users who never asked
for TurboQuant at all.

This repo has paid for that exact shape once: FP8 KV shipped default-on and
unmeasured (wave43-BU) and every V4 request died. #98 cites that incident
by name as its own reason not to default itself on; this takes the same
choice.

resolve_for_model already distinguished explicit from ambient-default in
order to pick error-vs-fallback, so the fix is to let that same 'forced'
flag also pick WHICH head-dim set applies:

  compiled  -> widens what you may ask for   (TURBOQUANT_CUDA_HEAD_DIMS)
  measured  -> widens what you get unasked   (TURBOQUANT_DEFAULT_HEAD_DIMS)

TURBOQUANT_DEFAULT_HEAD_DIMS is [128] and moves when a hardware gate
passes, not when a kernel compiles. Nothing about the kernels changes: the
six 'if (hs != 128) return;' no-ops (D18 instance 3) and the 48 KB dynamic
smem cap fix are untouched, and explicit --pa-cache-type turboquant still
reaches every instantiated width including V4's 512.

The unsupported-geometry message now separates 'no kernel exists' from 'a
kernel exists but is unmeasured, so the default will not choose it for
you' — the second tells the user to opt in rather than to give up.

Tests: the old fixture asserted the regression (its expected-to-fall-back
list moved [64,96,192,256] -> [96,192,320,1024] precisely because 64 and
256 stopped falling back). Replaced with an explicit-path test over every
instantiated width, plus two pins on the default path. Both pins were
mutation-checked by reverting the const to the full set: both fail, and the
first revision of one passed vacuously because widening the const emptied
its loop, so it now asserts the loop is non-empty before iterating.
@heydryft

Copy link
Copy Markdown
Contributor Author

Rebase attempt onto release/openrouter-ready — STOPPED, left open

Rebasing this branch's own two commits (02d70f070, afe56c579) onto the rebased #94 conflicts in mistralrs-core/src/models/deepseek4.rs (4 hunks) plus mistralrs-core/src/kv_cache/mod.rs and mistralrs-core/src/lib.rs.

I did not resolve them. mistralrs-core/src/models/ is held by the ArcGraph agent right now, and the conflicting hunks sit exactly in the code that agent is rewriting — the kv_cache_append device span, the append_kv_mqa call, and the Tensor::arange capture region. The last eight commits to touch this file are all ArcGraph/ArcInfer, and #190 is open against the same lines. Any resolution I wrote would race them.

Prerequisite: #94 (feat/turboquant-hd512-v4) is rebased and landing; rebase this on top of that once the ArcGraph work in deepseek4.rs has settled.

Nothing was force-pushed to this branch — it is untouched.

heydryft added a commit that referenced this pull request Aug 21, 2026
…lackwell (#94)

* feat(turboquant): instantiate the CUDA kernels at head_dim 64/128/256/512

TurboQuant's paged kernels were hardcoded to head_dim 128 — `constexpr int
HS = 128`, a single `SGN[128]` sign table, `rotate128`, and six `extern "C"`
launchers each opening with `if (hs != 128) return;`. That early return is a
silent no-op: the output buffer is left uninitialized, so a mismatched head
dim produces garbage *fast* rather than failing.

Kernels
- `tq_attn`, `tq_attn_clocked`, `tq_cache_k`, `tq_cache_v` and the two clocked
  cache variants now take HS as a template parameter. NT stays 128 for every
  HS, so the warp-reduction topology is unchanged; wider heads give each
  thread DPT = HS/NT output dims instead of adding threads. At HS=128 the
  emitted work is the same as before.
- The per-lane packed-K gather is now one vector load (`tq_load_klane`)
  instead of BPL scalar byte loads. The x=16 cache swizzle keeps each lane's
  run contiguous and naturally aligned, which `packed_k_bytes_are_swizzle_aligned`
  asserts.
- Every `if (hs != 128) return;` is replaced by a dispatch over the
  instantiated widths.

D16 dual-arch
- `TQ_PREFETCH` software-pipelines the K gather on SM90/SM100/SM103 and runs
  flat below Hopper — same arithmetic, scheduled for the part it runs on.
- Launches now opt in to the arch's real dynamic shared-memory budget. The
  logits array scales with context length and CUDA caps dynamic shared memory
  at 48 KB without an explicit opt-in, so contexts past ~12k tokens could not
  launch at all on any arch.
- `ARC_CUDA_ARCHS=90,100,103` emits cubins for Hopper *and* Blackwell rather
  than only the build box's GPU.

Tables
- Codebooks and sign diagonals are dimension-dependent, so 512 needs its own.
  All 20 tables are emitted by `emit_turboquant_cuda_tables`, which calls the
  same `generate_signs`/`get_codebook` the Rust compressor uses — correct by
  construction, not by transcription. The generator reproduces the shipped
  d=128 table byte-for-byte.
- `turboquant::cuda_tables` re-derives every table at test time and diffs the
  checked-in header, and pins the Rust head-dim list against both the kernel's
  dispatch switch and the FFI wrapper's list. Nothing previously asserted that
  the tree's four copies of the sign table agreed with the generator or with
  each other.

Gates
- `TURBOQUANT_HEAD_DIM = 128` exact-match becomes membership in
  `TURBOQUANT_CUDA_HEAD_DIMS`; K and V are compressed independently and no
  longer have to be equal widths.

Not measured on a GPU: the kernels are written and gated but this box has no
nvcc, so nothing here is a claim about hardware behaviour.

* test(gpu): wave63 gate — dual-arch compile proof for the head_dim 512 kernels

Proves the four head-dim instantiations compile for sm_90a AND sm_100a/sm_103a
(D16), not just the build box's GPU, and that the shipping decode path did not
regress. Deliberately contains no quality A/B: TurboQuant quality is settled by
prior measurement and re-proving it would spend GPU budget on a known number.

* fix(build): watch turbo_paged_attention.cu for changes

Every other .cu in this crate had a rerun-if-changed line; this one did not.
Emitting any rerun-if-changed disables cargo's watch-the-whole-package default,
so edits to the largest CUDA file in the crate were relying entirely on
cudaforge's object cache to trigger a rebuild.

* fix(turboquant): 64-bit cache offsets, per-device smem opt-in, clocked-kernel parity

Three problems from review of the head_dim generalization.

1. 32-bit overflow on cache byte offsets. `pb*kbs`, `pb*vbs`, `bi*cbs` and
   `bi*vbs` were all `int`. At head_dim 512 with nkvh=8 and block_size=32,
   `kbs` is 65536 B, so `pb*kbs` wraps once the K cache passes 2 GB — an
   entirely reachable allocation on an H200, and 4x sooner than at 128. Now
   computed in `long long` at all eight sites.

2. The shared-memory opt-in cached a process-wide `granted` flag per kernel
   instantiation. The requestable maximum differs by architecture, so one flag
   cannot be right across a mixed fleet, and a failed request left `granted`
   at 0 and fell through to a launch that then failed with no diagnostic. The
   cache is gone; the call is cheap and now always attempted.

3. `tq_attn_clocked` kept the scalar byte gather while `tq_attn` moved to a
   vector load plus the arch-gated pipeline. Identical results, but the phase-2
   stamps would have described a kernel nobody launches — at head_dim 512 that
   is 8 scalar loads measured against the 1 vector load actually executed. The
   loop now mirrors `tq_attn`.

Verified without nvcc by type-checking the kernel bodies against a CUDA shim
(clang++ -Wall -Wextra -Wshadow, both __CUDA_ARCH__ branches): clean. Still not
run on a GPU.

* fix(gate): D18 exit codes, branch-independent preflight, source-landmine guard

The gate exited 1 when gpu_box_preflight.sh was absent. That read as a code
failure for a couple of minutes when it was really 'this machine cannot
answer'. Per D18 rule 2: 0 pass, 1 genuine failure (the only signal to act
on), 2 environment could not answer.

Nine conditions now exit 2 — missing/failing/incomplete preflight, no repo, no
nvcc, unreachable remote, un-checkout-able branch, no archive to inspect, no
cuobjdump. Four exit 1, all genuine: nvcc compile failure, a native build
failure, a test failure, and the archive missing an arch it claims to carry.

The preflight is now looked up at ARC_PREFLIGHT, then /root/arc-tools, then
/usr/local/lib/arc, and only last inside the repo. It lives on
wave61/box-preflight-shared-prefix, so a repo-relative path vanishes the moment
a runner checks out any other branch — which is exactly how this gate came to
refuse to run.

Guarded the sourced-preflight landmine directly. gpu_box_preflight.sh ends on

    [ "$_arc_pf_sourced" = "1" ] || exit 1

so when SOURCED on a *failing* box that test succeeds, short-circuits the
'|| exit 1', and becomes the last command — 'source pf || handler' returns 0
and the handler never fires. The gate ignores the return value entirely and
reads _ARC_PF_FAILED, treating an unset flag as 'did not reach a verdict'
rather than as a pass.

Also reordered so the sm_90a+sm_100a+sm_103a compile is STEP 1 and needs no
full workspace build, and it now asserts on cuobjdump output that both cubins
are actually in the archive — a build that merely succeeds is not the claim.

Exit paths verified locally against a stub reproducing the real preflight's
sourced-exit structure: all four environment cases return 2.

* fix(paged): keep the TurboQuant DEFAULT at head_dim 128; widen only what is asked for

Widening the acceptance gate from head_dim == 128 to every instantiated
width was correct for what a user may REQUEST, but PagedCacheType::
TurboQuant carries #[default] (cache_engine.rs:22), so it also silently
moved standard-layout models at head_dim 64/256/512 onto kernels that have
never executed — and silently dropped prefix caching with them, since
supports_prefix_cache() is false for every TurboQuant variant. Two
regressions, neither visible at the call site, for users who never asked
for TurboQuant at all.

This repo has paid for that exact shape once: FP8 KV shipped default-on and
unmeasured (wave43-BU) and every V4 request died. #98 cites that incident
by name as its own reason not to default itself on; this takes the same
choice.

resolve_for_model already distinguished explicit from ambient-default in
order to pick error-vs-fallback, so the fix is to let that same 'forced'
flag also pick WHICH head-dim set applies:

  compiled  -> widens what you may ask for   (TURBOQUANT_CUDA_HEAD_DIMS)
  measured  -> widens what you get unasked   (TURBOQUANT_DEFAULT_HEAD_DIMS)

TURBOQUANT_DEFAULT_HEAD_DIMS is [128] and moves when a hardware gate
passes, not when a kernel compiles. Nothing about the kernels changes: the
six 'if (hs != 128) return;' no-ops (D18 instance 3) and the 48 KB dynamic
smem cap fix are untouched, and explicit --pa-cache-type turboquant still
reaches every instantiated width including V4's 512.

The unsupported-geometry message now separates 'no kernel exists' from 'a
kernel exists but is unmeasured, so the default will not choose it for
you' — the second tells the user to opt in rather than to give up.

Tests: the old fixture asserted the regression (its expected-to-fall-back
list moved [64,96,192,256] -> [96,192,320,1024] precisely because 64 and
256 stopped falling back). Replaced with an explicit-path test over every
instantiated width, plus two pins on the default path. Both pins were
mutation-checked by reverting the const to the full set: both fail, and the
first revision of one passed vacuously because widening the const emptied
its loop, so it now asserts the loop is non-empty before iterating.

* docs: correct the public record for the widened TurboQuant kernels

This PR branched before #101 ("correct the public record") landed, so it
had never seen that text. Merging master in makes four README claims
false — three falsified by this PR's own kernels, one that was already
wrong when it was written.

Falsified by this PR:

  - L31  "the paged kernel exists at head_dim 128 only"
  - L31  "there is no kernel at head_dim 512, so DeepSeek V4 cannot use it"
         DeepSeek V4 reports KvCacheLayout::Standard at head_dim 512
         (normal_loaders.rs), so the 512 instantiation genuinely reaches it.
  - L141 an explicit `--pa-cache-type turboquant` off-128 "is a hard error"
         It is now accepted at any instantiated width; the hard error moved
         out to the uninstantiated set.

Already wrong before this PR:

  - L141 "TurboQuant is not the default anywhere"

    `defaults::PAGED_CACHE_TYPE` is `PagedCacheType::TurboQuant` and the
    CLI's `--pa-cache-type` has no clap default, so leaving it unset on
    CUDA gives a standard-layout head_dim-128 model TurboQuant KV with no
    flag — and silently drops prefix caching, which no TurboQuant variant
    supports. That was true on master; the narrowing in a3a5faa bounds it
    to 128 but does not remove it. Stating the opposite understated the
    risk to anyone reading the README to decide whether to opt out, so it
    is corrected in place rather than quietly reworded, and the Rust
    example's comment now says `Auto` opts *out* rather than implying
    TurboQuant is off until asked for.

Also refreshes the `--pa-cache-type` help text, which still described the
explicit path as 128-or-error.

No behaviour change: documentation and one clap doc comment.

* fix(build): stop emitting a duplicate -gencode for the build box's own arch

`mistralrs-paged-attn/build.rs` appended one `-gencode` per entry in
`ARC_CUDA_ARCHS` on top of the one cudaforge derives itself, and applied
the same `a` suffix for cap >= 90 that cudaforge's `GpuArch::auto_suffix`
does. Verified against cudaforge 0.1.5 rather than assumed:

  compute_cap.rs  auto_suffix(90).to_gencode_arg()
                    -> "-gencode=arch=compute_90a,code=sm_90a"
  builder.rs:401  let gencode_arg = gpu_arch.to_gencode_arg();
  builder.rs:405  command.arg(&gencode_arg)...
  builder.rs:412  for arg in &self.extra_args { command.arg(arg); }

cudaforge emits its own gencode first and then appends `extra_args`
verbatim, so on an H200 building `ARC_CUDA_ARCHS=90,100,103` nvcc received
`-gencode=arch=compute_90a,code=sm_90a` twice, byte-identical.

Fixed by hoisting the existing `get_compute_cap()` binding above the loop
(it is what decides which requested arches are genuinely extra) and
skipping the one cudaforge already covers. The later duplicate binding is
deleted; there is now one.

This is the dedupe half of #108's `arc_target::build::split_primary()`,
done locally because `arc-target` does not exist on this branch. The other
half — `verify_and_export()`, the `cuobjdump` gate that fails a build whose
archive is missing a requested arch — genuinely needs that crate and is
left to #108, which should replace this block with the two calls. Until it
lands, `arc-tools/wave64_v4_turboquant_kv_gate.sh` STEP 2 asserts the arch
list externally with `cuobjdump --list-elf`, so a missing arch is caught by
the gate even though it is not yet caught by the build.

Not verified here: whether nvcc rejects or tolerates the duplicate flag.
macOS cannot run the cuda build path. Emitting it once is correct either
way, which is why this is worth doing without waiting on that answer.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant