Skip to content

feat(qtip): Stage 3 — dispatch and selection for K=9, default OFF - #172

Merged
heydryft merged 3 commits into
release/openrouter-readyfrom
feat/qtip-k9-dispatch
Aug 21, 2026
Merged

heydryft merged 3 commits into
release/openrouter-readyfrom
feat/qtip-k9-dispatch

Conversation

@heydryft

Copy link
Copy Markdown
Contributor

Stage 3 of 3. Stacked on #170, which is stacked on #168 — merge in that order (D20). Base is feat/qtip-k8v4l12-format.

A K=9/V=4/L=12 layer is now served rather than refused, and there is a way to ask for the rung. K=8 gets nothing beyond staying the compiled control — it is quality-closed (−0.00698, −0.00307 even with a converged trellis-Lloyd codebook, against a ±0.0008 band) and serving it would be wasted work. K=9 is the target: Δw_cos +0.00402, 5× better than the shipped control, same 32,768 B table, same L=12/V=4.

Wired

  • CPU decode for the family, routed through trellis_v4l12::Rung — which already owns the extraction arithmetic and is bit-exactly gated against the GPU kernel. Deliberately not restated: a format with two implementations has none.
  • dequantize_weights / dequantize_w / forward (2-D).
  • forward_v4l12_gemv_cuda — the family's single GPU path, the single-token fused decode+gemv the kernel exists for. Multi-token falls back to the CPU decode, not to a K=4 kernel: slow and right beats fast and wrong.
  • Row-scale hoist OFF. It is the one lever that costs bit-exactness against the CPU reference, and there is no device measurement to trade that for.

Still refused, loudly

Every path with no V=4 kernel: 3-D gather_forward, dequantize_expert, 3-D dequantize_weights, and qtip_packed (whose QtipPackedView has no geometry field, so a consumer would read K=9 bytes as K=4 nibbles). A half-wired dispatch that falls through is worse than one that refuses — at 2 bpw a K=4 misread does not fault, it serves wrong weights. the_k9_path_is_not_the_k4_decoder_in_disguise decodes the same bytes both ways and requires the answers to differ, so "it returned numbers" cannot pass for correctness.

Selection, and default OFF

ARC_QTIP_GEOMETRY = k4v2l16 | k<K>v4l12, mirroring ARC_QTIP_CODEBOOK. Unknown values refused, never resolved.

QtipGeometry::DEFAULT stays K4V2L16. Nothing in this family has touched a GPU — the kernel has never been compiled at any K. The precedent is specific: the last KV-storage change that shipped default-on with only CPU validation killed every request on the first V4 forward that met a real device. Its arithmetic was proven and its device-time cost was not.

Guards shown red (S1–S7) — three holes, all mine

  1. S2 and S5 sit behind #[cfg(feature = "cuda")], which no test lane compiles, so both mutations passed silently. Now pinned over the source — that a family layer never reaches the K=4 CUDA path, and that the hoist stays Off. A guard that only exists on hardware nobody has is not a guard.
  2. S7 was vacuous: the env test reimplemented the parse rule inline and asserted the replica, so the real reader could resolve an unsupported K to the default and stay green. from_env now delegates to a pure from_spec, and the test calls that.
  3. I deleted my own five new tests mid-edit — a second replacement span swallowed what the first had inserted, leaving two dangling doc citations and a suite that looked green at 337 passing. Restored and verified present by name, not by count.

Also: the new source guard fired on itself. include_str!("mod.rs") includes the test module, so a negative assertion naming the string it forbids matched its own text. Scoped to the serving half. Second time this class bit me here (the first was a barrier guard counting __syncthreads in a comment); the note is in the code.

Scoped clippy clean; cargo test -p mistralrs-quant 343 + 5 passing; mistralrs-core and arc-engine compile.

What remains

  • Nothing can produce a K=9 artifact yet — the quantizer is feat(qtip/bake): geometry-parametric beam encoder, ready to bake K=8/V=4/L=12 #169's scope. This PR makes one loadable and servable once it exists.
  • The tail-clamp format decision is still deferred, correctly: padding the row stride is expensive to reverse and the census is pricing padded-vs-clamped at K=9/K=10.
  • No compiled measurement, at any K. The nvcc lanes remain queued.

@github-actions

github-actions Bot commented Aug 19, 2026 •

Copy link
Copy Markdown
Code Metrics Report
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
 Language              Files        Lines         Code     Comments       Blanks
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
 C Header                  5          305          210           52           43
 CSS                       2         1181         1036           34          111
 CUDA                     81        29475        20161         6275         3039
 Dockerfile                1           39           22            8            9
 JavaScript               16         3546         2676          482          388
 Jinja2                    7          694          656            5           33
 JSON                     75         4896         4893            0            3
 Makefile                  1            6            5            0            1
 Metal Shading Lan|       33        12224         9431         1142         1651
 PowerShell                1          300          227           30           43
 Python                  147        15371        12677          824         1870
 Shell                    42        10304         6858         2769          677
 Plain Text                4         3801            0         2479         1322
 TOML                     33         1498         1294           54          150
 YAML                      3           25           23            2            0
─────────────────────────────────────────────────────────────────────────────────
 HTML                      4         2687         2604           43           40
 |- CSS                    2          543          479           37           27
 |- JavaScript             1         1233         1215           12            6
 (Total)                             4463         4298           92           73
─────────────────────────────────────────────────────────────────────────────────
 Jupyter Notebooks         4          122           83           23           16
 |- Markdown               1           60           30           22            8
 |- Python                 1          122          113            1            8
 (Total)                              304          226           46           32
─────────────────────────────────────────────────────────────────────────────────
 Markdown                213        48932            0        38081        10851
 |- BASH                  72         1655         1203          331          121
 |- C                      4           19           19            0            0
 |- CUDA                   2           84           56           16           12
 |- JSON                  19          779          779            0            0
 |- PowerShell             1            1            1            0            0
 |- Python                23         1008          787          113          108
 |- Rust                  68         2063         1727           78          258
 |- TOML                   6          207          164            0           43
 |- YAML                   5           41           36            5            0
 (Total)                            54789         4772        38624        11393
─────────────────────────────────────────────────────────────────────────────────
 Rust                    681       340988       293252        18096        29640
 |- Markdown             504        32562          471        28063         4028
 (Total)                           373550       293723        46159        33668
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
 Total                  1353       516771       363188        99077        54506
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

@heydryft

Copy link
Copy Markdown
Contributor Author

Retargeted at the integration branch

Base changed: master → release/openrouter-ready.

The queue is being restructured to the shape the owner asked for: one PR open against master (#194), with everything else merging into a single integration branch. Agents branch off release/openrouter-ready and PR back into it; #194 is the single gate from there to master.

Two things had to land on master first for this to be workable, and both have:

  1. The base-branch CI lane used to hard-fail any PR not targeting master (ci(ArcGate): let the base-branch lane permit the single integration branch — 7 PRs are red for addressing, not quality #216). Seven PRs were red for their addressing, not their quality. The lane now accepts master or release/openrouter-ready — an exact-match allowlist, nothing else — and its intent is intact: release/openrouter-ready reaches master only through Arc → OpenRouter-ready: the single integration PR (everything else merges into this branch) #194, whose own base is master, so nothing enters master without a full run whose base is master.
  2. release/openrouter-ready was re-cut from current master. It had diverged badly — merging it as-is would have reverted TCFRAG (perf(ArcQuant/ArcKernels): TCFRAG-2B — put the qtip2b trellis GEMV on the tensor cores #203) and re-broken the flag-polarity work (fix(ArcGate): read boolean ARC_* flags by value, so =0 means off #212). It is now master plus only the work it uniquely carried.

This PR was not closed and is not considered stale. An audit of the queue found the overwhelming majority of it to be real work that was never merged, not noise.

What you need to do: rebase onto release/openrouter-ready and resolve conflicts against it rather than against master. CI lanes will now actually run and report on this PR instead of failing at the base check.


📚 Stack order — this is the TOP

This PR used to be based on its parent PR's branch, which is exactly the pattern the base-branch CI lane was built to refuse: a base that can be rebased or abandoned underneath it, so a green here described a tree that might never exist.

Both are now retargeted at release/openrouter-ready and are siblings rather than a stack. That makes the merge order load-bearing and no longer enforced by the base ref:

Merge the bottom PR FIRST, then this one. Until the bottom lands, this PR's diff against the integration branch will also contain the bottom's commits.

heydryft and others added 3 commits August 21, 2026 17:46
A K=9/V=4/L=12 layer is now SERVED rather than refused, and there is a way to
ask for the rung. K=8 stays exactly as it was: the compiled control, nothing
spent on serving it.

WIRED
  * CPU decode for the whole V=4/L=12 family, routed through
    `trellis_v4l12::Rung` — which already owns the extraction arithmetic and is
    bit-exactly gated against the GPU kernel. Not restated here: a format with
    two implementations has none.
  * `dequantize_weights` / `dequantize_w` / `forward` (2-D) for the family.
  * `forward_v4l12_gemv_cuda`: the family's single GPU path, the single-token
    fused decode+gemv the kernel exists for. Row-scale hoist OFF — it is the
    one lever that costs bit-exactness against the CPU reference, and there is
    no device measurement to trade that for.

STILL REFUSED, LOUDLY
Every path with no V=4 kernel: 3-D `gather_forward`, `dequantize_expert`,
3-D `dequantize_weights`, and `qtip_packed` (whose `QtipPackedView` carries no
geometry field, so a consumer would read K=9 bytes as K=4 nibbles). A
half-wired dispatch that falls through is worse than one that refuses — at
2 bpw a K=4 misread does not fault, it just serves wrong weights.

SELECTION
`ARC_QTIP_GEOMETRY` = `k4v2l16` | `k<K>v4l12`, mirroring `ARC_QTIP_CODEBOOK`.
Unknown values are refused, never resolved.

DEFAULT OFF, and this is the load-bearing line: `QtipGeometry::DEFAULT` stays
`K4V2L16`. Nothing in this family has touched a GPU — the kernel has never been
compiled at any K — and the precedent is specific: the last KV-storage change
that shipped default-on with only CPU validation killed every request on the
first V4 forward that met a real device. Its arithmetic was proven and its
device-time cost was not. Ship the path, leave the default alone.

Guards shown red (S1-S7). THREE HOLES, all mine:
  * S2/S5 sit behind `#[cfg(feature = "cuda")]`, which no test lane compiles,
    so both mutations passed. Now pinned over the SOURCE — a guard that only
    exists on hardware nobody has is not a guard.
  * S7: the env test reimplemented the parse rule inline and asserted the
    replica, so the real reader could silently resolve an unsupported K to the
    default and stay green. `from_env` now delegates to a pure `from_spec` and
    the test calls that.

I also deleted my own five new tests mid-edit — a second replacement span
swallowed what the first had just inserted, leaving two dangling doc citations
and a suite that looked green. Restored and verified present by name.

And the source guard fired on itself: `include_str!("mod.rs")` includes the
test module, so a negative assertion naming the string it forbids matched its
own text. Scoped to the serving half. That is the second time this class bit
me here; the note is in the code.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpVNMpb13HkUXqSqbN1o9H
The K=9 CPU decode sized its rows at the data length. Rows now carry tail
padding so the kernel's compile-time-width extraction needs no clamp, and that
padding is part of what the artifact stores — so a decoder that assumes the
data length disagrees with the loader about where row N starts.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpVNMpb13HkUXqSqbN1o9H
`include_str!("mod.rs")` yields CRLF on a Windows checkout, so
`split_once("#[cfg(test)]\nmod tests {")` never matched and the guard hit
its `expect` instead of asserting anything — the guard was not failing,
it was not running. Normalise line endings before the split; every
assertion after it already goes through the whitespace-squeezer.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@heydryft
heydryft force-pushed the feat/qtip-k9-dispatch branch from 260f0b2 to 90dd68d Compare August 21, 2026 16:47
@heydryft
heydryft merged commit 39cc0b5 into release/openrouter-ready Aug 21, 2026
18 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant