feat(qtip): Stage 3 — dispatch and selection for K=9, default OFF - #172
Conversation
Code Metrics Report━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Language Files Lines Code Comments Blanks ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ C Header 5 305 210 52 43 CSS 2 1181 1036 34 111 CUDA 81 29475 20161 6275 3039 Dockerfile 1 39 22 8 9 JavaScript 16 3546 2676 482 388 Jinja2 7 694 656 5 33 JSON 75 4896 4893 0 3 Makefile 1 6 5 0 1 Metal Shading Lan| 33 12224 9431 1142 1651 PowerShell 1 300 227 30 43 Python 147 15371 12677 824 1870 Shell 42 10304 6858 2769 677 Plain Text 4 3801 0 2479 1322 TOML 33 1498 1294 54 150 YAML 3 25 23 2 0 ───────────────────────────────────────────────────────────────────────────────── HTML 4 2687 2604 43 40 |- CSS 2 543 479 37 27 |- JavaScript 1 1233 1215 12 6 (Total) 4463 4298 92 73 ───────────────────────────────────────────────────────────────────────────────── Jupyter Notebooks 4 122 83 23 16 |- Markdown 1 60 30 22 8 |- Python 1 122 113 1 8 (Total) 304 226 46 32 ───────────────────────────────────────────────────────────────────────────────── Markdown 213 48932 0 38081 10851 |- BASH 72 1655 1203 331 121 |- C 4 19 19 0 0 |- CUDA 2 84 56 16 12 |- JSON 19 779 779 0 0 |- PowerShell 1 1 1 0 0 |- Python 23 1008 787 113 108 |- Rust 68 2063 1727 78 258 |- TOML 6 207 164 0 43 |- YAML 5 41 36 5 0 (Total) 54789 4772 38624 11393 ───────────────────────────────────────────────────────────────────────────────── Rust 681 340988 293252 18096 29640 |- Markdown 504 32562 471 28063 4028 (Total) 373550 293723 46159 33668 ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Total 1353 516771 363188 99077 54506 ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ |
268ad8b to
634262e
Compare
5de37b6 to
3d83c5d
Compare
Retargeted at the integration branchBase changed: The queue is being restructured to the shape the owner asked for: one PR open against Two things had to land on
This PR was not closed and is not considered stale. An audit of the queue found the overwhelming majority of it to be real work that was never merged, not noise. What you need to do: rebase onto 📚 Stack order — this is the TOPThis PR used to be based on its parent PR's branch, which is exactly the pattern the Both are now retargeted at Merge the bottom PR FIRST, then this one. Until the bottom lands, this PR's diff against the integration branch will also contain the bottom's commits. |
A K=9/V=4/L=12 layer is now SERVED rather than refused, and there is a way to
ask for the rung. K=8 stays exactly as it was: the compiled control, nothing
spent on serving it.
WIRED
* CPU decode for the whole V=4/L=12 family, routed through
`trellis_v4l12::Rung` — which already owns the extraction arithmetic and is
bit-exactly gated against the GPU kernel. Not restated here: a format with
two implementations has none.
* `dequantize_weights` / `dequantize_w` / `forward` (2-D) for the family.
* `forward_v4l12_gemv_cuda`: the family's single GPU path, the single-token
fused decode+gemv the kernel exists for. Row-scale hoist OFF — it is the
one lever that costs bit-exactness against the CPU reference, and there is
no device measurement to trade that for.
STILL REFUSED, LOUDLY
Every path with no V=4 kernel: 3-D `gather_forward`, `dequantize_expert`,
3-D `dequantize_weights`, and `qtip_packed` (whose `QtipPackedView` carries no
geometry field, so a consumer would read K=9 bytes as K=4 nibbles). A
half-wired dispatch that falls through is worse than one that refuses — at
2 bpw a K=4 misread does not fault, it just serves wrong weights.
SELECTION
`ARC_QTIP_GEOMETRY` = `k4v2l16` | `k<K>v4l12`, mirroring `ARC_QTIP_CODEBOOK`.
Unknown values are refused, never resolved.
DEFAULT OFF, and this is the load-bearing line: `QtipGeometry::DEFAULT` stays
`K4V2L16`. Nothing in this family has touched a GPU — the kernel has never been
compiled at any K — and the precedent is specific: the last KV-storage change
that shipped default-on with only CPU validation killed every request on the
first V4 forward that met a real device. Its arithmetic was proven and its
device-time cost was not. Ship the path, leave the default alone.
Guards shown red (S1-S7). THREE HOLES, all mine:
* S2/S5 sit behind `#[cfg(feature = "cuda")]`, which no test lane compiles,
so both mutations passed. Now pinned over the SOURCE — a guard that only
exists on hardware nobody has is not a guard.
* S7: the env test reimplemented the parse rule inline and asserted the
replica, so the real reader could silently resolve an unsupported K to the
default and stay green. `from_env` now delegates to a pure `from_spec` and
the test calls that.
I also deleted my own five new tests mid-edit — a second replacement span
swallowed what the first had just inserted, leaving two dangling doc citations
and a suite that looked green. Restored and verified present by name.
And the source guard fired on itself: `include_str!("mod.rs")` includes the
test module, so a negative assertion naming the string it forbids matched its
own text. Scoped to the serving half. That is the second time this class bit
me here; the note is in the code.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SpVNMpb13HkUXqSqbN1o9H
The K=9 CPU decode sized its rows at the data length. Rows now carry tail padding so the kernel's compile-time-width extraction needs no clamp, and that padding is part of what the artifact stores — so a decoder that assumes the data length disagrees with the loader about where row N starts. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SpVNMpb13HkUXqSqbN1o9H
`include_str!("mod.rs")` yields CRLF on a Windows checkout, so
`split_once("#[cfg(test)]\nmod tests {")` never matched and the guard hit
its `expect` instead of asserting anything — the guard was not failing,
it was not running. Normalise line endings before the split; every
assertion after it already goes through the whitespace-squeezer.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
260f0b2 to
90dd68d
Compare
Stage 3 of 3. Stacked on #170, which is stacked on #168 — merge in that order (D20). Base is
feat/qtip-k8v4l12-format.A K=9/V=4/L=12 layer is now served rather than refused, and there is a way to ask for the rung. K=8 gets nothing beyond staying the compiled control — it is quality-closed (−0.00698, −0.00307 even with a converged trellis-Lloyd codebook, against a ±0.0008 band) and serving it would be wasted work. K=9 is the target: Δw_cos +0.00402, 5× better than the shipped control, same 32,768 B table, same L=12/V=4.
Wired
trellis_v4l12::Rung— which already owns the extraction arithmetic and is bit-exactly gated against the GPU kernel. Deliberately not restated: a format with two implementations has none.dequantize_weights/dequantize_w/forward(2-D).forward_v4l12_gemv_cuda— the family's single GPU path, the single-token fused decode+gemv the kernel exists for. Multi-token falls back to the CPU decode, not to a K=4 kernel: slow and right beats fast and wrong.Still refused, loudly
Every path with no V=4 kernel: 3-D
gather_forward,dequantize_expert, 3-Ddequantize_weights, andqtip_packed(whoseQtipPackedViewhas no geometry field, so a consumer would read K=9 bytes as K=4 nibbles). A half-wired dispatch that falls through is worse than one that refuses — at 2 bpw a K=4 misread does not fault, it serves wrong weights.the_k9_path_is_not_the_k4_decoder_in_disguisedecodes the same bytes both ways and requires the answers to differ, so "it returned numbers" cannot pass for correctness.Selection, and default OFF
ARC_QTIP_GEOMETRY=k4v2l16|k<K>v4l12, mirroringARC_QTIP_CODEBOOK. Unknown values refused, never resolved.QtipGeometry::DEFAULTstaysK4V2L16. Nothing in this family has touched a GPU — the kernel has never been compiled at any K. The precedent is specific: the last KV-storage change that shipped default-on with only CPU validation killed every request on the first V4 forward that met a real device. Its arithmetic was proven and its device-time cost was not.Guards shown red (S1–S7) — three holes, all mine
#[cfg(feature = "cuda")], which no test lane compiles, so both mutations passed silently. Now pinned over the source — that a family layer never reaches the K=4 CUDA path, and that the hoist staysOff. A guard that only exists on hardware nobody has is not a guard.from_envnow delegates to a purefrom_spec, and the test calls that.Also: the new source guard fired on itself.
include_str!("mod.rs")includes the test module, so a negative assertion naming the string it forbids matched its own text. Scoped to the serving half. Second time this class bit me here (the first was a barrier guard counting__syncthreadsin a comment); the note is in the code.Scoped clippy clean;
cargo test -p mistralrs-quant343 + 5 passing;mistralrs-coreandarc-enginecompile.What remains
nvcclanes remain queued.