perf(ArcKernels/ArcAttention/ArcQuant): session-10 measured default flips — GEMV_WIDE ON, QK_FUSED_COHORT ON, CUBLAS_MIN_M 512→64 (stacked: B=256 575→758 tok/s, b=1 40.8→43.8) - #223
Merged
Conversation
…43.80 vs 40.76 tok/s) Session-10 hardware A/B, box arc-s10-ledger (H200), candidate 4d03b9e, leg 3: 3x512 produced tokens, seeded t=0.7 p=0.95, engagement proven by [arc-fp8-dispatch] path=gemv_wide; canary1 byte-identical, canary2 coherent and arithmetically correct (numerics within the stated f32 re-association bound). The ledger's ~1.6x prediction did NOT materialise end-to-end; the default stands on the measured +7.5%. Polarity flips from 'only "1" enables' to 'only "0" disables' — still value-read, never presence-read (#212). Test renamed accordingly (only_literal_zero_disables) with the assertion values swapped. Registry entry added (capability_reachability.rs, ArcKernels). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WK8VgocBUrr5djjE1ZfCNf
…entical at B=256 (599.35 vs 575.12 tok/s) Session-10 hardware A/B, box arc-s10-ledger (H200), candidate 4d03b9e, leg 4: fused qk_norm_rope ENGAGED at B=256 (the default arm DECLINED and took the eager chain), seeded canaries byte-identical, 0/512 request errors, aggregate 599.35 vs 575.12 tok/s. The +4.2% is inside the ~±5% inter-leg variance measured the same session (leg 4b no-op delta), so the flip stands on engaged + bit-identical + not-slower, exactly the bar the capability entry set. Doc comment and registry entry updated as the old comment promised; decline message now names the kill switch (ARC_QK_FUSED_COHORT=0). Polarity stays value-read (#212). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WK8VgocBUrr5djjE1ZfCNf
…-arm sweep this constant's own doc demanded Session-10, box arc-s10-ledger (H200), candidate 4d03b9e, arc-tools/fp8_threshold_sweep.sh, LADDER {8,64,256}, one binary, flatten fix present, >=1000-token floor. Clean rows (tok/s): M=64: cublaslt 552.4 | wmma 328.2 | tiled 223.1 -> cuBLASLt wins M=256: cublaslt 715.9 | wmma 620.4 | tiled 285.8 -> cuBLASLt wins First clean M where dequant+cuBLASLt beats min(WMMA, tiled): 64. M=8 rows were VOID on the token floor, so 5..63 stays with the native arms — a void row makes no claim. fp8_gemv_warp still owns M<=4 (the M=1 floor row, cuBLASLt 0.72x, still stands). Rows quoted in the doc per its own rule: a threshold is a claim about every value it excludes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WK8VgocBUrr5djjE1ZfCNf
Code Metrics Report━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Language Files Lines Code Comments Blanks ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ C Header 5 305 210 52 43 CSS 2 1181 1036 34 111 CUDA 81 29475 20161 6275 3039 Dockerfile 1 39 22 8 9 JavaScript 16 3546 2676 482 388 Jinja2 7 694 656 5 33 JSON 75 4896 4893 0 3 Makefile 1 6 5 0 1 Metal Shading Lan| 33 12224 9431 1142 1651 PowerShell 1 300 227 30 43 Python 147 15371 12677 824 1870 Shell 42 10304 6858 2769 677 Plain Text 4 3801 0 2479 1322 TOML 33 1498 1294 54 150 YAML 3 25 23 2 0 ───────────────────────────────────────────────────────────────────────────────── HTML 4 2687 2604 43 40 |- CSS 2 543 479 37 27 |- JavaScript 1 1233 1215 12 6 (Total) 4463 4298 92 73 ───────────────────────────────────────────────────────────────────────────────── Jupyter Notebooks 4 122 83 23 16 |- Markdown 1 60 30 22 8 |- Python 1 122 113 1 8 (Total) 304 226 46 32 ───────────────────────────────────────────────────────────────────────────────── Markdown 213 48932 0 38081 10851 |- BASH 72 1655 1203 331 121 |- C 4 19 19 0 0 |- CUDA 2 84 56 16 12 |- JSON 19 779 779 0 0 |- PowerShell 1 1 1 0 0 |- Python 23 1008 787 113 108 |- Rust 68 2063 1727 78 258 |- TOML 6 207 164 0 43 |- YAML 5 41 36 5 0 (Total) 54789 4772 38624 11393 ───────────────────────────────────────────────────────────────────────────────── Rust 681 340988 293252 18096 29640 |- Markdown 504 32562 471 28063 4028 (Total) 373550 293723 46159 33668 ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Total 1353 516771 363188 99077 54506 ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ |
…red: engaged, byte-identical, not slower Session-10, box arc-s10-ledger (H200), candidate 4d03b9e. Leg 10: seeded canaries byte-identical to the BF16 arm, b=1 41.49 vs 40.76 tok/s. The lane has no engagement log line (silent-success class, noted for follow-up), so engagement was proven by nsys kernel capture instead: arc_kv_fp8_quantize_kernel x16,297 + arc_kv_fp8_dequantize_kernel x32,594 in one b=1 run. 4-gate stacked verification (this + the three earlier flips, one binary): B=256 744.7 tok/s (vs 757.5 three-gate, inside the ±5% band) with canaries IDENTICAL through the stack; b=1 44.68 tok/s — the session's best b=1. Unlike wave43-BU (same default, shipped unrun, broke every forward), this default stands on a measurement. ARC_V4_FP8_KV=0 restores BF16; the ARC_V4_CAPTURE_PROBE veto is unconditional and now tested for the default arm too (write_kv_inplace has no U8 variant). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WK8VgocBUrr5djjE1ZfCNf
Contributor
Author
|
Fourth flip added: 450208d — ARC_V4_FP8_KV default ON. Engagement was proven by nsys kernel capture (arc_kv_fp8_quantize_kernel x16,297 + arc_kv_fp8_dequantize_kernel x32,594 in one b=1 run) since the lane has no engagement log line (silent-success class; follow-up named). 4-gate stacked re-verification, one binary @4d03b9e25: B=256 744.7 tok/s (2 reps, within the ±5% band of the 3-gate 757.5), b=1 44.68 tok/s (session best), canaries IDENTICAL through the full stack, 0 errors. Ladder totals vs pre-wave 5516036: B=256 355.4 → 744.7–757.5; b=1 39.3 → 44.7. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Parent systems: ArcKernels, ArcQuant, ArcInfer/ArcAttention. Session-10 box run, box
arc-s10-ledger(H200, exclusive, PTX-JIT gate passed, provenance asserted per leg), candidate4d03b9e25= release/openrouter-ready after the six wave-1 lanes. One binary, env-toggle A/Bs, seeded canaries (greedy banned, D4), produced-token rates with a >=256/seq floor.The three flips, each measured this session
ARC_FP8_GEMV_WIDEdefault ONpath=gemv_wide, canary coherent (f32 re-association bound). The ledger's ~1.6x prediction did NOT materialise end-to-end.ARC_QK_FUSED_COHORTdefault ONARC_FP8_CUBLAS_MIN_M512 → 64fp8_threshold_sweep.sh, M∈{8,64,256}, >=1000-token floor). Clean rows: M=64 cublaslt 552.4 vs wmma 328.2 vs tiled 223.1; M=256 715.9 vs 620.4 vs 285.8. First clean crossover: 64. M=8 rows VOID on the floor ⇒ 5..63 stays native (a void row makes no claim). M<=4 staysfp8_gemv_warp.Stacked verification (all three ON together, same binary, env-equivalent to these defaults)
Also decided by the same ladder (no code change here)
ARC_NO_FP8_WMMA=1control: tiled is 46% slower at B=256 with an identical canary — the never-run WMMA kernel is confirmed as the correct default arm (its "silently wrong GEMM" fear is REFUTED on hardware).ARC_SAMPLE_ON_DEVICE(sample share 7.4%→57% of step — prediction inverted, measured loudly), TCFRAG (init OOM beside a loaded V4), capture/device-loop (verify_failed=trueat replay perf(stage2): absorbed-MLA V4 decode + GPU radix top-k/top-p sampler #10 — the +1-step divergence is still open): all measured FAIL/unengaged, all stay OFF, evidence in the session results.ARC_V4_FP8_KV: canaries byte-identical and not slower, but the lane has NO engagement instrument at all (silent-success class) — no flip on unproven engagement; needs an engagement line first.Verification
arc-s10-ledger, exclusivity-bracketed (compute-apps == server pid before AND after each leg), running-revision assertion per server, PTX-JIT gate (kernel returned 42 through a JIT-only path) at bootstrap.cargo check -p mistralrs-quant -p mistralrs-coreclean; polarity tests updated both ways (only_literal_zero_disables,only_literal_zero_disables_the_cohort,the_batched_cohort_is_on_by_default_since_the_s10_measurement);capability_reachability11/11 with both registry entries updated (QK cohort entry text now records the measurement; new ArcKernels entry for gemv_wide).🤖 Generated with Claude Code
https://claude.ai/code/session_01WK8VgocBUrr5djjE1ZfCNf