Skip to content

perf(ArcKernels/ArcAttention/ArcQuant): session-10 measured default flips — GEMV_WIDE ON, QK_FUSED_COHORT ON, CUBLAS_MIN_M 512→64 (stacked: B=256 575→758 tok/s, b=1 40.8→43.8) - #223

Merged
heydryft merged 4 commits into
release/openrouter-readyfrom
arcbox/s10-default-flips
Aug 22, 2026

Conversation

@heydryft

Copy link
Copy Markdown
Contributor

Parent systems: ArcKernels, ArcQuant, ArcInfer/ArcAttention. Session-10 box run, box arc-s10-ledger (H200, exclusive, PTX-JIT gate passed, provenance asserted per leg), candidate 4d03b9e25 = release/openrouter-ready after the six wave-1 lanes. One binary, env-toggle A/Bs, seeded canaries (greedy banned, D4), produced-token rates with a >=256/seq floor.

The three flips, each measured this session

commit flag evidence
8c81489 ARC_FP8_GEMV_WIDE default ON b=1 43.80 vs 40.76 tok/s (+7.5%), engagement path=gemv_wide, canary coherent (f32 re-association bound). The ledger's ~1.6x prediction did NOT materialise end-to-end.
0ef138b ARC_QK_FUSED_COHORT default ON ENGAGED at B=256 (default arm DECLINED to eager), seeded canaries byte-identical, 599.35 vs 575.12 tok/s. +4.2% is inside the measured ~±5% inter-leg variance — the flip stands on engaged + bit-identical + not-slower, the exact bar the capability entry set.
e570d30 ARC_FP8_CUBLAS_MIN_M 512 → 64 The three-arm sweep this constant's own doc demanded (fp8_threshold_sweep.sh, M∈{8,64,256}, >=1000-token floor). Clean rows: M=64 cublaslt 552.4 vs wmma 328.2 vs tiled 223.1; M=256 715.9 vs 620.4 vs 285.8. First clean crossover: 64. M=8 rows VOID on the floor ⇒ 5..63 stays native (a void row makes no claim). M<=4 stays fp8_gemv_warp.

Stacked verification (all three ON together, same binary, env-equivalent to these defaults)

  • B=256 decode aggregate 757.54 tok/s (5 reps 748.8–761.4, effective_B=256 verified, 0/1280 errors), per-user p50 3.35 tok/s, TTFT p50 4.87 s, $1.80/Mtok @ $4.92/hr.
  • b=1 steady state 43.77 tok/s (5 runs, spread 0.04).
  • Ladder context: pre-wave (5516036) B=256 was 355.4; candidate baseline 575.1; stacked 757.5 — 2.13× in one session, with the wave-1 merges and these flips both contributing.

Also decided by the same ladder (no code change here)

  • ARC_NO_FP8_WMMA=1 control: tiled is 46% slower at B=256 with an identical canary — the never-run WMMA kernel is confirmed as the correct default arm (its "silently wrong GEMM" fear is REFUTED on hardware).
  • Ragged pair, ARC_SAMPLE_ON_DEVICE (sample share 7.4%→57% of step — prediction inverted, measured loudly), TCFRAG (init OOM beside a loaded V4), capture/device-loop (verify_failed=true at replay perf(stage2): absorbed-MLA V4 decode + GPU radix top-k/top-p sampler #10 — the +1-step divergence is still open): all measured FAIL/unengaged, all stay OFF, evidence in the session results.
  • ARC_V4_FP8_KV: canaries byte-identical and not slower, but the lane has NO engagement instrument at all (silent-success class) — no flip on unproven engagement; needs an engagement line first.

Verification

  • Box: every number above from arc-s10-ledger, exclusivity-bracketed (compute-apps == server pid before AND after each leg), running-revision assertion per server, PTX-JIT gate (kernel returned 42 through a JIT-only path) at bootstrap.
  • Local: cargo check -p mistralrs-quant -p mistralrs-core clean; polarity tests updated both ways (only_literal_zero_disables, only_literal_zero_disables_the_cohort, the_batched_cohort_is_on_by_default_since_the_s10_measurement); capability_reachability 11/11 with both registry entries updated (QK cohort entry text now records the measurement; new ArcKernels entry for gemv_wide).

🤖 Generated with Claude Code

https://claude.ai/code/session_01WK8VgocBUrr5djjE1ZfCNf

heydryft and others added 3 commits August 22, 2026 01:41
…43.80 vs 40.76 tok/s)

Session-10 hardware A/B, box arc-s10-ledger (H200), candidate 4d03b9e,
leg 3: 3x512 produced tokens, seeded t=0.7 p=0.95, engagement proven by
[arc-fp8-dispatch] path=gemv_wide; canary1 byte-identical, canary2 coherent
and arithmetically correct (numerics within the stated f32 re-association
bound). The ledger's ~1.6x prediction did NOT materialise end-to-end; the
default stands on the measured +7.5%.

Polarity flips from 'only "1" enables' to 'only "0" disables' — still
value-read, never presence-read (#212). Test renamed accordingly
(only_literal_zero_disables) with the assertion values swapped. Registry
entry added (capability_reachability.rs, ArcKernels).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WK8VgocBUrr5djjE1ZfCNf
…entical at B=256 (599.35 vs 575.12 tok/s)

Session-10 hardware A/B, box arc-s10-ledger (H200), candidate 4d03b9e,
leg 4: fused qk_norm_rope ENGAGED at B=256 (the default arm DECLINED and
took the eager chain), seeded canaries byte-identical, 0/512 request
errors, aggregate 599.35 vs 575.12 tok/s. The +4.2% is inside the ~±5%
inter-leg variance measured the same session (leg 4b no-op delta), so the
flip stands on engaged + bit-identical + not-slower, exactly the bar the
capability entry set. Doc comment and registry entry updated as the old
comment promised; decline message now names the kill switch
(ARC_QK_FUSED_COHORT=0). Polarity stays value-read (#212).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WK8VgocBUrr5djjE1ZfCNf
…-arm sweep this constant's own doc demanded

Session-10, box arc-s10-ledger (H200), candidate 4d03b9e,
arc-tools/fp8_threshold_sweep.sh, LADDER {8,64,256}, one binary, flatten
fix present, >=1000-token floor. Clean rows (tok/s):

  M=64:  cublaslt 552.4 | wmma 328.2 | tiled 223.1  -> cuBLASLt wins
  M=256: cublaslt 715.9 | wmma 620.4 | tiled 285.8  -> cuBLASLt wins

First clean M where dequant+cuBLASLt beats min(WMMA, tiled): 64. M=8 rows
were VOID on the token floor, so 5..63 stays with the native arms — a void
row makes no claim. fp8_gemv_warp still owns M<=4 (the M=1 floor row,
cuBLASLt 0.72x, still stands). Rows quoted in the doc per its own rule:
a threshold is a claim about every value it excludes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WK8VgocBUrr5djjE1ZfCNf
@github-actions

Copy link
Copy Markdown
Code Metrics Report
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
 Language              Files        Lines         Code     Comments       Blanks
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
 C Header                  5          305          210           52           43
 CSS                       2         1181         1036           34          111
 CUDA                     81        29475        20161         6275         3039
 Dockerfile                1           39           22            8            9
 JavaScript               16         3546         2676          482          388
 Jinja2                    7          694          656            5           33
 JSON                     75         4896         4893            0            3
 Makefile                  1            6            5            0            1
 Metal Shading Lan|       33        12224         9431         1142         1651
 PowerShell                1          300          227           30           43
 Python                  147        15371        12677          824         1870
 Shell                    42        10304         6858         2769          677
 Plain Text                4         3801            0         2479         1322
 TOML                     33         1498         1294           54          150
 YAML                      3           25           23            2            0
─────────────────────────────────────────────────────────────────────────────────
 HTML                      4         2687         2604           43           40
 |- CSS                    2          543          479           37           27
 |- JavaScript             1         1233         1215           12            6
 (Total)                             4463         4298           92           73
─────────────────────────────────────────────────────────────────────────────────
 Jupyter Notebooks         4          122           83           23           16
 |- Markdown               1           60           30           22            8
 |- Python                 1          122          113            1            8
 (Total)                              304          226           46           32
─────────────────────────────────────────────────────────────────────────────────
 Markdown                213        48932            0        38081        10851
 |- BASH                  72         1655         1203          331          121
 |- C                      4           19           19            0            0
 |- CUDA                   2           84           56           16           12
 |- JSON                  19          779          779            0            0
 |- PowerShell             1            1            1            0            0
 |- Python                23         1008          787          113          108
 |- Rust                  68         2063         1727           78          258
 |- TOML                   6          207          164            0           43
 |- YAML                   5           41           36            5            0
 (Total)                            54789         4772        38624        11393
─────────────────────────────────────────────────────────────────────────────────
 Rust                    681       340988       293252        18096        29640
 |- Markdown             504        32562          471        28063         4028
 (Total)                           373550       293723        46159        33668
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
 Total                  1353       516771       363188        99077        54506
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

…red: engaged, byte-identical, not slower

Session-10, box arc-s10-ledger (H200), candidate 4d03b9e. Leg 10: seeded
canaries byte-identical to the BF16 arm, b=1 41.49 vs 40.76 tok/s. The lane
has no engagement log line (silent-success class, noted for follow-up), so
engagement was proven by nsys kernel capture instead:
arc_kv_fp8_quantize_kernel x16,297 + arc_kv_fp8_dequantize_kernel x32,594
in one b=1 run. 4-gate stacked verification (this + the three earlier
flips, one binary): B=256 744.7 tok/s (vs 757.5 three-gate, inside the
±5% band) with canaries IDENTICAL through the stack; b=1 44.68 tok/s —
the session's best b=1.

Unlike wave43-BU (same default, shipped unrun, broke every forward), this
default stands on a measurement. ARC_V4_FP8_KV=0 restores BF16; the
ARC_V4_CAPTURE_PROBE veto is unconditional and now tested for the default
arm too (write_kv_inplace has no U8 variant).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WK8VgocBUrr5djjE1ZfCNf
@heydryft

Copy link
Copy Markdown
Contributor Author

Fourth flip added: 450208d — ARC_V4_FP8_KV default ON. Engagement was proven by nsys kernel capture (arc_kv_fp8_quantize_kernel x16,297 + arc_kv_fp8_dequantize_kernel x32,594 in one b=1 run) since the lane has no engagement log line (silent-success class; follow-up named). 4-gate stacked re-verification, one binary @4d03b9e25: B=256 744.7 tok/s (2 reps, within the ±5% band of the 3-gate 757.5), b=1 44.68 tok/s (session best), canaries IDENTICAL through the full stack, 0 errors. Ladder totals vs pre-wave 5516036: B=256 355.4 → 744.7–757.5; b=1 39.3 → 44.7.

@heydryft
heydryft merged commit 4033b8f into release/openrouter-ready Aug 22, 2026
18 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant