Skip to content

fix(ArcLab/ArcQuant): make prefill measurable, then ship the two kernels it proves were switched off - #196

Merged
heydryft merged 3 commits into
release/openrouter-readyfrom
agent/prefill-honest-bench-v2
Aug 21, 2026
Merged

heydryft merged 3 commits into
release/openrouter-readyfrom
agent/prefill-honest-bench-v2

Conversation

@heydryft

Copy link
Copy Markdown
Contributor

Prefill was unmeasurable: pp N printed 0.000±0.000 into a results table and exited 0. This makes it measurable, then uses the measurements to switch on two things that were already built, already correct, and shipping disabled.

Three premises this PR retracts, with evidence

1. "TTFT ≈ 24 s / 11.8 ms per prompt token" does not reproduce. Measured at d7742670a before any change here: TTFT 3.675 s, 1.780 ms/prompt-token (pp2048, b=1, H200, --no-paged-attn). The 24 s figure is stale by many commits.

2. "71.3% of the prefill step is the QTIP MoE expert gather" is no longer true. nsys (--cuda-graph-trace=node, collection started at process start) at pp2048 with the grouped GEMM engaged:

kernel share instances avg
flash_attn_sinks_kernel 60.3% 86 21.8 ms (max 54.0)
qtip2b_grouped_gemm_kernel 17.1% 129 4.11 ms
ucopy_bf16 10.0% 2527 —

GPU-busy is 99.4% (3.106 s kernel / 3.126 s wall). Attention is the prefill bottleneck now, not the MoE gather.

3. The reported rc=139 SIGSEGV did not reproduce in ~26 process exits — pre-patch binary, post-patch, under nsys, and the CLI mistralrs bench path, all rc=0. The teardown fix here is implemented from the existing static diagnosis in memory/mission/BACKLOG.md, but it is not claimed to have cured anything, because the symptom never appeared.

The 0.000, root-caused and proven

A prefill-only request (max_len = 1) finishes inside step(): pipeline/sampling.rs calls update_time_info(), builds group.get_usage() and dispatches Response::Done there. The default scheduler arm stamped prompt_timestamp/total_prompt_time only after step() returned, so usage was built with prompt_timestamp == None, update_time_info skipped total_prompt_time, and get_usage took its == 0 branch. The PagedAttention arm already stamps before step() with a comment explaining exactly this — it was never applied to the arm that serves (PagedAttention is banned here: it shadows the graph arm).

Proof, same box, same model, same prompt, one commit apart:

pre-patch : | pp 2048 | 0.000±0.000 |  inf±NaN | 1 | 0 |          exit 0
patched   : | pp 2048 | 561.711     | 1.780    | 1 | 561.7 | ...  exit 0

The harness now exits 2 (environment failure, never 1) on any zero-token / zero-rate / zero-time row and prints no table; carries a second engine-independent wall clock (agrees with the engine to 0.8%); stops reusing one request id across every concurrent copy and repetition; and prints the grouped-GEMM launch counters, which nothing in the workspace read — so every prior kernel comparison this harness could produce was unproven by construction.

Two things that were built, correct, and switched off

Tuned grouped-GEMM kernel was unreachable. grouped_variant() fell back to QTIP_GROUPED_VARIANT_BASELINE with no env var set, and the only callers of set_grouped_variant in the workspace are in examples/qtip_grouped_curve.rs. Every production prefill ran variant 0.

variant 0 (was default): 561.7 prompt tok/s | 1.780 ms/tok | TTFT 3.675 s
variant 2 (now default): 671.3 prompt tok/s | 1.490 ms/tok | TTFT 3.080 s   => +19.5%

The tile-fill gate refused the grouped GEMM below 683 tokens. Both arms pinned to variant 2 so the only difference is which kernel runs, each arm proved from runtime launch counters ([0,0,0] GEMV vs [0,0,129] = 43 layers × 3 expert matrices):

N GEMV grouped gain m-tile fill
16 97.6 113.5 +16% 7.5%
32 110.7 141.0 +27% 8.8%
64 145.1 217.7 +50% 12%
128 162.6 289.6 +78% 20%
512 246.0 691.0 +181% 75%

The gate's justifying A/B — "1.00× at N=128 and N=512, 2.41× at N=1024" — does not reproduce. Grouped wins at every length measured, the margin grows monotonically, and it never inverts, including at 7.5% tile fill, two orders of magnitude below the amortization the gate was protecting. A gate defended by a measurement nobody can reproduce is worse than no gate.

Tile fill cannot be the deciding quantity, and the mechanism says why: the grouped kernel stages each woken expert's packed bytes once per m-tile instead of once per (token, expert) pair, and computes on tensor cores where the GEMV is scalar. Padding an under-full tile is cheap against both; fill only bounds the wasted mma, which is the smaller term.

DECODE_REGIME_MAX_TOKENS stays as a floor (the RUN-161 graph-capture safety rule, not a performance claim), and ARC_QTIP_ONDEVICE_MOE_MAX_TOKENS still pins the GEMV arm so the A/B stays runnable. Scope is the qtip2b rung only — the LUT rung in mod.rs keeps its own boundary, unmeasured here.

A peer measured forced-grouped in decode and saw 0.47× at B=1. Consistent with this: the only genuinely bad point is n=1. The threshold should be ~16, not 683 and not 128.

Net effect (H200, b=1, --no-paged-attn, shipped defaults, zero env vars)

pp before before TTFT after after TTFT gain
128 162.6 tok/s 0.815 s 289.6 0.471 s 1.78×
512 246.0 tok/s 2.110 s 946.4 0.570 s 3.85×
2048 561.7 tok/s 3.675 s 670.8 3.082 s 1.19×
8192 226.9 tok/s 36.13 s — — attention-bound

b=8 aggregate (pre-change): 1080 / 990 / 588 / 224 prompt tok/s at 128 / 512 / 2048 / 8192. Prefill batches enormously at short prompts (128: 1080 vs 163 = 6.6×) and barely at long ones (2048: +5%), and the engagement log shows why — eight 128-token prompts land in one 1024-token step, which crossed the old kernel-switch boundary. The dead zone was per-step tokens, not per-request tokens.

Follow-ups found, not fixed here

  • mistralrs-cli/src/commands/bench.rs:136 still rates tokens requested (prompt_len / elapsed), so the CLI can print a high number for a zero-token run. The fix/bench-counts-produced-tokens fix never reached it.
  • mistralrs-quant/src/qtip/mod.rs:3653 (LUT rung) carries the same tile-fill gate retracted here, unmeasured — flagged deliberately, not changed blind.
  • flash_attn_sinks_kernel at 60.3% with no prefill.cuh in the tree is now the largest single item in prefill.

🤖 Generated with Claude Code

@github-actions

github-actions Bot commented Aug 20, 2026 •

Copy link
Copy Markdown
Code Metrics Report
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
 Language              Files        Lines         Code     Comments       Blanks
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
 C Header                  5          305          210           52           43
 CSS                       2         1181         1036           34          111
 CUDA                     81        29475        20161         6275         3039
 Dockerfile                1           39           22            8            9
 JavaScript               16         3546         2676          482          388
 Jinja2                    7          694          656            5           33
 JSON                     75         4896         4893            0            3
 Makefile                  1            6            5            0            1
 Metal Shading Lan|       33        12224         9431         1142         1651
 PowerShell                1          300          227           30           43
 Python                  147        15371        12677          824         1870
 Shell                    42        10304         6858         2769          677
 Plain Text                4         3801            0         2479         1322
 TOML                     33         1498         1294           54          150
 YAML                      3           25           23            2            0
─────────────────────────────────────────────────────────────────────────────────
 HTML                      4         2687         2604           43           40
 |- CSS                    2          543          479           37           27
 |- JavaScript             1         1233         1215           12            6
 (Total)                             4463         4298           92           73
─────────────────────────────────────────────────────────────────────────────────
 Jupyter Notebooks         4          122           83           23           16
 |- Markdown               1           60           30           22            8
 |- Python                 1          122          113            1            8
 (Total)                              304          226           46           32
─────────────────────────────────────────────────────────────────────────────────
 Markdown                213        48932            0        38081        10851
 |- BASH                  72         1655         1203          331          121
 |- C                      4           19           19            0            0
 |- CUDA                   2           84           56           16           12
 |- JSON                  19          779          779            0            0
 |- PowerShell             1            1            1            0            0
 |- Python                23         1008          787          113          108
 |- Rust                  68         2063         1727           78          258
 |- TOML                   6          207          164            0           43
 |- YAML                   5           41           36            5            0
 (Total)                            54789         4772        38624        11393
─────────────────────────────────────────────────────────────────────────────────
 Rust                    681       340988       293252        18096        29640
 |- Markdown             504        32562          471        28063         4028
 (Total)                           373550       293723        46159        33668
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
 Total                  1353       516771       363188        99077        54506
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

@heydryft
heydryft force-pushed the agent/prefill-honest-bench-v2 branch from af638ca to cd34ec4 Compare August 20, 2026 23:13
@heydryft

Copy link
Copy Markdown
Contributor Author

Added: the CLI bench could still fabricate a number (cd34ec4)

mistralrs bench computed prompt_len / elapsed and gen_len / elapsed — both numerators are what the caller asked for. A run that produced nothing still printed a confident throughput number, and the faster it failed the better it scored.

Structural cause: run_single_bench returned Result<()>, discarding the engine's Usage, so the call sites had nothing to divide by except the request. It now returns the Usage and both legs rate usage.prompt_tokens / usage.completion_tokens. Since that alone would convert a fabricated number into a silent zero, assert_produced_tokens exits 2 on zero and warns (without failing) when the count differs from the request — decode legitimately stops early on EOS.

Verified on the H200, patched build, artifact assertion 1 (pre-patch binary: 0):

Prefill (128 tokens)  324.3 ± 0.0   394.73 ms (TTFT)     CLI_EXIT=0

against 173.3 ± 0.0 / 738.77 ms from the same CLI binary before the two quant defaults in this branch landed. Caveat: the exit-2 path is not exercised live — I could not manufacture a zero-token run on working hardware. The guard is verified only in the negative direction (it does not fire on a good run).

Also flagged, deliberately not fixed: the LUT rung's copy of the retracted gate

mistralrs-quant/src/qtip/mod.rs carries the same grouped_gemm_tiles_amortize predicate this PR retracts for qtip2b. It is left in place with a comment recording why: the LUT rung's fallback is the host-synced per-expert dequantize loop, not the fused gather GEMV — different traffic, different economics — and it was not measured. Changing it on a sibling's numbers would repeat exactly the unreproducible-measurement failure this PR is about. The comment records the A/B that would settle it.

Correction to an earlier finding of mine — "prefix caching is disabled by default" is not a config bug

I previously relayed cache_engine.rs disabling prefix caching under TurboQuant as a one-line config fix. That is wrong and it should not be taken as one. The disable is a correctness guard:

TurboQuant blocks are packed 4-bit K / 3-bit V codebook indices with per-token norms held in separate tensors (turbo_paged_attention.cuh:15-19); gather_kv_cache has no dequantize path and no norms argument, so a prefix-cache hit cannot be served.

The conflict is already detected (PagedCacheType::prefix_cache_conflict), already surfaced naming both flags (engine/mod.rs:256-266), and already unit-tested. Flipping it on would serve garbage from packed blocks rather than restore caching.

The premise underneath is correct and worth acting on — the default command line does land in the conflict (--pa-cache-type unset ⇒ TurboQuant, PagedAttention is the CUDA default, --prefix-cache-n defaults to 16), so default serving runs with no prefix cache. But the fix is a kernel change: give gather_kv_cache a TurboQuant dequantize path plus a norms argument. That is a real and valuable piece of work, and it is not one line.

@heydryft

Copy link
Copy Markdown
Contributor Author

🔴 STOP-EVERYTHING: the fused head_dim=512 flash-sinks path computes something different from the reference, and it is ON by default

Found while scoping the attention port. This is a correctness bug in V4 serving today, not a performance item, and it outranks everything else in this PR.

The mechanism (static, confirmed by reading)

mistralrs-core/src/attention/backends/sinks.rs:161-170 — the CUDA arm — calls:

return mistralrs_paged_attn::flash_attn_sinks(q, k, v, Some(sinks), sdpa_params.softmax_scale, window_size);

There is no mask argument. V4 builds a dense [t_q, n_keys] additive 0/−inf mask at models/dsv4_attention.rs:806-826 over the concatenation [raw sliding-window KV ++ compressed KV], and it is then dropped on the floor. sdpa_sliding_window returns None for CSA/HCA (dsv4_attention.rs:426), so window_size = 0 and the kernel falls back to a full causal scan (flash_attn_sinks.cu:167-172) over an axis where relative distance is meaningless — compressed entry j stands for absolute position j * ratio.

The codebase predicted this exactly, at dsv4_attention.rs:421-425:

"it is a landmine that arms itself the moment a 512-wide flash-sinks kernel lands, and silently, because the fused path takes no mask to disagree with."

That kernel has landed (flash_attn_sinks.cu:600 takes head_dim=512) and flash_512_enabled() (sinks.rs:71) defaults ON. My nsys profile in this PR shows flash_attn_sinks_kernel is 60.3% of all prefill GPU time, so the armed path is the one production runs.

The measurement (H200, aeonmind/DeepSeek-V4-Flash-UQFF-qtip2b)

Same binary, same prompt, temperature: 0.0, /v1/completions, one environment variable apart. ARC_FLASH_512=0 selects the unfused reference in the same binary, and the arcflash log line confirms which arm ran on each leg.

6,649 prompt tokens:

leg arcflash log output
ARC_FLASH_512=1 (default) sinks path is FUSED 'orem, etc. etc. etc. etc. etc. etc. etc. …'
ARC_FLASH_512=0 (reference) sinks path is UNFUSED '> \n> #+end_src\n> \n> * Footnote\n> \n> Think about how to assign the letters to the numbers.\n> \n> My previous approach was to use the first letter of the english name of each number'

485 prompt tokens — they differ here too, so this is not confined to long context:

leg output
FUSED (default) '…\n\n## 2.2.2.2.2.2.2.2.2.2.2.2.2.2.2.2.2.2.2'
UNFUSED (reference) '. \ncinder . cinder . cinder . cinder . cinder …'

What this does and does not establish

Established: the two paths do not agree, at either length, at temperature 0, with the arm proved from the runtime log on both legs. At 6,649 tokens the reference produces coherent English and the default produces pure repetition.

Not established: that the unfused path is correct. I proved disagreement, not which side is right — though the unfused path is the reference the mask was written for, and it is the one producing coherent text. A F32 logit bit-compare is the next instrument; the text diff is only what I could get in the time I had.

Consequence for this PR's numbers: the prefill throughput measurements are A/B comparisons where both arms ran the same attention path, so the relative results (variant 0 vs 2, GEMV vs grouped) stand. But the absolute rates — the 8192 row especially — are rates for a kernel that is computing the wrong thing, and should not be quoted as product numbers until this is resolved.

Reproduce

mistralrs serve -p 8123 -m <V4_SRC> --from-uqff <UQFF>/qtip2b-0.uqff   # default = FUSED
ARC_FLASH_512=0 mistralrs serve -p 8123 -m <V4_SRC> --from-uqff <UQFF>/qtip2b-0.uqff  # UNFUSED
# same prompt, temperature 0, /v1/completions; grep the server log for "sinks path is"

Suggested immediate action

Default flash_512_enabled() to off until the fused path takes the mask, and give flash_attn_sinks a custom-mask argument (FlashInfer's MaskMode::kCustom, not kCausal — causality here is raw sliding-window ∧ compressed block-causality ∧ caller padding, which a relative-distance window cannot express). I have not made that change in this PR because it is a serving-behaviour change that deserves its own review, not a rider on a benchmark fix.

@heydryft

Copy link
Copy Markdown
Contributor Author

✅ Settled, fixed, and verified — and there was no correctness/speed tradeoff to make

1. Direction: the fused path is the wrong one

Compared by magnitude, not bit-equality — two implementations of the same math never agree bit-for-bit, so bit-inequality proves nothing. (My first attempt used bit-exact F32 and produced an uninformative "63,853/65,536 words differ"; that was my instrument error, corrected.) Control is a 1-ULP perturbation of the OUTPUT; a shared-input perturbation moves both paths together and can never fire.

H200, head_dim=512, output magnitude 0.1875:

comparison max abs diff
fused vs reference (causal only) 0.012848
fused vs reference (masked) 0.774902
ref(causal) vs ref(masked) 0.771973 ← the mask's entire effect

The fused result sits 60.3× closer to the unmasked reference, and its error against the masked reference (0.7749) is the same size as the mask's whole effect (0.7720). The mask is not approximated — it is discarded. The fused path computes plain causal attention; the reference is right.

2. Fixed at the interface, not at the caller

flash_attn_sinks now takes mask so that it cannot be forgotten, and hard-errors when one is supplied. A kernel that cannot honour a mask must refuse work that carries one; the next caller gets a compile error (new argument), then a runtime error — never silently wrong numbers. The dispatch gates the fused arm on mask.is_none() and warns once when it declines.

Three GPU regression gates, all passing:

  • fused_kernel_refuses_a_mask_instead_of_dropping_it
  • fused_sinks_512_ignores_the_mask — documents the defect quantitatively
  • masked_sinks_attention_matches_the_masked_reference — max abs diff 0.00000000

End to end with the fused path enabled (ARC_FLASH_512=1, the production default), 6,649-token prompt: the server logs sinks kernel DECLINED: an explicit attention mask is present, and the output is now byte-identical to the unfused reference. The 'orem, etc. etc. etc.' degeneracy is gone.

3. The default decision makes itself — correct is also faster

pp, b=1 before (wrong answers) after (correct answers) change
128 289.6 tok/s, TTFT 0.471 s 292.2 tok/s, TTFT 0.466 s +0.9%
512 946.4 tok/s, TTFT 0.570 s 1103.4 tok/s, TTFT 0.492 s +16.6%
2048 670.8 tok/s, TTFT 3.082 s 1205.4 tok/s, TTFT 1.728 s +79.7%, TTFT −44%

fused_declined_logged=1 on every slot, so the fallback is proven engaged. The hand-written scalar sinks kernel — 60.3% of prefill — was losing to cuBLAS matmul + softmax. This is the same finding as attention sitting at ~1% of tensor-core peak: it is scalar where it claims to be fused.

No flag flip is needed. The gate is now correctness-driven rather than a default anyone has to remember: whenever a mask is present the fused arm declines, and for V4's CSA/HCA that is always.

4. GSM8K 96.0% is not void

Measured 2026-08-15. The fused 512 path first became reachable 2026-08-17 23:24 (17cecfa62 — the only commit introducing hd == 512 into the fused sinks gate). The eval predates the bug by two days and stands.

⚠️ But every absolute prefill rate taken after 2026-08-17 with default settings is a rate for a kernel computing the wrong thing — including the ones earlier in this PR. The corrected numbers are the table above.

@heydryft

Copy link
Copy Markdown
Contributor Author

The B=256 decode profile — the operating point nobody had ever profiled

H200, DeepSeek-V4-Flash qtip2b, --no-paged-attn, ARC_PROFILE=1, 40 recorded steps after 8 warmup. ARC_TIME_DECODE deliberately unset — it inserts per-layer device.synchronize() calls, which would serialise the pipeline and destroy the exact split being measured. Profiler self-test passed at load: launch_wall=174560 ns, device=11059616 ns, ratio=63.4x — CUDA events are measuring execution, not launches.

Achieved concurrency: geom {b: 256, t: 1} — a genuine 256-wide step, asserted, not assumed. 20,480 tokens over 40 steps = 512 tokens/step = 256 sequences × 2 (MTP depth 1).

1. 🔑 Host vs device — the step is DEVICE-bound

before RoPE fix after RoPE fix
wall/step 1082.3 ms 794.0 ms
device (stream-elapsed) / wall 91.4% 88.8%
sync (host blocked on device) / wall 10.3% 18.6%
aggregate 473.0 tok/s 644.8 tok/s

Two independent instruments agree. nsys kernel-busy over the same window is 16.28 s of kernel time against 18.26 s of decode wall ≈ 89%, against the profiler's 88.8% device/wall. The earlier "host-bound, GPU idle" claim is refuted by measurement — the GPU is busy ~89% of the step, and the host now waits on the GPU (18.6% sync) more than the reverse.

That said, there was a large per-sequence host cost, and it was partly hidden behind the GPU rather than absent. See below.

2. Kernel shares at B=256 decode (--cuda-graph-trace=node)

share ms/step launches/step kernel
66.01% 467.3 288 fp8_gemm::fp8_matmul_tiled<bf16,32,32,32>
8.69% 61.5 129 qtip2b_grouped_gemm_kernel
8.41% 59.5 1 sm90_xmma_gemm (lm_head)
4.32% 30.6 898 ucopy_bf16
2.70% 19.1 12 cutlass bf16 gemm
1.11% 7.8 3,295 copy2d_bf16

⚠️ The fp8-45% / qtip-25% shares do not reproduce. I measure fp8 66.0% / qtip-grouped 8.7%. I also could not find those numbers anywhere in FACTS.md, so I cannot identify their provenance. The only ARC_TIME_DECODE component profile in the tree (FACTS.md:1382) is a different decomposition (mla_attn 49% / mhc 16% / 16% / moe 16%) and its own header says cudnn build — RE-MEASURE on no-cudnn — so it is compromised twice over: sync-serialised and built with cudnn, which is −62% on decode. Please stop quoting the 45/25 split.

fp8_matmul_tiled at 66% is where the next session's money goes.

3. The per-sequence term, located and removed

Same defect as PR #198, found independently here from the profile: layers.rs gated RoPE's batched path on seqlen_offsets.len() == 1 — the length of the vector, where the property needed is the distinctness of its values.

Before the fix, mla_attn.qk_norm_rope.rope alone spent 728.6 ms/step of HOST time against 163.5 ms/step of device time — 67% of a 1082 ms step, 16.9 ms per layer per step. forward_inverse_tail carried the identical bug.

per_sequence == 0 is confirmed, by a different route. I did not need the counter: that the fix produces any speedup at all proves the offsets were uniform. If they were ragged, uniform_seqlen_offset returns None, the loop still runs, and nothing changes. It changed by 1.36×.

Measured, not predicted: 1082.3 → 794.0 ms/step, 473.0 → 644.8 tok/s (1.36×). That exceeds the 19.9% the author scoped; 26.6% of wall came back.

On the launch-count prediction: rope_i_bf16 is now 156/step against 129 predicted — essentially confirmed. uneg_bf16 collapsed to 50/step from the predicted 11,008. But total launches are ~8,916/step, not 301 — the rope family is fixed; the remainder is dominated by copy2d_bf16 (3,295/step) and ucopy_bf16 (898/step), which are data movement and untouched by this fix.

4. MoE arm A/B at B=256 decode

Arm proved from the engagement log (grouped engaged: 1 vs 0):

wall/step aggregate
grouped GEMM (default after the gate retraction) 794.0 ms 644.8 tok/s
GEMV arm pinned (ARC_QTIP_ONDEVICE_MOE_MAX_TOKENS) 1596.0 ms 320.8 tok/s

Grouped is 2.01× at decode B=256. This also reconciles the peer's B=1 result of 0.47×: at B=1, n_tokens = 1 ≤ DECODE_REGIME_MAX_TOKENS, so the GEMV arm runs regardless and the gate retraction does not apply. The two findings do not conflict.

Caveats

  • The nsys trace covers the whole process; with -p 0 the single prefill step is ~1,024 tokens and decode dominates, but the shares carry a small prefill contamination.
  • t(B) = 159 ms + 8.33 ms·B does not describe this binary — it predicts 2,292 ms/step at B=256 and I measure 794 ms. The shape (a per-sequence cost) was right and is now located; the constants are not portable across the fixes in this branch.

@heydryft

Copy link
Copy Markdown
Contributor Author

Rebased onto release/openrouter-ready — 3 of 7 commits survive; here is what happened to the rest

Landing: 368db7275 (prefill measured 0.000 tok/s), 469674424 (grouped-GEMM launch counters), 6e4fc909e (tuned grouped-GEMM as the default, +19.5% prefill).

Dropped as already-applied (identical patch-id already on the integration branch): 145bbbc0a, 11cf8cd38, d5c55fe73, 8dbdd1a87.

Dropped as superseded — resolved toward the newer fix:

  • 500bca01e "retract the tile-fill gate on the qtip2b MoE gather" — conflicted in qtip/bitshift.rs against fix(ArcMoE): the grouped-GEMM crossover was 40x too high — on BOTH rungs #197, already on the integration branch. After resolving to the integration branch's version the commit was empty: same dispatch, same boundary. Its version was kept deliberately, because this branch's side still spelled the disable flag std::env::var("ARC_NO_QTIP_ONDEVICE_MOE").is_ok() — the presence check that fix(ArcGate): read boolean ARC_* flags by value, so =0 means off #212 fixed. Taking this branch's text would have re-armed the =0-means-ON bug six hours after it was closed. The integration branch's crate::env_flag_is_set(...) stands. Its comment already carries this branch's prefill table verbatim (+16/+27/+50/+78/+181%), so no measurement was lost.
  • cd34ec4ef "the CLI bench rated tokens REQUESTED" — the integration branch already carries a strictly stronger form of the same fix: a BenchRun struct that also returns text_len, and rate_from_produced() which returns an error rather than assert_produced_tokens()'s process::exit(2). Residual was a now-dead helper and an unused Usage import.

Dropped as out of scope — NOT superseded, still wanted:

Verified after rebase: the n >= 683 gate is still absent from bitshift.rs and qtip/mod.rs (only retraction comments remain), TCFRAG still opt-in (== Some("1")), env_flag_is_set count unchanged at 39, EXPECTED_KERNEL_COUNT untouched at 44.

Nirupam Bhowmick added 3 commits August 21, 2026 18:09
…rdown

Three defects, all on the path that makes prefill measurable at all.

1. `pp N` reported `0.000±0.000` for every prompt benchmark.

   A prefill-only request (`max_len = 1`) finishes *during* the prompt step:
   `pipeline/sampling.rs` calls `seq.update_time_info()`, builds
   `group.get_usage()` and dispatches `Response::Done` from inside `step()`.
   But the default (non-paged) scheduler arm stamped `prompt_timestamp` and
   `total_prompt_time` only AFTER `step()` returned. So at the moment the usage
   was built `prompt_timestamp` was still `None`, `update_time_info` skipped
   `group.total_prompt_time`, and `get_usage` took its `== 0` branch and
   returned `avg_prompt_tok_per_sec: 0.0`.

   The PagedAttention arm already stamps before `step()`, with a comment saying
   exactly why. The fix was simply never applied to the arm that serves --
   PagedAttention is banned here (it shadows the graph arm and measures zero
   tokens). Stamp on the default arm too.

2. The harness printed that zero into a results table and exited 0.

   A zero is not a measurement. Every row now has to carry evidence the engine
   processed the tokens the row claims to be about -- non-zero prompt tokens,
   the count actually asked for, a non-zero timed prefill, and a finite positive
   rate. A row that cannot exits 2 (environment failure, never 1) and prints no
   table. Adds a second, engine-independent instrument: wall-clock seconds and
   wall-clock prompt tok/s, so the engine's own numbers are only believed when
   an outside clock agrees.

   Also stops reusing one request id for every concurrent copy and repetition:
   `next_request_id()` was called once and the request cloned, so
   `concurrency * repetitions` sequences were in flight sharing one engine
   handle.

3. The binary aborted at teardown AFTER printing valid results.

   `Request::Terminate` only asks; nothing waited. `engine_handler` had no
   `.join()` anywhere in the workspace, so `main` returned while the engine
   thread was still releasing device memory and CUDA objects, and libc `exit()`
   ran CUDA's atexit handler concurrently with it -- a textbook
   `corrupted double-linked list` / SIGSEGV. `Drop for MistralRs` now joins each
   engine thread with a bounded 30s timeout, and says so loudly if it has to
   detach instead.

   Second half of the same race: `DedicatedDecodePath::new` binds the CUDA
   primary context and warns that skipping it gives "a SEGV in libcuda MOVAPS
   ... no error code, just a fault". Its `Drop` then made the same class of raw
   calls (`cuGraphExecDestroy`, many `cudaFree`) with no bind at all, made legal
   on a foreign thread by `unsafe impl Send`/`Sync`. It now binds exactly as
   `new()` does.

   This matters beyond tidiness: the abort libelled good runs. One chain already
   discarded an expensive valid measurement because its harness gated on rc==0.

Parent system: ArcLab (harness) / ArcInfer (engine timing, teardown).
`grouped.rs` already states the rule: "a harness that reports a per-variant
timing MUST show this counter advancing on the variant it claims to have
measured, and NOT advancing on the others, or the number describes some other
kernel." The counters existed and were public. Nothing in the workspace ever
read them, so every grouped-vs-GEMV and variant-vs-variant comparison this
harness could produce was unproven by construction.

The bench now prints `grouped_launch_counts()` and the selected variant next to
every results table, and says so explicitly when the total is zero -- which is
the common case, because the grouped GEMM declines to run below its tile-fill
boundary (~683 tokens at V4's top-6-of-256 routing) and the measurement is then
the per-pair gather GEMV, not the kernel the reader will assume.

Parent system: ArcLab.
…9.5% prefill)

The trellis grouped GEMM has three variants compiled into every binary. The
tuned ones were measured long ago at +39.4% (v1) and +41.6% (v2) per m-tile,
bit-identical in output, and then left unreachable: `grouped_variant()` fell
back to `QTIP_GROUPED_VARIANT_BASELINE` whenever `ARC_QTIP_GROUPED_VARIANT` was
unset, and the only callers of `set_grouped_variant` in the entire workspace are
inside `examples/qtip_grouped_curve.rs`. So every production prefill since the
variants landed has run variant 0, and the faster kernels existed only for a
benchmark nobody ran in serving.

✅ MEASURED END-TO-END, H200, DeepSeek-V4-Flash qtip2b, pp2048 b=1,
--no-paged-attn, on `arc-prefill`:

    variant 0 (was default): 561.7 prompt tok/s | 1.780 ms/tok | TTFT 3.675 s
    variant 2 (now default): 671.3 prompt tok/s | 1.490 ms/tok | TTFT 3.080 s
    => +19.5% prefill throughput, -16.2% TTFT

The arm is proved from the runtime launch counters, not from the build:
[129,0,0] on the control against [0,0,129] on the variant -- 43 layers x 3
expert matrices, and zero launches on the arms not under test. An
engine-independent wall clock agrees (557.2 -> 664.9 tok/s, +19.3%).

The end-to-end gain is smaller than the kernel gain because the expert gather is
~56% of a 2048-token prefill step; backing +41.6% on that share out of a measured
+19.5% overall is consistent.

Also: an unrecognised env value used to fall through to the baseline. With the
baseline now the slow arm, a typo would silently cost ~20% of prefill, so
unknown values warn and keep the default, and `baseline`/`0` is now an explicit
arm so the A/B knob still selects the control.

Parent system: ArcQuant / QTIP.
@heydryft
heydryft force-pushed the agent/prefill-honest-bench-v2 branch from b2805d2 to a2937ff Compare August 21, 2026 17:10
@heydryft
heydryft merged commit c94d026 into release/openrouter-ready Aug 21, 2026
18 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant