docs(ArcGate): record #167's moved dispatch gate as an open question needing a hardware A/B - #192
Merged
Merged
Conversation
#167 is the only one of the seven PRs merged in the 2026-08-20 queue sweep that changes behaviour for a flagless user, and it landed on a green-CI gate that cannot see performance. Recording it in the register so it stays a visible open question rather than a merged assumption. The gate moved the 9-682 token band from grouped GEMM to gather-GEMV -- batched decode at B>=9 and every prefill chunk under 683 tokens, which is where chunked prefill lives. The path it now prefers carries 3.15x redundant reads, so it may be slower exactly where it fires. Names the A/B that settles it (B=16/32/64, chunks 128/256/512) and the coupling with the open #99, whose GROUPED_TILE_M -> 64 on SM90+ moves the threshold ~4x and which must read grouped_tile_m_for_cc or the gate and kernel will disagree on tile size. Also records the caveat covering the whole sweep: every number attached to the seven merged PRs was measured on that branch's own base, not on current master, and none may be quoted until re-measured. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Code Metrics Report━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Language Files Lines Code Comments Blanks ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ C Header 5 305 210 52 43 CSS 2 1181 1036 34 111 CUDA 74 25734 18232 4703 2799 Dockerfile 1 39 22 8 9 JavaScript 16 3546 2676 482 388 Jinja2 7 694 656 5 33 JSON 74 4600 4597 0 3 Makefile 1 6 5 0 1 Metal Shading Lan| 33 12224 9431 1142 1651 PowerShell 1 300 227 30 43 Python 145 15139 12482 811 1846 Shell 40 9948 6637 2655 656 Plain Text 4 3801 0 2479 1322 TOML 33 1498 1294 54 150 YAML 3 25 23 2 0 ───────────────────────────────────────────────────────────────────────────────── HTML 4 2687 2604 43 40 |- CSS 2 543 479 37 27 |- JavaScript 1 1233 1215 12 6 (Total) 4463 4298 92 73 ───────────────────────────────────────────────────────────────────────────────── Jupyter Notebooks 4 122 83 23 16 |- Markdown 1 60 30 22 8 |- Python 1 122 113 1 8 (Total) 304 226 46 32 ───────────────────────────────────────────────────────────────────────────────── Markdown 204 45280 0 35143 10137 |- BASH 72 1655 1203 331 121 |- C 3 17 17 0 0 |- CUDA 2 84 56 16 12 |- JSON 18 708 708 0 0 |- PowerShell 1 1 1 0 0 |- Python 23 1008 787 113 108 |- Rust 66 2051 1716 77 258 |- TOML 6 207 164 0 43 |- YAML 5 41 36 5 0 (Total) 51052 4688 35685 10679 ───────────────────────────────────────────────────────────────────────────────── Rust 673 331228 285444 16819 28965 |- Markdown 491 29538 471 25464 3603 (Total) 360766 285915 42283 32568 ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Total 1324 495625 352655 90563 52407 ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Issues are disabled on this repo, so this is the tracking record for the one merged change in the 2026-08-20 queue sweep that is not inert.
docs/engineering/OPEN_QUESTIONS.mdis the right home: it is explicitly the register of things "deliberately not done, with the reasoning that led to each decision" and "the honest list of what has not been measured. Nothing on that list may be claimed, quoted, or reasoned from."What it records
#167 (master
b2841f5eb) moved the MoE dispatch gate from token count to m-tile occupancy. b=1 decode and full prefill are unchanged, but 9 ≤ n ≤ 682 now prefers gather-GEMV over grouped GEMM — batched decode at B≥9 and every prefill chunk under 683 tokens, which is where chunked prefill operates.It merged on a green-CI gate, and CI cannot see performance. The path it now prefers is a no-dedup GEMV with 3.15× redundant reads, so it may be a regression exactly where it fires. Exposure is performance, not correctness — both paths compute the same thing.
The entry names the A/B that settles it (B = 16/32/64; prefill chunks 128/256/512), states plainly that it needs hardware, and records the coupling with the open #99: raising
GROUPED_TILE_Mto 64 on SM90+ moves this threshold ~4×, so #99 must make the gate readgrouped_tile_m_for_ccrather than the constant — otherwise the gate and the kernel disagree about tile size, which is a wrong-answer bug rather than a perf bug.It also carries the caveat covering the whole sweep: every performance number attached to the seven merged PRs was measured on that branch's own base, not on current master. None is a current fact until re-measured — including the 15.0 → 32.6 tok/s ladder figures on the still-open #185–#189.
Blast radius
Documentation only. One file, one new subsection under §2 "Ready to run, not yet run". No code, no kernel, no format, no dependency movement.
🤖 Generated with Claude Code