Repository navigation
ci: run the full lane on stacked PRs, and refuse to report green when it did not (D18 #11) - #106
Conversation
… it did not (D18 #11) `ci.yml` and `cuda_compile_check.yaml` both triggered on `pull_request: branches: [master]`. A PR whose base is another PR's branch matched neither, so not one lane ran — no compile, test, rustfmt, clippy, MSRV or nvcc. The only workflow without a branch filter is analysis.yaml (`pull_request_target`), so those PRs showed exactly one check, `comment`, and rendered green. Six open PRs were in that state at once: the stack #95 -> #100 -> #102 -> #103 -> #104, plus #98. Check-run census: 16 checks on master-based PRs, 1 on every stacked one. This compounds with the macOS gap — a stacked PR touching cuda-gated Rust got neither the nvcc lane nor a local type-check, since cargo check on macOS cannot see those arms. Two changes, because the trigger fix alone is not enough: 1. `pull_request.branches: ['**']` on both workflows. Stacked PRs are normal here and must get the full lane. 2. A `ci-complete` guard job that needs every lane, runs `if: always()`, and fails unless each reports success — skipped and cancelled included — and asserts the lane count. A branch filter is only one route to 'the lanes did not run'; a skipped job, a renamed job, an empty matrix or an over-narrow paths filter all reproduce it. Unlike the individual lanes, ci-complete cannot pass by not existing, which makes it the correct required status check for branch protection. flash_attn_compile_check.yaml is left dispatch-only by design.
Widening the pull_request trigger multiplies this workflow across every stacked PR, and ci.yml is 12 jobs — the largest multiplier in the repo, and the only workflow with no concurrency group. cuda_compile_check.yaml has carried the identical block for a while; this is the same shape on purpose. Without it every push leaves its superseded runs sitting in a shared queue. Measured today rather than assumed: 38 in-flight runs, of which 13 were on SHAs that were no longer any PR's head, cancelled by hand to free capacity. perf/qtip-grouped-gemm-arch alone held two queued CI runs, one obsolete. The 'head_ref || run_id' idiom is load-bearing. On pull_request, head_ref is the source branch, so a new push cancels that PR's own in-flight run. On push, schedule and workflow_dispatch, head_ref is empty and run_id is unique per run, so master pushes and scheduled runs are never cancelled by each other — only PR iterations are.
|
Two additions, plus a live instance that replaces the hypothesis in the description with a dated example. #109 — the defect, happening right nowPR #109 (
So the only check that fires on its own is the cosmetic tokei line-counter. No compile, no test, no rustfmt, no clippy, no MSRV — and the CUDA lane exists purely because someone is manually re-dispatching it on every push. Three of those hand-dispatched runs had already been superseded and were cancelled today. Worth stating precisely, because the shorter version of this story ("#109 got only The
|
Code Metrics Report━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Language Files Lines Code Comments Blanks ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ C Header 5 305 210 52 43 CSS 2 1181 1036 34 111 CUDA 72 24328 17592 4018 2718 Dockerfile 1 39 22 8 9 JavaScript 16 3546 2676 482 388 Jinja2 7 694 656 5 33 JSON 74 4600 4597 0 3 Makefile 1 6 5 0 1 Metal Shading Lan| 33 12224 9431 1142 1651 PowerShell 1 300 227 30 43 Python 143 14830 12217 797 1816 Shell 23 5756 3964 1412 380 Plain Text 4 3801 0 2479 1322 TOML 33 1485 1292 43 150 YAML 3 25 23 2 0 ───────────────────────────────────────────────────────────────────────────────── HTML 4 2687 2604 43 40 |- CSS 2 543 479 37 27 |- JavaScript 1 1233 1215 12 6 (Total) 4463 4298 92 73 ───────────────────────────────────────────────────────────────────────────────── Jupyter Notebooks 4 122 83 23 16 |- Markdown 1 60 30 22 8 |- Python 1 122 113 1 8 (Total) 304 226 46 32 ───────────────────────────────────────────────────────────────────────────────── Markdown 188 37465 0 28743 8722 |- BASH 69 1614 1187 311 116 |- C 2 12 12 0 0 |- CUDA 2 84 56 16 12 |- JSON 18 708 708 0 0 |- PowerShell 1 1 1 0 0 |- Python 23 1008 787 113 108 |- Rust 65 2048 1713 77 258 |- TOML 6 207 164 0 43 |- YAML 4 38 33 5 0 (Total) 43185 4661 29265 9259 ───────────────────────────────────────────────────────────────────────────────── Rust 662 306982 265873 13815 27294 |- Markdown 477 22543 471 19392 2680 (Total) 329525 266344 33207 29974 ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Total 1276 450597 329477 73114 48006 ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ |
|
Bypass record. Admin-merged without review — no second reviewer on this fork. The bypass covered the review requirement only; every required status check was green, and unusually for this session, so was the guard this PR adds. Required contexts at merge time (the six Rust lanes, pre-switch) all
Eight lanes, matching Follow-up applied: Transition note: branches created before this merge cannot emit a |
The defect
ci.ymlandcuda_compile_check.yamlboth triggered on:A PR whose base is another PR's branch — a stack — matches neither, so not one lane ran: no compile, no test, no rustfmt, no clippy, no MSRV, no nvcc.
The only workflow without a branch filter is
analysis.yaml, which usespull_request_target. So those PRs showed exactly one check —comment— and rendered green.Six open PRs were in that state simultaneously: the stack #95 → #100 → #102 → #103 → #104, plus #98. Check-run census at the time: 16 checks on every master-based PR, 1 on every stacked one.
This compounds with a known gap: a stacked PR touching
#[cfg(feature = "cuda")]Rust got neither the nvcc lane nor a local type-check, becausecargo checkon macOS structurally cannot see those arms. #99's missingqtip_grouped_tile_mre-export is exactly that hole — it survived a local check and was only caught once the nvcc lane was forced to run.The fix
1.
pull_request.branches: ['**']on both workflows. Stacked PRs are normal in this repo and must get the full lane.2. A
ci-completeguard job. The trigger fix alone is not enough, because a branch filter is only one route to "the lanes did not run" — a skipped job, a renamed job, a matrix that expands to nothing, or an over-narrowpaths:filter all reproduce it exactly.ci-completedepends on every lane, runsif: always(), and fails unless each reportssuccess(treatingskippedandcancelledas failures). It also asserts the lane count, so deleting a lane trips it rather than quietly shrinking what "green" covers.Unlike the individual lanes,
ci-completecannot pass by not existing. That makes it the correct required status check for branch protection, and I'd suggest switching the protection rule to it.flash_attn_compile_check.yamlis deliberately leftworkflow_dispatch-only — the flash-attn kernel compile is heavy and that's an intentional choice documented in the file.Why this is filed as doctrine
Recorded as D18 instance 11 in
memory/mission/KERNEL_RULES.md. Every prior D18 instance was a single code path lying about itself; this one is the verification layer being absent for a whole class of PR while presenting as a pass. Same mechanical form as the other ten: the absence of a signal was read as a specific signal — here, "no CI configured for this base branch" rendered as "checks passed."Corollary worth keeping for tooling and agents: count the checks before trusting the conclusion.
0 failuresout of1 checkis not a pass.