Skip to content

fix(ArcGraph): Tensor::arange in the captured V4 forward is a dangling host pointer (4 call sites) - #190

Closed
heydryft wants to merge 4 commits into
release/openrouter-readyfrom
agent/graph-static-input
Closed

heydryft wants to merge 4 commits into
release/openrouter-readyfrom
agent/graph-static-input

Conversation

@heydryft

Copy link
Copy Markdown
Contributor

Standalone bug fix, orphaned in this session with no PR. Opening it because it is small, self-contained, and fixes a real correctness bug on the capture path.

Tensor::arange inside the captured V4 forward allocates on the host and hands the kernel a pointer that is dangling by replay time. Four call sites across layers.rs, models/deepseek4.rs, and models/dsv4_attention.rs.

This is the same bug class that CLAUDE.md already warns about — "Never use Tensor::{from_vec,arange} in hot loops" — but the capture path makes it worse than a sync: the pointer is not just slow, it is invalid on every replay after the first.

Blast radius

  • DEFAULT-FLIP: none. No new env var, no default changed.
  • KERNEL: none. No .cu added or modified.
  • FORMAT/ABI: none.
  • DEPS: candle 88d86a2 — identical to master. Candle-neutral.
  • Flagless user: the arange results are precomputed/device-resolved instead of built per-forward. Same values, fewer host allocations. Nothing turns on.

Relationship to the rest of the ArcGraph work

arcgraph/capture-final strictly contains these two commits, so if that branch lands first this becomes redundant. But capture-final also carries a candle rev bump to 9211966 (master is on 88d86a2), 41 commits, and is the base of live in-flight work — so it is a much bigger decision. This fix does not need to wait for it, and it is worth landing on its own.

⚠️ Note for reviewers: mistralrs-core is outside the scoped Clippy lane CI gates on, so this code needs the by-hand check rather than relying on the lint job.

🤖 Generated with Claude Code

…OST pointer

V4 decode capture reached "forward RECORDED" and then SIGSEGV'd on its very
first cuGraphLaunch, with a device-side assert from candle's index_select:

    ids[id_i] < src_dim_size   (T = __nv_bfloat16, I = unsigned int)

That dtype pair is the bf16 cos table indexed by U32 positions, i.e.
`self.cos.index_select(&positions, 0)` in
`DeepSeekV2RotaryEmbedding::forward_at_positions` — NOT the TD-MoE
`tid2eid.index_select` (tid2eid is i64, so it cannot produce T=bf16), and not
the embedding (its ids come from the already-address-stable graph-mode buffer).

`positions` was built one line earlier by `compressed_kv_from_rows` with
`Tensor::arange(0u32, t_c, dev)`. `arange` materialises a transient host `Vec`
and uploads it via `CudaDevice::clone_htod` -> an ASYNC `cuMemcpyHtoDAsync`.
Under `cuStreamBeginCapture` that copy is not executed, it is RECORDED: the
graph stores the HOST POINTER and re-reads it on the first launch and on every
replay. The `Vec` is freed when the expression returns, so the graph copied
freed host memory into `positions` and handed garbage indices to index_select.

This is not an address-staleness bug — it fires on the first launch, before any
replay — which is why the static input_ids buffer did not prevent it. candle
already documents this exact mechanism and works around it in
`CudaDevice::htod_info` (it leaks the host source while capturing); nothing
protects `clone_htod`, which is the path `arange` takes.

Fix: serve the strided compressed positions from a per-(ratio, device) table
built once, outside capture, and take a zero-copy `narrow` view each step. The
values are a pure function of (ratio, capacity) and never change. This also
removes a host round trip from every compressor layer of every decode step.

Keyed per ratio because V4 interleaves CSA and HCA compression ratios; a
single-slot cache would thrash and reintroduce the copy on every layer.
A sweep of every host->device constructor reachable from the B=1 decode forward
found three more `Tensor::arange` calls with the same defect as the compressor
positions, all in `dsv4_attention`, all unconditional on every layer of every
step:

  :610  kp = arange(raw_base, raw_base + t_k)   absolute key positions
  :613  qp = arange(q0, q0 + t_q)               query positions (ONE element
                                                at decode -- a whole clone_htod
                                                for 4 bytes, 43x per token)
  :628  bp = arange(0, t_c)                     compressed-block indices

Each is a contiguous integer ramp cast to F32, so all three are views of a
single cached device ramp: `layers::positions_f32(start, len, dev)` narrows
`[0.0, 1.0, 2.0, …]`. Capacity grows by powers of two from 8192, so a typical
run allocates the ramp once during warmup and never again -- rebuilding inside
capture is precisely the hazard being removed.

Fixing only the compressor site would not have been enough: any one of these
would have re-armed the same dangling-host-pointer crash on the first launch.

Also removes three H2D copies per layer per token from the decode hot path
(129 per token at 43 layers), independent of graph capture.
@heydryft

Copy link
Copy Markdown
Contributor Author

Recording a decision on the two ArcGraph branches this one relates to, so they are not left as an implicit 'maybe later'.

arcgraph/capture-final (41 commits ahead of master) and arcgraph/samestep-probe (44 ahead) are both unmerged with no PR. Decision: they should be proposed as ONE PR by the agent that owns the ArcGraph capture stack, not by this queue-clearing pass. Two concrete reasons, not reluctance:

  1. They carry a candle rev bump — both pin 9211966, master pins 88d86a2. A dependency rev move needs the diff between those two candle revs audited and justified in the PR body. I have not audited it, and a rev bump that rides in silently on a 41-commit branch is exactly the thing that should not happen to the moat.
  2. arcgraph/rope-device-pos is being authored on top of samestep-probe right now. Proposing its base underneath live work invites a rebase conflict for whoever is mid-edit.

agent/graph-static-input — this PR — is the piece that can move independently: 2 commits, candle-neutral at 88d86a2, no default flipped, no kernel touched. arcgraph/capture-final strictly contains these two commits, so if the big stack lands first this becomes redundant and should be closed rather than merged. Either outcome is fine; what matters is that the dangling-host-pointer fix is not lost in a branch nobody proposed.

Owner: whoever finishes arcgraph/rope-device-pos. Action: one PR for capture-final + samestep-probe with the candle 88d86a2 -> 9211966 diff justified in the body.

@heydryft

Copy link
Copy Markdown
Contributor Author

🚫 UNVERIFIED — NOT MERGEABLE. Do not read the check list on this PR as a pass.

Every CI run on this PR is in CANCELLED state with no SUCCESS alongside it. I cancelled them deliberately to free a starved Actions queue (27 runs deep, master's own post-merge confirmation was not dispatching a single job). That was capacity triage, not a retry — but it produces the same artifact either way: a check list with no FAILURE in it that also proves nothing.

The standing rule is that CANCELLED is acceptable only when the same job also has a SUCCESS. These runs fail that test. An agent scanning for red and finding none here would be reading a vacuous signal — the exact class this session has hit repeatedly.

This PR has never had a real CI verdict. Treat it as red until it does.

I am re-triggering CI staggered (not all six at once — simultaneous dispatch is what exhausted the runners). Once this PR shows a genuine SUCCESS on the full matrix plus both CUDA lanes, this note is discharged and the blast-radius discussion above becomes the only thing standing between it and merge.

To re-trigger by hand: gh run rerun -R aeonmindai/arc <run-id>, or push any commit.

@heydryft

Copy link
Copy Markdown
Contributor Author

Cause identified, and CI re-triggered — this one CAN reach a genuine verdict.

The single FAILURE across #185-#190 was CI complete, the rollup gate, not six separate problems and not infrastructure. Its log:

LANE_RESULTS: cancelled cancelled cancelled cancelled cancelled cancelled cancelled cancelled cancelled
##[error]A CI lane reported 'cancelled'. Refusing to report green.

That is ci-complete working precisely as intended: it caught my capacity-triage cancellation and turned it into a loud red rather than an ambiguous blank. Correcting one thing I had assumed: no re-trigger had actually run — those were still the original cancelled run IDs, so nothing had been re-verified.

This PR targets master, so unlike #186-#189 it can pass the base-branch gate and produce a real verdict. I have re-triggered its runs (staggered against the other, to avoid re-creating the runner starvation that caused the cancellations in the first place).

The UNVERIFIED note stands until this shows a genuine SUCCESS on CI complete. A green here discharges the CI objection only — the blast-radius objection above is a separate gate and is unaffected by any CI result.

@github-actions

github-actions Bot commented Aug 20, 2026 •

Copy link
Copy Markdown
Code Metrics Report
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
 Language              Files        Lines         Code     Comments       Blanks
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
 C Header                  5          305          210           52           43
 CSS                       2         1181         1036           34          111
 CUDA                     74        25734        18232         4703         2799
 Dockerfile                1           39           22            8            9
 JavaScript               16         3546         2676          482          388
 Jinja2                    7          694          656            5           33
 JSON                     74         4600         4597            0            3
 Makefile                  1            6            5            0            1
 Metal Shading Lan|       33        12224         9431         1142         1651
 PowerShell                1          300          227           30           43
 Python                  145        15139        12482          811         1846
 Shell                    40         9948         6637         2655          656
 Plain Text                4         3801            0         2479         1322
 TOML                     33         1498         1294           54          150
 YAML                      3           25           23            2            0
─────────────────────────────────────────────────────────────────────────────────
 HTML                      4         2687         2604           43           40
 |- CSS                    2          543          479           37           27
 |- JavaScript             1         1233         1215           12            6
 (Total)                             4463         4298           92           73
─────────────────────────────────────────────────────────────────────────────────
 Jupyter Notebooks         4          122           83           23           16
 |- Markdown               1           60           30           22            8
 |- Python                 1          122          113            1            8
 (Total)                              304          226           46           32
─────────────────────────────────────────────────────────────────────────────────
 Markdown                204        45330            0        35183        10147
 |- BASH                  72         1655         1203          331          121
 |- C                      3           17           17            0            0
 |- CUDA                   2           84           56           16           12
 |- JSON                  18          708          708            0            0
 |- PowerShell             1            1            1            0            0
 |- Python                23         1008          787          113          108
 |- Rust                  66         2051         1716           77          258
 |- TOML                   6          207          164            0           43
 |- YAML                   5           41           36            5            0
 (Total)                            51102         4688        35725        10689
─────────────────────────────────────────────────────────────────────────────────
 Rust                    673       331228       285444        16819        28965
 |- Markdown             491        29538          471        25464         3603
 (Total)                           360766       285915        42283        32568
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
 Total                  1324       495675       352655        90603        52417
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

Same root cause as #181, and the identical fix: `.typos.toml` already exempted
the bare Rust identifier `arange` (candle's/numpy's `Tensor::arange`) under
[type.rust.extend-identifiers], but that only matches WHOLE identifiers — the
word inside `matches_the_arange_expression_it_replaced` still tripped the gate.
Add the word-level exemption rather than renaming, because `arange` is a real
API name, not a misspelling of 'arrange'.

This branch and #181 both carry those test functions, so both hit it. Whichever
lands first makes the other's copy redundant.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@heydryft heydryft closed this Aug 20, 2026
@heydryft heydryft reopened this Aug 20, 2026
Two prior attempts (a push to the branch and a close/reopen of the PR) fired
only the pull_request_target Analysis workflow; Continuous integration and
CUDA compile check never dispatched for 489134b, leaving the PR displaying
only 'comment' — an absence that reads as a pass. Empty commit to force a
synchronize on a fresh sha.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@heydryft

heydryft commented Aug 20, 2026 •

Copy link
Copy Markdown
Contributor Author

🔴 UNVERIFIED note STANDS — and the reason is a trap worth naming, because nothing in the repo currently guards it.

This PR displays comment ✓ and nothing else. That is not a pass. Zero lanes ran.

Root cause: mergeable=CONFLICTING, mergeStateStatus=DIRTY. GitHub runs pull_request workflows against the merge ref (head merged into base). A conflicting PR has no merge ref, so Continuous integration and CUDA compile check cannot dispatch at all. Only Analysis fires, because it is pull_request_target and runs against the base. Result: a check list with no red in it, and no evidence in it either.

I confirmed this is a mechanism, not a fluke — three separate attempts to force dispatch all produced only Analysis:

  1. a push to the branch (the .typos.toml fix, 489134b91),
  2. a close/reopen of the PR,
  3. an empty commit on a fresh sha (240982c8c).

Zero Continuous integration runs exist for any sha after fb58fc63c.

🔑 This is the same failure shape as the stacked-PR incident that motivated the base-branch lane — "ran zero lanes, displayed only comment, and read as green while nothing had been compiled, tested, linted or nvcc'd" — but it arrives through a different door. base-branch catches PRs that target a non-master base. It cannot catch this one, because this PR does target master; the lanes simply never dispatch. A conflicting PR is currently indistinguishable from a passing one at a glance. Worth a guard of its own.

Why it conflicts: it did not when opened. The seven merges I landed this session moved mistralrs-core/src/models/deepseek4.rs and dsv4_attention.rs underneath it — the two files that now conflict.

What it needs: a merge from master resolving those two files. I am not doing it here. Both are V4 model files on the capture path, and the fix this PR carries is itself about dangling host pointers during capture — resolving that conflict without auditing the interaction is precisely the kind of guess that should not be made on a queue-clearing pass, and there is no GPU to check it against.

Its value is unchanged and still worth landing: 4 Tensor::arange call sites that allocate on the host and hand the kernel a pointer that is dangling by replay time. Candle-neutral at 88d86a2, no default flipped, no kernel touched. It needs an owner with the capture path in context, not a rebase by someone clearing a queue.

@heydryft
heydryft changed the base branch from master to release/openrouter-ready August 21, 2026 15:37
@heydryft

Copy link
Copy Markdown
Contributor Author

Retargeted at the integration branch

Base changed: master → release/openrouter-ready.

The queue is being restructured to the shape the owner asked for: one PR open against master (#194), with everything else merging into a single integration branch. Agents branch off release/openrouter-ready and PR back into it; #194 is the single gate from there to master.

Two things had to land on master first for this to be workable, and both have:

  1. The base-branch CI lane used to hard-fail any PR not targeting master (ci(ArcGate): let the base-branch lane permit the single integration branch — 7 PRs are red for addressing, not quality #216). Seven PRs were red for their addressing, not their quality. The lane now accepts master or release/openrouter-ready — an exact-match allowlist, nothing else — and its intent is intact: release/openrouter-ready reaches master only through Arc → OpenRouter-ready: the single integration PR (everything else merges into this branch) #194, whose own base is master, so nothing enters master without a full run whose base is master.
  2. release/openrouter-ready was re-cut from current master. It had diverged badly — merging it as-is would have reverted TCFRAG (perf(ArcQuant/ArcKernels): TCFRAG-2B — put the qtip2b trellis GEMV on the tensor cores #203) and re-broken the flag-polarity work (fix(ArcGate): read boolean ARC_* flags by value, so =0 means off #212). It is now master plus only the work it uniquely carried.

This PR was not closed and is not considered stale. An audit of the queue found the overwhelming majority of it to be real work that was never merged, not noise.

What you need to do: rebase onto release/openrouter-ready and resolve conflicts against it rather than against master. CI lanes will now actually run and report on this PR instead of failing at the base check.

Retargeted, not closed — the supersession claim does not hold

This PR was put forward for closure as superseded by fdf0b6c19 (#205, "ArcGraph: land the CUDA-graph capture lane"). I checked that before acting on it, and it is only half true, so the PR stays open.

fdf0b6c19 is on master (git merge-base --is-ancestor fdf0b6c19 origin/master → 0). But this branch introduces two helpers, and only one of them shipped:

helper on master? evidence
compress_positions (+ COMPRESS_POS_CHUNK) ✅ landed via fdf0b6c19 layers.rs:2435, applied at one site in dsv4_attention.rs
positions_f32 (+ IOTA_MIN_LEN) ❌ not on master git grep positions_f32 origin/master → no match

positions_f32 is this PR's headline — the cached F32 device ramp that replaces Tensor::arange in the captured V4 forward. Master still calls Tensor::arange at mistralrs-core/src/models/dsv4_attention.rs lines 372, 396, 888, 896 and 950. fdf0b6c19 replaced exactly one arange site, with the U32 compress_positions, not with this.

So closing this would have discarded real, unmerged work — the four-call-site fix the title is about. Retargeted at release/openrouter-ready instead.

⚠️ One caveat worth carrying forward for whoever rebases this: the capture-lane findings recorded since this PR was opened list the mask-arange as innocent on the performance question. That does not make the change worthless — the argument here is correctness under CUDA-graph capture (a recorded H2D memcpy holding the host pointer of an already-freed Vec), which is a different claim from "it is slow". Rebase it on the merits of that claim, and re-check it against the capture lane as it now exists on master.

@heydryft

Copy link
Copy Markdown
Contributor Author

Rebased onto release/openrouter-ready: this is already landed. Closing.

Rebasing this branch onto the integration branch leaves nothing — both
substantive commits apply empty, and the only non-empty residue is a
.typos.toml change the integration branch has explicitly rejected.

The two commits are already on release/openrouter-ready, under different SHAs:

here on release/openrouter-ready
2743a4400 Tensor::arange in the captured forward is a dangling HOST pointer c1d2b27f4, same title
fb58fc63c the other three arange calls in the captured V4 forward 64fe1d678, same title

On the positions_f32 question. It is true that git grep positions_f32 origin/master returns nothing, and that Tensor::arange still appears at
dsv4_attention.rs:372, 391, 396, 888, 896, 950. But neither fact means the fix
is missing — the helper was deliberately deleted, by this lane, in
6fcb87f47 on the integration branch:

layers::positions_f32 is REMOVED, with its IOTA_F32 cache. #206
(02edd31cc) deleted two of its three call sites outright — at t_q == 1 the
raw-window mask is provably all-ones, so the kp/qp chain that fed it is gone —
and this rebase drops the third swap in favour of master's spelling. That left
a pub fn with ZERO callers workspace-wide, and two tests exercising nothing
but itself, which is a vacuous test by the repo's own definition.

So the capture path got a stronger fix than this PR's, not a weaker one.
#206 (a7b0890d9, the decode raw-window mask is all ones — 467
launches/token
) does not merely move the two arange calls off the host, it
deletes them: at t_q == 1 — every decode step, i.e. the captured forward —
raw_valid is now Tensor::ones(..) and the compressed-branch block indices
come from layers::compress_positions(t_c, 1, dev), a zero-copy view of a table
built once outside capture.

The surviving Tensor::arange calls are not in the captured forward. Lines
888/896/950 all sit on the t_q > 1 arm (let decode_row = t_q == 1;), and
372/391/396 are in ragged_union_mask. Merging this PR would take the
positions_f32 side of those hunks and thereby revert #206's identity, revert
the graph_positions handling that keeps fixed-capacity graph decode correct,
and fail to compile besides — qp is now Option<Tensor>, and this branch's
code does &qp + 1.0.

What is left after rebase: five lines of .typos.toml adding arange to
[default.extend-words]. 6fcb87f47 chose the opposite fix on purpose —
renaming the test to matches_the_strided_range_expression_it_replaced, "which
is the precedent set in 745c871dd and e4eb59dfb" — and no identifier in the
tree currently contains arange inside a longer snake_case name, so the entry
would be dead on arrival.

Prefill-path arange in ragged_union_mask (3 calls) and on the t_q > 1
attention arm (3 calls) is still real, but it is not a capture-correctness
issue and is not what this PR patched. Worth a separate change if the prefill
budget wants it.

(No behaviour was changed on the integration branch by this analysis; the
branch was rebased locally only, to measure the residue.)

@heydryft

Copy link
Copy Markdown
Contributor Author

Superseded — see the analysis above. Rebases to an empty diff plus a typos allow-list entry the integration branch rejected.

@heydryft

Copy link
Copy Markdown
Contributor Author

📌 Correcting the note that sent me here — so this does not get re-opened

This PR was handed to me with the instruction: "Its positions_f32 helper is
genuinely not on master (git grep positions_f32 origin/master → no match)
while master still calls Tensor::arange at dsv4_attention.rs:372,396,888,896, 950. A prior claim that this was superseded was verified false. Land the
4-call-site fix."

Both observations are true. The conclusion drawn from them is backwards, and
the "prior claim" it overrules was right.
Writing that here, on the PR, so the
next person reading the same stale note does not re-open this.

The reasoning error is worth naming precisely, because it will recur: absence
of the helper
was read as absence of the fix. It was the opposite. The helper
is gone because the fix landed in a stronger form —
6fcb87f47, on the integration branch, deleted positions_f32 on purpose:

layers::positions_f32 is REMOVED, with its IOTA_F32 cache. #206
(02edd31cc) deleted two of its three call sites outright — at t_q == 1 the
raw-window mask is provably all-ones, so the kp/qp chain that fed it is gone —
and this rebase drops the third swap in favour of master's spelling. That left
a pub fn with ZERO callers workspace-wide, and two tests exercising nothing
but itself, which is a vacuous test by the repo's own definition.

And the surviving Tensor::arange calls are not in the captured forward.
Lines 888/896/950 sit on the t_q > 1 arm (let decode_row = t_q == 1;);
372/391/396 are in ragged_union_mask. On every decode step — the captured path,
the one this PR exists for — raw_valid is now Tensor::ones(..) and the
compressed block indices come from layers::compress_positions(t_c, 1, dev), a
zero-copy view of a table built once outside capture. #206 does not move those
two host→device copies off the host; it deletes them.

Merging this would have (a) reverted #206's all-ones identity, (b) reverted the
graph_positions handling that keeps fixed-capacity graph decode correct, and
(c) not compiled — qp is now Option<Tensor> and this branch does &qp + 1.0.

This is the second time today a "missing from master" claim was wrong in this
direction.
The check that would have caught both is cheap: before concluding a
fix is absent, git log --grep the integration branch for the commit subject,
not just git grep for a symbol. Here it returns c1d2b27f4 and 64fe1d678 —
this PR's own two commits, already landed.

Rebased onto release/openrouter-ready, this branch produces an empty diff
plus five lines of .typos.toml that 6fcb87f47 explicitly rejected in favour
of renaming the test (the precedent set by 745c871dd / e4eb59dfb), and which
no identifier in the tree currently needs.

Staying closed. The genuine residue — six prefill-path arange calls in
ragged_union_mask and on the t_q > 1 arm — is real but is not a
capture-correctness issue and is not what this PR patched. That is a separate
change, if the prefill budget wants it.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant