Skip to content

[Review pipeline] Zero PRs mergeable in fast-mlsirm/LineageWeave: hash-pinning, cross-repo status 403, dispatch throughput #1212

Description

@seonghobae

Problem

Zero PRs are currently mergeable in ContextualWisdomLab/fast-mlsirm (60 open) or ContextualWisdomLab/LineageWeave (12 open) because the required branch-protection review (required_approving_review_count: 1) has not been satisfied by any bot on any current-head PR in either repo (reviewDecision is REVIEW_REQUIRED, CHANGES_REQUESTED, or empty on every open PR — none APPROVED). PR authors cannot self-approve (GitHub platform rule), so this is a hard blocker on the whole PR backlog in both repos, not a per-PR defect.

Investigated via gh run view/gh api .../check-runs/.../annotations on recent opencode-review-dispatch.yml and *-hourly-review-repair.yml runs. Three independent, concrete root causes found:

1. LineageWeave's uv.lock cannot be hash-pinned because two dependencies are git-sourced

scripts/ci/materialize_base_python_requirements.py requires the base branch's uv.lock to export as a fully hash-pinned requirements closure before the sandboxed coverage-evidence job (and therefore the whole approval decision) can run. LineageWeave's uv.lock pins fast-mlsirm and RankWeave via source = { git = "https://github.com/ContextualWisdomLab/....git?rev=<sha>#<sha>" }. uv export cannot emit a --hash=sha256:... for a git dependency (there is no fixed sdist/wheel artifact to hash — installing from git means running the package's build backend), so every export of LineageWeave's main lock fails _is_fully_hash_pinned_requirement, and _export_uv_lock raises "uv export for tracked base lock uv.lock was not fully hash-pinned". Every dispatch run against a LineageWeave PR fails at the coverage-evidence job for this reason (confirmed on PR #373, #387, #258, #349 dispatch runs 2026-08-21 21:3x UTC), regardless of what the PR itself changes.

Reproduced: LineageWeave's pinned fast-mlsirm commit (5006c38...) is 477 commits ahead of the latest tagged release (v0.6.0), and fast-mlsirm's release-tag.yml only runs gh release create — it does not build or attach an installable sdist/wheel asset, and the package is not published to PyPI (pypi.org/pypi/fast-mlsirm/json → 404) despite the org's PYPI_API_KEY secret already being available. Same shape of problem likely applies to RankWeave.

Recommended fix (do not relax the sandbox's hash-pinning requirement — that requirement is a real supply-chain control, since installing an unpinned/unhashed git dependency in an untrusted-PR sandbox means running arbitrary build-backend code from a mutable ref class):

  • Add a maturin build --release (or equivalent) step to fast-mlsirm's and RankWeave's release workflow that builds sdist + wheels and attaches them as GitHub Release assets (or publishes to PyPI using the existing PYPI_API_KEY).
  • Cut a release close to the commit LineageWeave currently depends on.
  • Switch LineageWeave's pyproject.toml entries for fast-mlsirm/RankWeave from git sources to a versioned PyPI dependency (or a direct release-asset URL dependency, which uv can hash), then regenerate uv.lock.

2. Cross-repo opencode-review status publish always fails with 403

In opencode-review-dispatch.yml, the step that publishes the optional opencode-review commit status back to the target repo (e.g. repos/ContextualWisdomLab/LineageWeave/statuses/<sha>) using the opencode-app exchanged token fails every time with gh: Resource not accessible by integration (HTTP 403) (confirmed on runs against appguardrail#969, LineageWeave#373). The OpenCode GitHub App does not appear to have statuses:write granted on these target repos, or is not installed there with that permission. This step is non-authoritative (the doc string says "the exact-head formal review remains authoritative") but it fails loudly on every run and wastes a job step; either grant the App the permission on all org repos it dispatches to, or drop the step.

3. Review-dispatch throughput is far below the PR creation rate

fast-mlsirm-hourly-review-repair.yml (and the generic pr-review-merge-scheduler.yml) invoke scripts/ci/pr_review_fix_scheduler.py / pr_review_merge_scheduler.py with MAX_DISPATCHES: 1 per run, i.e. at most ~1 new review dispatch per repo per scheduled run. Against a backlog of 60 open PRs in fast-mlsirm alone, this cannot keep the queue current even before accounting for root cause #1 above. Separately, the scheduler's own GitHub GraphQL calls (listing open PRs) intermittently fail after only 3 retries with a 1s/2s/4s backoff (Transient GitHub GraphQL error on attempt 1..3/4), aborting the entire scheduled run with no partial progress (confirmed failures at 2026-08-21T22:00 and 21:01 UTC for fast-mlsirm). Recommend widening the retry budget/backoff for transient GraphQL errors before concluding failure, and revisiting MAX_DISPATCHES once root cause #1 is fixed and the backlog is under control (raising throughput without fixing #1 would just burn compute on runs that fail at coverage-evidence anyway).

Evidence

  • gh run view 32529248149 --repo ContextualWisdomLab/.github (LineageWeave#373 dispatch, coverage-evidence failure annotation: Could not materialize base Python locks: uv export for tracked base lock uv.lock was not fully hash-pinned)
  • gh run view 32531069679 --repo ContextualWisdomLab/.github (fast-mlsirm-hourly-review-repair, Transient GitHub GraphQL error x3 then abort)
  • gh run view 32536566379 --repo ContextualWisdomLab/.github (cross-repo status 403 against appguardrail#969)
  • gh pr list --repo ContextualWisdomLab/fast-mlsirm --state open --json reviewDecision / same for LineageWeave: zero APPROVED among 72 combined open PRs.

Suggested scope for follow-up PRs (do not bundle into one PR)

  1. fast-mlsirm, RankWeave: add release-asset (or PyPI) publishing to the release workflow.
  2. LineageWeave: once (1) ships a release, switch the two git-sourced pyproject.toml entries to the released, hash-pinnable versions and regenerate uv.lock.
  3. .github: fix or remove the cross-repo opencode-review status-publish step (403).
  4. .github: widen the scheduler's GraphQL retry budget; revisit MAX_DISPATCHES after (1)-(2) land.

Activity

  1. seonghobae commented on Aug 22, 2026

    @seonghobae
    ContributorAuthor

    Progress on recommended-fix item 1 (fast-mlsirm/RankWeave release-asset publishing): opened ContextualWisdomLab/fast-mlsirm#1207, which adds a publish-pypi.yml workflow (sdist + linux/macos/windows wheels via maturin-action, published to PyPI + attached as release assets on release: published).

    This PR is itself blocked by the exact gate this issue describes — it can't get the required bot approval until the pipeline is fixed, which is exactly the chicken-and-egg this issue flags. Flagging here so it's visible to whoever reviews/repairs the pipeline: merging #1207 (or applying it directly) unblocks a real fix path rather than just working around the symptom.

    Once #1207 lands and a release is cut, LineageWeave's fast-mlsirm/RankWeave git-source uv.lock entries can switch to the released PyPI versions (item in the original recommended-fix scope), which resolves root cause #1 for every LineageWeave PR at once. RankWeave still needs the equivalent workflow — not done yet, tracking as follow-up.

  2. seonghobae commented on Aug 22, 2026

    @seonghobae
    ContributorAuthor

    Checked RankWeave separately (the other git-sourced dependency in LineageWeave's uv.lock): unlike fast-mlsirm, RankWeave is already on PyPI and already has a working publish.yml (trusted-publisher/OIDC based, no token secrets needed) — but PyPI currently only has 0.1.0 published while LineageWeave's git pin is exact-identical to RankWeave's v0.18.0 tag (0 commits ahead/behind). The last publish attempt (2026-08-06, run 31067753363) failed at the trusted-publishing OIDC exchange:

    invalid-publisher: valid token, but no corresponding publisher — claims: sub: repo:ContextualWisdomLab/RankWeave:environment:pypi, workflow_ref: .../publish.yml@refs/tags/v0.18.0

    This is a PyPI project-settings problem (Trusted Publisher config on the rankweave PyPI project no longer matches the repo/workflow/environment claims RankWeave's Actions run presents) — not something fixable from the GitHub side or via gh/git. It needs whoever administers the rankweave PyPI project to fix/re-add the trusted publisher at https://pypi.org/manage/project/rankweave/settings/publishing/ (matching repo ContextualWisdomLab/RankWeave, workflow publish.yml, environment pypi), then re-run publish.yml (workflow_dispatch with release_tag=v0.18.0) to catch PyPI up.

    Once that's done, LineageWeave's RankWeave git-source uv.lock entry can switch to rankweave==0.18.0 from PyPI immediately (no code/build work needed on RankWeave's side, unlike fast-mlsirm).

  3. seonghobae commented on Aug 22, 2026

    @seonghobae
    ContributorAuthor

    New finding via static analysis (actionlint): pr-review-fix-scheduler.yml's "Resolve immutable called-workflow source" step references a nonexistent context, unconditionally fails

    Ran actionlint (v1.7.12) against the review-pipeline workflows in this repo to look for issues the runtime logs hadn't already surfaced. It flagged .github/workflows/pr-review-fix-scheduler.yml:225-228:

    pr-review-fix-scheduler.yml:225:36: property "workflow_repository" is not defined in object type
      {check_run_id: number; container: ...; services: ...; status: string} [expression]
              WORKFLOW_REPOSITORY: ${{ job.workflow_repository }}
    pr-review-fix-scheduler.yml:226:29: property "workflow_sha" is not defined ...
    pr-review-fix-scheduler.yml:227:29: property "workflow_ref" is not defined ...
    pr-review-fix-scheduler.yml:228:35: property "workflow_file_path" is not defined ...
    

    job.workflow_repository / job.workflow_sha / job.workflow_ref / job.workflow_file_path are not properties of the GitHub Actions job context (which only exposes container, services, and status). Each of these ${{ }} expressions evaluates to an empty string at runtime rather than erroring — so the four env vars the "Resolve immutable called-workflow source" step (line 221) sets are always empty, unconditionally, on every run that reaches this step.

    That step immediately does:

    if [ "$WORKFLOW_REPOSITORY" != "$expected_repository" ]; then
      ...
      exit 1
    fi

    Since $WORKFLOW_REPOSITORY is always "" and $expected_repository is "ContextualWisdomLab/.github", this check can never pass — the step fails every single time it executes.

    Confirmed against real runs, not just static analysis: this step doesn't appear at all in pr-review-fix-scheduler.yml runs from 2026-08-16 (older workflow revision, before this step existed — job logs show only pr_review_fix_scheduler.py --self-test, no "Resolve immutable..." step). The step first appears in the 2026-08-20 repository_dispatch-triggered runs (32360044573 failure, 32360412990 cancelled) — i.e. every run that has hit this code path since it was added has failed here. pr-review-merge-scheduler.yml, opencode-review-dispatch.yml, pr-review-autofix.yml, and pr-auto-rebase.yml do not contain this pattern (checked via the same actionlint/grep pass) — the blast radius looks isolated to pr-review-fix-scheduler.yml's repository_dispatch (auto-repair-dispatch) path specifically, not the main review/approval/merge flow this issue already tracks.

    Likely intended fix (not applied — this is a security-relevant workflow-identity check and I don't want to guess-patch it): GitHub Actions doesn't expose "which exact commit of the currently-executing workflow file is running" under job.*. The relevant built-ins are github.workflow_ref / github.workflow_sha (current workflow file's ref/SHA) and, for a workflow invoked as a reusable workflow via uses:, github.job_workflow_ref / github.job_workflow_sha on the caller's side. Whoever owns this script should confirm which of those (or a different mechanism entirely, e.g. parsing github.workflow_ref server-side) was actually intended before changing it, since the four-way split (repository/sha/ref/file_path) suggests the author meant to parse a combined owner/repo/.github/workflows/file.yml@ref string rather than read four separate fields.

    Evidence commands:

    • actionlint -shellcheck= -pyflakes= pr-review-fix-scheduler.yml (reproduces the four warnings above)
    • gh run list --repo ContextualWisdomLab/.github --workflow pr-review-fix-scheduler.yml --json databaseId,conclusion,event,createdAt
    • gh api repos/ContextualWisdomLab/.github/actions/jobs/95141441945/logs (2026-08-16 success — step absent) vs the 2026-08-20 failing run logs (step present, fails)
  4. seonghobae commented on Aug 22, 2026

    @seonghobae
    ContributorAuthor

    Opened #1221 for root cause #3 (the actionlint-caught nonexistent job.workflow_* context references in pr-review-fix-scheduler.yml). Also found and fixed the same broken pattern in exact-artifact-sbom-attestation.yml (3 occurrences) while I was in there — its trusted-verifier checkout and signer-repository resolution had the identical bug, though it happened to fail closed downstream via gh attestation verify rather than silently bypassing signer verification. Full test suite green (1330 passed / 1 skipped, unrelated env-only failure).

  5. added
    area: authAuthentication, authorization, identity, or tenant isolation
    area: ci-cdCI, GitHub Actions, checks, release, or supply chain
    area: securitySecurity boundary, hardening, or vulnerability prevention
    status: triagedOpen issue has an organization taxonomy assignment
    type: featureNew or expanded product capability
    on Aug 22, 2026
  6. seonghobae commented on Aug 22, 2026

    @seonghobae
    ContributorAuthor

    Root cause #3 landed: #1221 merged. Note this ended up being a two-part fix — the original PR only replaced the nonexistent job.workflow_* context references with github.workflow_ref/github.workflow_sha, but those still don't identify a called reusable workflow's own file (they reflect the top-level caller inside a workflow_call target), so the scheduler's identity guard would have kept failing closed on every real hourly-caller invocation even after that first push. Caught via Devin review and fixed properly before merge: validate github.repository instead, since every current caller uses a local same-repo uses: ./.... Watching fast-mlsirm#1207 and TEPP#153 as bellwethers for whether main-branch PRs start getting real bot APPROVE decisions now.

  7. seonghobae commented on Aug 22, 2026

    @seonghobae
    ContributorAuthor

    New finding after #1214/#1235/#1236/#1240/#1241 (all merged, confirmed working end-to-end in production)

    Every GitHub-API-layer issue in the mention-dispatch pipeline is now fixed: retry-with-backoff on the shared installation-token rate limit (scoped safely to avoid resending non-idempotent writes), a wall-clock time budget so the sweep job exits cleanly instead of hitting its 15-minute Actions timeout, unbuffered output so diagnostics actually reach the log, and isolation so one repo's listing failure doesn't crash the whole cycle. Confirmed live: a recent sweep run cleanly logged "Agent mention sweep stopped before its time budget (480s)...; 0 dispatch(es) and 7 isolated failure(s) so far." and exited in 9m21s instead of hanging or crashing.

    But the org-wide dispatch budget (1 per 5-minute cycle) is still dominated by a small set of repos (bandscope, pg-erd-cloud, newsdom-api) that get re-dispatched repeatedly, while fast-mlsirm/TEPP/LineageWeave never reach the front of the queue.

    Checked one repeat target closely — bandscope#895: dispatched at least 3 times in ~4 hours (its updatedAt keeps refreshing), yet gh api repos/ContextualWisdomLab/bandscope/pulls/895/reviews shows zero opencode-agent reviews ever posted — only 2 unrelated github-advanced-security comments from 2026-08-17.

    agent-mention-opencode-dispatch.yml only forwards the request to an external service (api.opencode.ai) via a 5-second validate-and-forward job. The actual review-generation happens entirely outside GitHub Actions — opaque to gh/anything I can inspect from here. If that external service is silently failing or never completing for some PRs while the dispatch step still reports success (consuming a budget slot each time regardless), that would explain the queue starvation even with a fully healthy dispatch mechanism.

    This isn't diagnosable or fixable with the tools available to me — flagging for whoever has access to OpenCode's own service-side dashboard/logs to check why bandscope#895 (and similar repeat-dispatch targets) never receive a posted review verdict despite repeated dispatch.

  8. seonghobae commented on Aug 23, 2026

    @seonghobae
    ContributorAuthor

    Root cause for the remaining gap is now identified — it's .github#624, not opaque

    Checkpoint at 2026-08-22T19:19:30Z above diagnosed the remaining gap after #1214/#1235/#1236/#1240/#1241 as an opaque external-service failure: dispatch reaches api.opencode.ai successfully but a posted review never follows for some PRs (e.g. bandscope#895, dispatched repeatedly with zero opencode-agent reviews ever posted).

    Another investigation thread (.github#624, "Migrate OpenCode review model pool off GitHub Models before 2026-07-30 retirement") root-causes exactly this symptom precisely — not opaque, just tracked in a different issue:

    • GitHub Models was fully retired 2026-07-30; org-attributed inference was already cut off from 2026-07-25. The stopgap PAT mitigation stopped working the day of cutoff.
    • OpenAI returns insufficient_quota; OpenRouter returns HTTP 402 (credit-exhausted). With GitHub Models also dead, the entire configured model pool for opencode-review is exhausted org-wide → OPENCODE_MODEL_POOL_OUTCOME: exhausted → MODEL_OUTPUT_UNAVAILABLE. No review can be produced, so no approval evidence, so every main-targeted PR org-wide stays REVIEW_REQUIRED/blocked regardless of how healthy the dispatch/scheduler mechanism is.
    • Migrate OpenCode review model pool off GitHub Models before 2026-07-30 retirement #624's most recent comment (2026-08-23T06:36:30Z) traces this deeper: even the default, non-retired nvidia-nim/nvidia/llama-3.3-nemotron-super-49b-v1.5 candidate fails, but via a different failure mode — CONTROL_REJECTED: no top-level current-run control JSON object was found → NO_CONCLUSION (a content-shape/schema-validation rejection, not a billing/reachability error). This means restoring OpenRouter credit or OpenAI quota alone may not be sufficient if the default NVIDIA NIM candidates keep failing the anti-replay control-JSON schema check first and burn most of the 30-attempt ceiling before ever reaching a paid fallback.

    Combined with the dispatch-mechanism fixes already landed here (#1214/#1235/#1236/#1240/#1241), the full picture is now: the dispatch/scheduler layer is healthy end-to-end; the sole remaining blocker is the model-pool exhaustion + a possible secondary control-JSON schema-rejection bug on the default candidate, both tracked in #624. Not attempting a fix from this thread either — no NVIDIA NIM/OpenAI credentials to reproduce locally, and #624's own investigator already correctly declined to guess-and-check on this deliberately hardened, actively-maintained shared scheduler logic. Recording the cross-link here so anyone reading either issue gets the complete picture without re-deriving it.

  9. seonghobae commented on Aug 23, 2026

    @seonghobae
    ContributorAuthor

    Follow-up to the cross-link comment above: opened #1246 fixing the content-shape/schema-rejection half of #624's newest finding (int-typed run_id/run_attempt in a model's control JSON always failing the anti-replay identity check, regardless of billing). Full evidence (tests proven to fail pre-fix, 100% coverage/docstring, safety reasoning) in the PR. The billing/model-pool-exhaustion half still needs a human with OpenRouter/OpenAI billing access — this PR alone won't fully unblock main-targeted approvals.

  10. seonghobae commented on Aug 24, 2026

    @seonghobae
    ContributorAuthor

    New root cause found: Strix's own benign MODEL QUALITY WARNING startup banner false-fails every clean scan on the org's configured default model

    Independent of the review-approval gap this issue tracks, the required strix check itself was blocking merges via a separate false positive. Strix prints a box-drawn MODEL QUALITY WARNING banner at startup whenever the configured model isn't on its own hardcoded "recommended frontier model" list — a static disclaimer about model choice, unrelated to the scan's outcome. Its literal WARNING text satisfies strix_quick_gate.sh's generic Fatal|Denied|Warn|Warning infra-failure matcher, so any clean, 0-vulnerability scan on the org's configured default model (nvidia_nim/nvidia/nemotron-3-super-120b-a12b, not on Strix's recommended list) fails closed even though the scan itself succeeded with zero findings.

    Reproduced directly: TEPP#214's strix check failed with Strix run emitted provider infrastructure or failure-signal output; failing closed. while its own transcript shows a complete pentest summary reporting "Low" risk and "Vulnerabilities 0". The same banner (6 occurrences across fallback attempts) also appears in fast-mlsirm#1237's strix job log, so this isn't TEPP-specific — it fires wherever the org's default model resolves to a non-"frontier" model, i.e. broadly.

    Fix opened: #1311 — sanitizes the banner out of both the console transcript and report-artifact logs before the infra-failure matchers run (does not touch vulnerability-severity classification; a real finding still fails closed as before). New regression test reproduces the exact TEPP#214 failure on the pre-fix script and passes post-fix.

  11. seonghobae commented on Aug 27, 2026

    @seonghobae
    ContributorAuthor

    Fresh downstream canary after #1002 integration shows a fail-open review-evidence regression that belongs to the central review pipeline, not OriginWeave product code.

    ContextualWisdomLab/OriginWeave#37 at exact head e1fca7c65ce738de3406b0e408708e08c283552e / live base main@f658f329c83a106b68385e17cb714c4147c12f49 has central coverage-evidence check run 98391429201 = success and opencode-review check run 98391443246 = success (workflow run 33033609031). The same-head Reviews API has no qualifying APPROVED or CHANGES_REQUESTED OpenCode verdict; the newest current-head formal review activity is COMMENTED/Devin only. Live organization ruleset 18156473 still requires one approval.

    That is the exact consumer failure #1002 was intended to prevent: a green opencode-review check exists without a substantive exact-head Reviews API verdict. Please treat this as a current regression canary, reproduce it at the first central causal boundary, add a RED contract that forbids success without the exact-head formal verdict, then repair/revalidate the central workflow without adding a leaf OriginWeave workaround or synthesizing approval. The OriginWeave lane will remain non-mergeable-by-policy until real exact-head approval exists.

  12. seonghobae commented on Aug 28, 2026

    @seonghobae
    ContributorAuthor

    Orgmetra exact-head canary: required OpenCode context fails only because no current-head verdict materialized

    Fresh downstream evidence from ContextualWisdomLab/Orgmetra#100 isolates the first failing boundary to the central review-verdict path, after required-workflow bootstrap and coverage evidence have already succeeded.

    • target PR/head: ContextualWisdomLab/Orgmetra#100@57cb9e461bf594e96c0b125349325f35cba95d20
    • live base: develop@9e3e4847510e1e612b48474ba42b177b8ed824df
    • required OpenCode workflow run: 33139299589
    • required-workflow-bootstrap job 98746318056: terminal SUCCESS
    • coverage-source-tree job 98746532800: terminal SUCCESS
    • coverage-evidence job 98746587038: terminal SUCCESS
    • opencode-review job 98746893380: terminal FAILURE
    • first/only failing step in that job: Fail closed without a current-head OpenCode verdict

    This canary does not prove whether the missing verdict is caused by dispatch throughput, mention/router concurrency, model/provider availability, Reviews API publication, or another central control-plane boundary. It does prove that no Orgmetra source/test/coverage repair is justified for this failure: the central required context reaches its explicit no-current-head-verdict fail-closed branch after evidence jobs succeed.

    RED acceptance: an exact target head with successful bootstrap/source-tree/coverage-evidence must never be represented as review-passing without a qualifying formal Reviews API verdict for that same SHA, but the owner path must produce a durable, machine-readable causal failure when that verdict cannot be materialized rather than leaving the leaf to guess or mutate product code.

    Required GREEN: identify and repair the first central dispatch/review-publication boundary that prevented a verdict; preserve reviewer identity, exact-head binding, fail-closed behavior, and coverage requirements; then rerun unchanged/current Orgmetra #100 and require a formal current-head OpenCode verdict plus terminal-success required context. Do not use a status-only substitute, stale review, self-approval, or downstream workflow shim.

    This comment advances the existing central review-pipeline owner path only. No .github source/ref/workflow/settings/PR state was mutated from the Orgmetra loop.

  13. seonghobae commented on Aug 28, 2026

    @seonghobae
    ContributorAuthor

    Fresh Inkspan canary adds a current consumer failure to the existing review-dispatch throughput lane without asserting a provider failure that has not yet been reached.

    Affected target: ContextualWisdomLab/inkspan#391

    • protected live base: 128a239f8b71ca16add4b9e15e21752d1ad63ff0
    • exact current head: 3edd7495a20f96b91eee97afa0f0dd62a8adc55c
    • required-workflow run: 33146959271
    • fail-closed verifier job: 98771103570

    Observed sequence: the required OpenCode workflow started at 06:08:51Z. Its source/coverage placeholder jobs completed successfully, then the opencode-review verifier ran at 06:17:55Z and failed at 06:18:02Z with the exact message No APPROVED or CHANGES_REQUESTED from opencode-agent on the current head. This required check is not a review and must not succeed until the authenticated dispatch posts a current-head verdict. The PR's current review inventory has no opencode-agent verdict for 3edd7495.... A fresh inspection of the 100 most recent central repository_dispatch runs after the verifier failure contained active OpenCode dispatches for other targets (including .github#897@4a5f29f...) but no dispatch titled for ContextualWisdomLab/inkspan#391@3edd7495...; the Inkspan PR comment stream likewise has no durable OpenCode receipt marker for this head.

    First causal boundary for this canary is therefore dispatch admission/throughput/receipt before model execution, not application source and not yet a demonstrated model-provider failure. This is falsifiable: if an authenticated dispatch for this exact target/head exists, record its run id/receipt and move the boundary to that run's first failing job; otherwise the scheduler must ensure an eligible exact-current head is dispatched early enough that the required verifier cannot expire first.

    Smallest acceptance: preserve fail-closed verifier semantics; do not synthesize a review/status. For the unchanged or successor Inkspan head, require a durable authenticated dispatch receipt + central run id bound to exact repository/PR/head, followed by an actual formal opencode-agent APPROVED or CHANGES_REQUESTED review on that same head before the required verifier can pass. Queue/backlog handling must not silently starve a current required head.

  14. seonghobae commented on Aug 28, 2026

    @seonghobae
    ContributorAuthor

    DiskSage exact-head canary for the central OpenCode dispatch/throughput lane (owner handoff only; no foreign source mutation): ContextualWisdomLab/disksage#267 is currently at exact head ccb395ed5b3ceb7ec3180012b06a3d30c2000dbb on base main@79067c1160ddedf7fc962cbf8067ce7e83c4564a. Required-workflow run 33187827737, job 98906719264 (opencode-review) failed closed. The first failing boundary is Fail closed without a current-head OpenCode verdict: the job queried Reviews API for opencode-agent/opencode-agent[bot] on the exact SHA and found no APPROVED or CHANGES_REQUESTED. Fresh /pulls/267/reviews likewise contains no OpenCode review for this head. Same-head central coverage-evidence and coverage-source-tree are successful, so this is not a DiskSage source/coverage failure. Acceptance: authenticated central dispatch submits a non-synthetic OpenCode APPROVED or CHANGES_REQUESTED review bound to ccb395ed…; then the unchanged-head opencode-review gate is rerun and reaches terminal success only for a valid verdict. Please preserve the exact-head/fail-closed semantics; no leaf status synthesis or gate relaxation is requested.

  15. seonghobae commented on Aug 31, 2026

    @seonghobae
    ContributorAuthor

    Fresh downstream canary from ContextualWisdomLab/Orgmetra#55 shows the review-pipeline throughput/liveness problem can now consume the target required-check lifetime even when exact-head coverage is already green and the central scheduler eventually accepts the same-head dispatch.

    Exact evidence:

    • target PR/head: Orgmetra#55 @ 27f09897507182f8ddcb090dd7273dcbaf182e36
    • target required opencode-review job 99631473020: started 2026-08-31T21:33:14Z, polled for an exact-current-head opencode-agent[bot] APPROVED/CHANGES_REQUESTED verdict every 30s for 180 attempts, then failed closed at 2026-08-31T23:04:28Z; exact-head coverage-evidence was already SUCCESS and no qualifying OpenCode formal review appeared
    • central repository_dispatch run 33441942494: created 2026-08-31T21:33:19Z, but scan-pr-queue job 99651996585 did not start until 2026-08-31T22:51:12Z (~78 min of Actions queue delay)
    • once scheduled, the central job did identify the same exact head and logged acceptance of the same-head OpenCode review dispatch at about 22:52:04Z
    • this left only ~12 minutes before the target's original 90-minute polling budget expired; the target therefore failed without a formal verdict even though the control-plane dispatch was eventually accepted

    This is not repairable by an Orgmetra source change or a no-op head churn. The first causal boundary is the central review scheduler/required-check liveness contract: the consumer wait budget starts before the central dispatch worker may actually be scheduled under organization Actions contention.

    RED acceptance: reproduce a same-head consumer where central queue latency consumes most of the target polling window, central dispatch is eventually accepted, but the target gate expires without a formal verdict.

    Required GREEN semantics: preserve fail-closed exact-head review requirements, but decouple the substantive verdict wait budget from pre-dispatch scheduler queue delay (for example by exposing durable accepted-dispatch state/timestamp that the target can observe and budgeting from that boundary, or an equivalent owner design). On the unchanged Orgmetra #55 head above, a post-repair run must obtain a formal current-head OpenCode verdict and let the required target gate reflect that verdict without branch churn, synthetic status, predecessor evidence, or relaxed review requirements.

    #1150 is useful read-only Actions queue-health evidence for this class of delay, but it intentionally does not repair review-scheduler liveness; this comment is routed here because this issue already owns OpenCode dispatch throughput/retry behavior.

  16. seonghobae commented on Aug 31, 2026

    @seonghobae
    ContributorAuthor

    A second fresh Orgmetra canary narrows the same review-pipeline liveness defect further: the fixed target OpenCode wait budget is consumed not only by central Actions queue latency, but by the scheduler's intentional prerequisite ordering before OpenCode can even be dispatched.

    Exact evidence:

    • target: ContextualWisdomLab/Orgmetra#57 @ 6ca554791595d925a76587378b543e7dbc3dc20b
    • target required opencode-review job 99636447752: started 2026-08-31T21:42:06Z, then polled every 30s for 180 attempts for an exact-head opencode-agent[bot] APPROVED/CHANGES_REQUESTED; it failed closed at 2026-08-31T23:13:24Z with no formal current-head verdict. Exact-head coverage-source-tree and coverage-evidence were already SUCCESS.
    • central merge-scheduler repository_dispatch run 33442679292: created about 2026-08-31T21:42:15Z; its scan-pr-queue job 99654383906 did not obtain a runner until 2026-08-31T22:53:25Z (~71 minutes later).
    • the central job then validated the unchanged exact 🛡️ Sentinel: [CRITICAL] Fix HTML Comment Breakout in JSON Serialization #57 head and at 22:53:39Z made this explicit decision: security_dispatch: current head has no completed Strix evidence; same-head Strix dispatched, contract_decision: WAIT.
    • therefore OpenCode was still not eligible/dispatched on that central pass. Instead central Strix run 33448297677 was created at 22:53:41Z for exactly Orgmetra#57@6ca554...; a fresh refetch still reports that Strix run as pending, with no target exact-head Strix check materialized yet.

    This rules out an Orgmetra source repair: the target OpenCode timer is already running while the central owner deliberately sequences missing Strix ahead of OpenCode, and organization queue delay can consume most of that timer before prerequisite security work even starts. With a 90-minute target polling budget, a correct fail-closed scheduler can therefore cause OpenCode failure before substantive OpenCode review has had a viable start boundary.

    RED acceptance: exact head has no completed Strix; target OpenCode waiter starts; delayed central scheduler eventually runs and correctly dispatches Strix-only/WAIT; target OpenCode wait budget expires before a formal OpenCode review can be dispatched and published.

    Required GREEN semantics: preserve exact-head fail-closed security and review requirements, but do not consume the substantive OpenCode verdict budget before prerequisite security evidence and an accepted OpenCode-dispatch phase make that verdict achievable. A durable phase/receipt boundary or equivalent owner design is acceptable; absent/cancelled/neutral/stale Strix and absent/stale/model-only OpenCode evidence must remain non-passing. After the owner repair reaches protected main, the unchanged #57 head should be able to progress through exact-head Strix and then obtain a formal current-head OpenCode verdict without branch churn, synthetic status, predecessor evidence, or relaxed gates.

  17. seonghobae commented on Sep 1, 2026

    @seonghobae
    ContributorAuthor

    Fresh independent downstream canary from ContextualWisdomLab/Orgmetra#56 reproduces the same central review-liveness defect with a prerequisite WAIT, not an Orgmetra source failure.

    Exact evidence:

    • target PR/head: Orgmetra#56 @ 68af42cb80807b6638d745d1687fd3c6a814d64f, live base develop@9e3e4847510e1e612b48474ba42b177b8ed824df
    • target Required OpenCode run 33262099911, attempt 2, job 99630478791: started 2026-08-31T21:21:25Z; authenticated OIDC/app-token dispatch succeeded, then the verifier polled Reviews API every 30s for 180 attempts and failed closed at 22:52:30Z because no exact-head opencode-agent[bot] APPROVED/CHANGES_REQUESTED verdict existed
    • corresponding central scheduler run 33440932661 was created at 21:21:29Z, but scan-pr-queue job 99648671891 did not start until 22:42:07Z (~80m38s control-plane queue delay)
    • once scheduled, it validated the exact Orgmetra repository/PR/head and at 22:42:22Z correctly returned WAIT: current head has no completed Strix evidence; same-head Strix dispatched
    • that prerequisite decision is correct, but it arrived after ~80 minutes of the target's fixed ~90-minute verdict budget had already elapsed, leaving only ~10 minutes before the target verifier failed
    • the newly dispatched exact-head Strix run is central .github run 33447442798; it is currently in_progress, hence non-passing and not a reason to synthesize an OpenCode verdict

    First causal boundary: the required-check lifetime is consumed while the central worker is queued and while exact-head prerequisites are legitimately unresolved. There is no correct Orgmetra-local code change or no-op head churn that can repair this ordering/liveness contract.

    Smallest GREEN acceptance: preserve fail-closed exact-head semantics; record a durable authenticated dispatch/phase receipt bound to repository/PR/head, do not consume the substantive OpenCode-verdict budget before the central worker has actually started and prerequisite state can be evaluated, and require a real formal opencode-agent APPROVED or CHANGES_REQUESTED review on that same exact head before the target required verifier can pass. A prerequisite WAIT, status-only record, or synthetic fallback must remain non-passing.

  18. seonghobae commented on Sep 1, 2026

    @seonghobae
    ContributorAuthor

    Orgmetra #59 independent RED: target waiter exhausted while its accepted central scheduler dispatch never left queue

    Fresh unchanged-head evidence from ContextualWisdomLab/Orgmetra#59@f361012212d4c90e20acb9edaa5302850a340573 isolates the liveness failure before any substantive OpenCode execution.

    Target required workflow 33279409143, attempt 2, job 99688748116 started at 2026-09-01T01:30:03Z. Its authenticated Request current-head OpenCode review execution step succeeded at 01:30:06Z, POSTing the merge-scheduler repository dispatch. The target then polled the Reviews API for a real exact-head opencode-agent APPROVED or CHANGES_REQUESTED verdict for 180 × 30 seconds and failed at 03:02:01Z because none materialized.

    The corresponding central dispatch window materialized Required PR Review Merge Scheduler runs 33458985379 (created 01:30:08Z) and 33458986030 (created 01:30:09Z) on central policy SHA 1186a9f4e5eda7683b23ae63d2c806831743432a. Fresh refetch after the target waiter had already failed still shows both runs queued; in each run the causal scan-pr-queue job (99704928025 / 99704929696) is itself still queued, while the other jobs are skipped. In other words, an accepted target dispatch had no opportunity to inspect exact-head Strix/prerequisite state or dispatch substantive OpenCode before the target's entire ~90-minute verdict window expired.

    This is stronger than a model/provider failure and independently reproduces the phase-budget defect already captured by the #56/#57 canaries: the target verifier's substantive-verdict timeout is not coupled to an actually-started central worker or prerequisite-ready phase. There is no correct Orgmetra-local source repair or no-op head change for this.

    RED acceptance: target exact-head dispatch succeeds; central scheduler run/job remains queued past the target verdict budget; no exact-head formal review can possibly materialize; target fails closed.

    Smallest GREEN semantics remain: preserve fail-closed exact-head Strix and formal-review requirements, but bind the required-check lifetime to durable authenticated phase/receipt state so queue/prerequisite time does not consume the substantive OpenCode verdict budget. A target may pass only after a real exact-head formal opencode-agent verdict exists; queued/pending/WAIT/status-only/model-only evidence stays non-passing. After the owner repair reaches protected central main, validate with a fresh central event on unchanged Orgmetra #59 head rather than rerunning an old trusted workflow snapshot.

  19. seonghobae commented on Sep 1, 2026

    @seonghobae
    ContributorAuthor

    Fresh independent canary from ContextualWisdomLab/Orgmetra#40 removes the earlier coverage ambiguity from this review-delivery lane.

    • target PR exact head: 6917e41f9053fab6f7e99f8185f2137e8fc5fca5; live base: develop@9e3e4847510e1e612b48474ba42b177b8ed824df
    • current exact-head coverage-source-tree check 99422518429: SUCCESS
    • current exact-head coverage-evidence check 99422518803: SUCCESS
    • required OpenCode job 99422517837: FAILURE
    • that job successfully obtained OIDC, exchanged a repository-scoped App token, and dispatched the central scheduler; it then polled exact-head formal reviews for 180×30s and failed only because no authenticated opencode-agent APPROVED or CHANGES_REQUESTED verdict arrived before the window expired
    • the earlier OpenCode COMMENTED/COVERAGE_BLOCKED review is therefore stale as a diagnosis and remains non-authoritative; Orgmetra's current coverage evidence itself is green

    This is not an Orgmetra source/coverage repair candidate. Please use this unchanged-head canary for the scheduler/verdict-delivery acceptance already tracked here: prerequisite/queue time must not consume the substantive verdict window, and GREEN requires a qualifying authenticated formal verdict on the same 6917e41f… head without a no-op consumer commit or gate weakening.

  20. seonghobae commented on Sep 1, 2026

    @seonghobae
    ContributorAuthor

    Fresh independent unchanged-head canary from ContextualWisdomLab/Orgmetra#42 confirms the current OpenCode failure is no longer the earlier coverage-wrapper diagnosis.

    • target exact head: fb03c0837b38424412fa774576a8ded0f9847896; live base: develop@9e3e4847510e1e612b48474ba42b177b8ed824df
    • coverage-source-tree check 99423491065: SUCCESS
    • coverage-evidence check 99423453333: SUCCESS
    • required OpenCode job/check 99423452804: FAILURE
    • the job successfully obtained OIDC, exchanged a repository-scoped App token, and dispatched the central scheduler, then polled the exact head for 180×30 seconds; it failed only because no authenticated opencode-agent APPROVED or CHANGES_REQUESTED formal review arrived
    • the historical COMMENTED/COVERAGE_BLOCKED review remains non-authoritative and its coverage diagnosis is stale on this unchanged head

    No correct Orgmetra source/coverage repair remains at this boundary. Please include #42 as another RED→GREEN acceptance canary for the existing scheduler/verdict-delivery contract: prerequisite/queue time must not consume the substantive verdict window, and GREEN requires a qualifying authenticated formal verdict on this same exact head without a no-op consumer commit, gate weakening, or predecessor evidence.

  21. added
    bugSomething isn't working
    type: bugDefect or incorrect behavior
    on Sep 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area: authAuthentication, authorization, identity, or tenant isolationarea: ci-cdCI, GitHub Actions, checks, release, or supply chainarea: dependenciesDependency or lockfile maintenancearea: securitySecurity boundary, hardening, or vulnerability preventionbugSomething isn't workingpriority: mediumNormal-priority or P2 workstatus: triagedOpen issue has an organization taxonomy assignmenttype: bugDefect or incorrect behaviortype: featureNew or expanded product capability

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions