Repository navigation
fix(strix): bare Fatal/Denied/Warn/Warning match lacks provider-context guard, causing false-positive fail-closed on clean scans #1291
Description
Activity
- addedtype: bugDefect or incorrect behaviorDefect or incorrect behaviorarea: ci-cdCI, GitHub Actions, checks, release, or supply chainCI, GitHub Actions, checks, release, or supply chainarea: securitySecurity boundary, hardening, or vulnerability preventionSecurity boundary, hardening, or vulnerability preventionpriority: highHigh-priority or P1 workHigh-priority or P1 work
on Aug 24, 2026 seonghobae commented
on Aug 24, 2026 ContributorAuthorMore actionsCanonical repair is now published on #1263 exact head
d6c34c59f18e959596ddc18f2c169d312e571758.Root cause was reproduced before the fix: a complete zero-finding fixture containing
OpenAIon one target-documentation line andFatal,Denied,Warn, andWarningon a separate target-taxonomy line exited 1 becausehas_detected_infrastructure_error()treated any bare failure word anywhere in the log as provider evidence.The smallest causal change requires the LLM-provider marker and failure word to occur on the same physical non-allowlisted line. Existing provider-failure controls were authenticated with explicit
litellm,openai, oranthropiccontext; target/repository narrative cannot combine separate lines to spoof the classification.Verification on the exact source tree: RED fixture failed with the historical false-positive output; after the fix, the negative fixture plus Fatal/Warning/Denied positive controls all exited 0 from the focused harness, while the three genuine provider-signal scenarios remained fail-closed as expected. Both shell files pass
bash -nand whitespace checks. Exact-head hosted workflows are newly queued/in progress and are non-passing until terminal.Do not close this issue from source publication alone; #1263 must first produce terminal exact-head evidence and then protected-main/consumer canaries must confirm the org-wide false-positive cascade is gone.
seonghobae commented
on Aug 24, 2026 ContributorAuthorMore actionsAcceptance head advanced to
a73831f60f83a61df95dbcb6999094fab80bb352after the full harness caught and repaired an altered-advisory regression in predecessord6c34c5….The final boundary keeps exact known HF advisory lines allowlisted, but an appended
Fatalsuffix is authenticated by the same-lineHF Hubprovider context and remains fail-closed. The clean target narrative still passes, and explicit litellm/openai/anthropic provider failures still block.Hosted exact-head evidence is pending; issue remains open.
seonghobae commented
on Aug 25, 2026 ContributorAuthorMore actionsFresh downstream canary from
ContextualWisdomLab/bandscope#783confirms this failure mode on the current central Strix source, not predecessor evidence.Exact identities:
- BandScope PR:
#783 - protected base:
acdbea6344fe1231c39535b575f4de35e4c607c9 - exact PR head:
1168c8f4257de5de036ea54bf5ee73edb83e775e - central reusable-workflow SHA:
8fd471a31399a914d9cb22a840f4a4c68e010ea6 - Strix workflow run:
32696787054 - Strix job:
97660756524
The current
strix_required_workflow_smoke.shpasses, so #1317/#1318's retired-fallback smoke mismatch is no longer the first boundary. The primary NVIDIA model then returns HTTP 429 on all three bounded attempts. The first configured fallback,nvidia_nim/nvidia/llama-3.3-nemotron-super-49b-v1.5, subsequently completes an authoritative scan and printsVulnerabilities 0 (No exploitable vulnerabilities detected)plus a completed penetration-test summary. Immediately afterward, however, the gate reportsStrix run emitted provider infrastructure or failure-signal output; failing closedand treats that successful fallback as contaminated, then proceeds toopenai-direct/gpt-5.4, which independently fails with404 page not found. The final required check therefore fails asSTRIX_PROVIDER_UNAVAILABLEdespite the completed zero-vulnerability fallback result.This proves a current attempt-scoping defect: provider-failure text from an earlier model attempt is being allowed to poison a later fallback attempt that itself completed with structured zero-vulnerability evidence. This is central
.githubownership; BandScope cannot correctly repair it without duplicating or weakening the required workflow.Acceptance test for the owning repair:
- primary model emits a classified provider/rate-limit failure;
- a later configured fallback completes and produces authoritative structured/report-bearing zero-vulnerability evidence for the exact PR head;
- earlier-attempt failure markers must not classify that completed fallback as infrastructure-failed;
- findings or provider failure from the fallback itself must still fail closed;
- keep separate regression coverage for the
openai-direct/gpt-5.4endpoint/model path so it does not return 404 when that fallback is actually needed.
No BandScope gate suppression, model-only success, or leaf workaround was introduced.
- BandScope PR:
seonghobae commented
on Aug 25, 2026 ContributorAuthorMore actionsFresh post-#1333 downstream canary: the same BandScope #783 exact head still fails under the newly protected central workflow, so the prior
8fd471a…evidence is now superseded by a newer central-source reproduction.Exact identities:
- BandScope PR/head:
#783@1168c8f4257de5de036ea54bf5ee73edb83e775e - protected BandScope base:
acdbea6344fe1231c39535b575f4de35e4c607c9 - central reusable-workflow SHA actually checked out by the job:
e3b7ece44ba8e891e4c948e9b5b75773f330cd0e(fix(strix): bound retry budget and retain attempt logs (#1333)) - Strix workflow run:
32696787054 - current Strix job/check:
97930578853 - trusted-workflow smoke: PASS
Observed causal chain from the exact job log:
- Primary
nvidia_nim/nvidia/nemotron-3-super-120b-a12bhits 429 and the new bounded retry logic runs through the configured attempts. - The first fallback,
nvidia_nim/nvidia/llama-3.3-nemotron-super-49b-v1.5, then completes a substantive penetration scan and printsVulnerabilities 0 (No exploitable vulnerabilities detected)with an executive summary / completed-scan output. - Despite that fallback-local completion, the gate immediately says
Strix run emitted provider infrastructure or failure-signal output; failing closedand treats the fallback as contaminated. - The next fallback
openai-direct/gpt-5.4independently returns404 page not found. - Final required result is
STRIX_PROVIDER_UNAVAILABLE; the outer retry cannot start another full attempt because the remaining budget (4701s) is below its reserve calculation based on the 5700s attempt budget.
So #1333 improved bounded retrying but did not close the attempt/fallback-evidence isolation defect. A prior primary provider failure can still poison a later fallback that itself reaches zero-vulnerability completion, and the direct-OpenAI gpt-5.4 path remains nonfunctional in this canary.
Acceptance for the central owner remains fail-closed but must be attempt-scoped:
- classify each primary/fallback attempt only from that attempt's own log/report state;
- when a fallback itself completes with an authoritative structured/report-bearing zero-vulnerability result, earlier-attempt provider markers must not downgrade it;
- any finding or provider failure produced by that fallback itself must still fail closed;
- add a live/contract regression for the exact
openai-direct/gpt-5.4mapping/endpoint so a configured fallback does not deterministically 404; - keep retry/budget behavior bounded, but do not require reserve for a full 5700-second scan when the chosen recovery operation has a smaller validated bound.
No BandScope-local suppression or workaround was added.
- BandScope PR/head:
Problem
While triaging why the required
strixcheck was failing on essentially every open PR acrossfast-mlsirm(including a trivial Dependabot Actions-version bump, PR #1311, where no plausible real finding exists), I traced one concrete failure to its root cause inscripts/ci/strix_quick_gate.sh.has_detected_infrastructure_error()(around line 2946) starts with:Unlike every other branch in the same function (
is_timeout_error,is_rate_limit_error,is_llm_token_limit_error,is_llm_api_connection_error,is_llm_service_unavailable_error,is_nvidia_nim_not_found_error, and the genericConnectionError|...branch at the bottom), this first branch has no requirement that the match occur in an LLM-provider context (noLLM_PROVIDER_ONLY_REGEX/PROVIDER_CONTEXT_REGEXco-requirement). It fires on the bare wordsFatal,Denied,Warn, orWarningappearing anywhere in the full Strix log — including inside Strix's own narrative/report text about the scanned target, which very plausibly contains these exact words as ordinary vocabulary (e.g. many CWL repos' ownCLAUDE.md/AGENTS.mdliterally instruct treatingTimeout,Fatal,Warn, orDeniedoutput as a hard failure — a convention Strix would naturally quote or reference while analyzing those repos; or Strix simply describing an access-control check as "Denied" or a log level as "Warn").Evidence
PR
fast-mlsirm#1311(a purebump rust-toolchain/CodeQL-action-version Dependabot PR — no plausible real finding):Strix run emitted provider infrastructure or failure-signal output; failing closed.— withrc=0and a genuine 0-vulnerability report already produced, this can only be the bare-word branch (none of the other, provider-context-gated branches would plausibly fire on a report that already completed with no LLM/provider errors visible in the summary box).nvidia_nim/nvidia/llama-3.3-nemotron-super-49b-v1.5(also flagged the same way) →openai-direct/gpt-5.6-luna, which then hit a second, independent bug: litellm rejected it outright (litellm.BadRequestError: LLM Provider NOT provided ... You passed model=openai-direct/gpt-5.6-luna). The final fallback exhausted, and the gate correctly failed closed on "provider infrastructure failures are not clean scan evidence" — but only because the cascade was triggered unnecessarily by the first false positive.I did not have time to fully trace the second bug (hyphen- vs. underscore-normalized
openai-direct/openai_directmodel-string handling differs betweenstrix.yml:757-758andstrix_quick_gate.sh:2311-2312, and whichever code path fedopenai-direct/gpt-5.6-luna— hyphenated — into this particular fallback attempt evidently skipped the hyphen→underscore conversion that the other path applies). Flagging it here since it compounds the same incident, but the bare-word false positive is the primary, reproducible finding and worth fixing independently.Why this matters
This repeats across nearly every open
fast-mlsirmPR I sampled (confirmed failing on 1237, 1194, 1311; sameblocked/no-fresh-review pattern strongly suggestive of the same cause on 1181, 1172, 1156, 1074, 1056, 1302, 1299, 1196) — i.e., it is very likely the single largest current source of required-check failures blocking merges org-wide, not a per-repo code issue.Suggested fix (not attempted here — this touches fail-closed security logic and the existing multi-thousand-line
test_strix_quick_gate.shsuite, and per this repo's own policy that deserves a dedicated, test-first change, not a drive-by patch)Require the bare
Fatal|Denied|Warn|Warningbranch to co-occur with an LLM/provider-context marker (mirroringLLM_PROVIDER_ONLY_REGEX/PROVIDER_CONTEXT_REGEXused by the other branches in the same function), or scope it to text that appears after the log's own completion marker is absent (i.e., don't apply it once aPenetration test completed+Vulnerabilities Nblock with no in-band provider-error marker has already been observed for that attempt). Add regression fixtures using real Strix report narrative text that legitimately contains these words in a non-infrastructure sense, alongside the existing genuine-infrastructure-failure fixtures, so the fix is provably narrowing false positives without reopening the fail-closed gap that #891 tracks.Related
strix_quick_gate.sh)Agent: Claude
Recorded: 2026-08-24T07:xx (see issue creation timestamp)