Skip to content

fix(noema-review): fail closed with typed blocker on invalid changed-line / malformed JSON model output #1637

Description

@seonghobae

Affected exact evidence

  • leaf repository: ContextualWisdomLab/LineageWeave
  • PR: #897
  • exact leaf head: 7180068eedf40689e267e85478a9ad46981bcb82
  • central reusable workflow source: ContextualWisdomLab/.github@0774e29acd7d4688fa2224f1c6fe6c56a03bbbb6
  • workflow run: 33547273531
  • failing job: 99987788190 (noema-review)

First failing boundary

The review sidecar and repository-scoped reviewer token are successfully provisioned and the target head remains exact-current. The model phase then fails during verdict-envelope validation/repair:

  1. initial model output references reviewed line 1, which is not an exact changed-side line;
  2. repair output is malformed JSON (Expecting ',' delimiter ... char 5224);
  3. the job exits 1. Raw model output is intentionally not written to public logs because a model can echo or hallucinate credential-shaped material.

This is a central review-control/model-output validation failure, not a LineageWeave source finding. Leaf product source must not be changed to satisfy it.

RED acceptance

Add deterministic central tests covering at least:

  • a structurally valid verdict that cites a line outside the exact changed-side line set;
  • repair output that is truncated/malformed JSON;
  • combined invalid-line + malformed-repair output;
  • preservation of exact repository / PR / expected_head_sha binding throughout both phases.

The test must demonstrate that invalid model evidence cannot become a synthetic source finding or a passing review.

Smallest acceptable behavior

For invalid source-line evidence, attempt only bounded schema/line repair. If repair remains invalid or unparsable, emit a typed model/infrastructure ABSTAIN / UNAVAILABLE (or existing equivalent) decision envelope that is distinct from semantic source findings. Preserve fail-closed merge readiness: the required check must not become green merely because semantic review was unavailable unless the central governance contract explicitly defines and separately satisfies a safe unavailable-review path.

Do not log raw model output. Do not broaden reviewer credentials. Do not weaken exact-head / live-source checks.

GREEN acceptance

  • valid model verdict JSON parses and every line-bound finding maps to an actual changed-side line;
  • invalid-line or malformed-repair output is represented as typed review-unavailable/model-output evidence rather than an invented source-code defect;
  • current exact-head binding is rechecked before publication;
  • central tests cover malformed/truncated output without logging the raw payload;
  • a fresh LineageWeave fix(security): fail closed on unavailable dependency review #897 exact-head Noema run after the central fix no longer fails for this model-envelope failure class.

Activity

  1. seonghobae commented on Sep 2, 2026

    @seonghobae
    ContributorAuthor

    Wardnet confirms the same central Noema failure class on a second repository/current-head candidate. Exact evidence: ContextualWisdomLab/wardnet#138@e6f05d77858e91c176cff25c4b11e790bc5dcdd1, reusable Noema run 33505617580, job/check 99848830883. The job checked out and reviewed the exact Wardnet head. The initial model envelope was parseable but cited src/lib.rs:1838, which is not an exact changed-side line for this PR; bounded repair then returned malformed/truncated JSON and failed with Expecting ',' delimiter: line 1 column 7909 (char 7908). No independently source-backed Wardnet finding was established. This should therefore remain model-output/infrastructure review-unavailable evidence, not be converted into a synthetic source defect.

    RED extension: include a fixture where the initial verdict is structurally valid JSON but cites an unchanged line, followed by malformed repair JSON, while repository/PR/head binding is otherwise exact. GREEN: unchanged wardnet#138@e6f05d... receives either a valid exact-changed-line verdict or a typed ABSTAIN/UNAVAILABLE envelope after the central repair; exact-head binding remains intact and no fabricated source finding is emitted.

  2. seonghobae commented on Sep 2, 2026

    @seonghobae
    ContributorAuthor

    Fresh Wardnet consumer evidence for this exact failure class:

    • consumer: ContextualWisdomLab/wardnet#138
    • exact head: e6f05d77858e91c176cff25c4b11e790bc5dcdd1
    • central trusted workflow source: ContextualWisdomLab/.github@5768f2bd29b0856ad49f18a0fda72b871eb95b46
    • Noema job: 99848830883
    • contextual-orchestrator sidecar pin: 8cd99f139915131ba0239bce12a5d6a5fd85394e

    Runner acquisition, repository-scoped reviewer token, exact-head validation, sidecar health, and provider-route preflight all succeeded. After roughly 92 minutes in the model/review stage, the terminal deterministic failure was:

    Noema LLM response was not valid JSON (Expecting ',' delimiter: line 1 column 3475 (char 3474))

    The response length was 3474 characters and its SHA-256 was 83410c1e444e3418; raw model output was correctly not logged.

    Wardnet repository CI/Fuzz/Security/SAST are terminal-success and current review threads are resolved on this same exact head. This is therefore a second independent consumer confirming #1637's central malformed-model-output class, not a Wardnet source finding.

    RED: deterministic malformed/truncated JSON output that survives a long model call and fails the current verdict parser.

    GREEN: emit a typed MODEL_OUTPUT_INVALID / review-unavailable decision within a bounded budget; never synthesize a source finding; never log raw model output; preserve exact repository/PR/head/live-source binding. After the central repair is protected, rerun unchanged wardnet#138@e6f05d77858e91c176cff25c4b11e790bc5dcdd1 and require schema-valid Noema evidence.

  3. seonghobae commented on Sep 2, 2026

    @seonghobae
    ContributorAuthor

    Fresh Wardnet Noema evidence adds a second terminal mode on the clean runner-only lineage. Historical trigger PR #149 was later closed as contaminated and superseded by current clean replacement #153, but the failed review job is still useful as a central control-plane reproduction because it bound the exact clean source head b663f9d200e5f385c7dd067d074940a02836c68e before #149's branch moved.

    Exact central run 33564088411, job 100043156284, trusted .github workflow source 7683f1da91f1fc9e046660169f1f7ac4aabcc3c6: runner acquisition, repository-scoped App token, exact-head validation, target visibility, contextual-orchestrator sidecar and gateway preflight all succeeded. The model phase then failed after ~58 minutes with:

    initial failure: Noema LLM response was not valid JSON (Expecting ',' delimiter: line 1 column 1913 (char 1912)); response length=1912, sha256=5e8cb4a165346f60; repair failure: NoemaRepairDeadlineExceeded: Noema repair exceeded 900-second absolute wall-clock deadline.

    Raw model output was correctly not logged. This is not Wardnet source evidence. It extends #1637's RED family from malformed repair JSON to malformed initial envelope + bounded repair transport deadline exhaustion.

    RED: schema-invalid/truncated initial verdict where bounded repair does not return a valid envelope before the absolute repair deadline. GREEN: preserve exact repo/PR/head binding and raw-output non-disclosure, but terminate with a typed MODEL_OUTPUT_INVALID / REPAIR_DEADLINE_EXCEEDED review-unavailable envelope rather than an opaque source-review failure; no synthetic finding and no passing semantic review. Current consumer is clean replacement wardnet#153@b663f9d...; after the central repair reaches protected truth, re-run that exact PR/head and require a schema-valid verdict or the typed unavailable contract.

  4. seonghobae commented on Sep 2, 2026

    @seonghobae
    ContributorAuthor

    Fresh LifeOS consumer reproduces #1637 on a distinct current head.

    • consumer: ContextualWisdomLab/life-os#218
    • exact head: 71460a518d7fd0fdd79a811730a5b3dd8a83aee5
    • central Noema run/job: 33544475940 / 99978486877
    • trusted .github workflow source: 09090e98ca6fe34ee0ab8a0a8145d7af0a0bb9aa
    • contextual-orchestrator sidecar pin: 8cd99f139915131ba0239bce12a5d6a5fd85394e

    Runner acquisition, repository-scoped reviewer token, exact-head validation, target visibility, sidecar provisioning, route preflight, and a live gateway chat/completions preflight all succeeded. The first terminal boundary was Prepare Noema model verdict: after ~81 minutes the model phase exited because the response was malformed JSON (Expecting ',' delimiter: line 1 column 4338 (char 4337)), response length 4337, sha256 prefix 6746fddec1ac963d; raw output was correctly withheld from public logs. Publication was skipped.

    This is not a LifeOS source finding. LifeOS repo CI, AppGuardrail, SAST, Commercial Readiness and the deterministic runner-contract checks are independently green on this head; Security Scan has a separate already-handoffered dependency-review availability incident.

    The same sidecar log also records non-terminal upstream evidence (request_too_large 413 during discovery/probe, Bytez discovery HTTP 500, several route 404/timeouts) while still reaching a healthy gateway preflight. Those signals should not be conflated with the first terminal Noema boundary above.

    GREEN consumer acceptance: after the central malformed-output repair reaches protected truth, rerun the then-current unchanged/successor LifeOS #218 head and require either schema-valid exact-head Noema evidence or the explicitly typed fail-closed review-unavailable envelope; never synthesize a source finding and never expose raw model output.

  5. seonghobae commented on Sep 2, 2026

    @seonghobae
    ContributorAuthor

    Fresh Wardnet reproduction of the same central Noema model-envelope failure class:

    • leaf: ContextualWisdomLab/wardnet
    • PR at trigger: #149
    • exact leaf head: b663f9d200e5f385c7dd067d074940a02836c68e
    • protected base: main@cc15cc2c34daf8c104eeb83d52a6a66f3cd6e128
    • central reusable workflow source: ContextualWisdomLab/.github@7683f1da91f1fc9e046660169f1f7ac4aabcc3c6
    • run: 33564088411
    • job: 100043156284 (noema-review)
    • hosted runner did acquire correctly: ubuntu-24.04, runner GitHub Actions 1001603712; exact-head validation, repository-scoped App credential, target visibility, and contextual-orchestrator sidecar provisioning all succeeded.

    First causal boundary is again the model verdict phase, not Wardnet source or runner acquisition. The sidecar eventually had 4 ready routes and the gateway preflight succeeded. Prepare Noema model verdict then failed after ~58 minutes with:

    Noema bounded repair transport was exhausted; initial failure: Noema LLM response was not valid JSON (Expecting ',' delimiter: line 1 column 1913 (char 1912)). ... response length=1912 chars, sha256=5e8cb4a165346f60.; repair failure: NoemaRepairDeadlineExceeded: Noema repair exceeded 900-second absolute wall-clock deadline

    Raw output remained correctly suppressed from the public log. This confirms #1637 is not LineageWeave-specific: malformed/truncated model JSON plus bounded-repair exhaustion recurs on an unrelated Wardnet control-plane-only diff. Leaf source must not be altered to satisfy it.

    Please extend the RED/GREEN acceptance already recorded here with a deterministic malformed initial envelope whose bounded repair consumes the absolute deadline. GREEN should classify exhausted malformed-envelope repair as typed review-unavailable/model-output evidence, preserve fail-closed merge readiness and exact-head binding, avoid raw model logging, and terminate within a bounded control-plane budget rather than consuming a full long-running reviewer slot before emitting the typed outcome. Wardnet-side revalidation is a fresh exact-head Noema run on the unchanged candidate after the central fix.

  6. seonghobae commented on Sep 2, 2026

    @seonghobae
    ContributorAuthor

    OriginWeave adds a distinct current-head Noema model-evidence failure to this owner class.

    Exact consumer evidence:

    • ContextualWisdomLab/OriginWeave#82@90c4e9a8f31eb94eb46243343160d3f6c96921ef
    • required Noema run 33497088638, attempt 2; job 99943871679
    • trusted central source actually materialized by that rerun: .github@1ddc31fb341a75ddafe8516b86c5d52e26669933
    • contextual-orchestrator pin: 8cd99f139915131ba0239bce12a5d6a5fd85394e

    The runner, repository-scoped cwl-noema-review App token, exact-head validation, target visibility, contextual-orchestrator sidecar, provider-route preflight, and a real orchestrator/free chat/completions preflight all succeeded. The sidecar reported 54 admitted free routes and 4 ready probes. The model/review phase then terminated with:

    Noema request_changes requires a confirmed probe on a published finding

    No review envelope was published. This is not an OriginWeave source finding and must not be converted into one. It is a deterministic rejection of model evidence after the model transport was already proven available.

    There is also an important source-binding detail: protected .github/main has advanced since the run's immutable trusted source. Commits affecting scripts/ci/noema_review_gate.py after 1ddc31f... include 430a007515e8510ded466636b65a34e06864ed0e (fix(noema): bound and classify malformed-verdict repair) and bb14b014eee31e6abdb5d2fffbb805aa29420eac (fix(noema): report actual rejected review location). Rerunning the old Actions run cannot validate those changes because the workflow continues to materialize its original workflow_sha.

    RED extension: an exact-bound verdict chooses request_changes but fails the required finding/probe publication invariant while model transport itself is healthy. GREEN: the current owner path must either repair/validate that model evidence into a schema-valid exact changed-side finding with the required confirmed probe, or terminate as typed MODEL_OUTPUT_INVALID / review-unavailable evidence without synthesizing a product defect. After protected central truth is current, re-trigger unchanged OriginWeave#82@90c4e9a8... through a fresh PR event so the new run materializes the then-current central workflow source; require a schema-valid published Noema verdict or the typed unavailable contract.

  7. seonghobae commented on Sep 2, 2026

    @seonghobae
    ContributorAuthor

    Fresh fast-mlsirm consumer evidence reproduces this central Noema malformed-envelope class without a leaf source finding.

    • consumer: ContextualWisdomLab/fast-mlsirm#1710
    • exact head: 2372f44856d9955b6e390840ae75069a62e24841
    • Noema run/job: 33512966400 / 99872901588
    • trusted central workflow source: ContextualWisdomLab/.github@c70b081dd93cf9ca53c2277ba95eab0e200cbe5c
    • contextual-orchestrator sidecar pin: 8cd99f139915131ba0239bce12a5d6a5fd85394e

    Runner acquisition, repository-scoped reviewer App token, live exact-head validation, target visibility, sidecar startup, route discovery, and gateway chat/completions preflight all completed. The orchestrator/free preflight found 5 ready routes out of 12; some Bytez/NIM discovery/probes were unavailable, but the gateway itself reached a successful preflight before Noema review began.

    The model/review stage then ran for about 78 minutes and failed only at the verdict-envelope parser:

    Noema LLM response was not valid JSON (Expecting ',' delimiter: line 1 column 2189 (char 2188)); response length=2188, sha256=6ba623098ff39512

    Raw model output was correctly withheld from the public log. The same exact fast-mlsirm head has repository CI/security/package evidence independently succeeding; this failed Noema job is therefore review-control/model-output evidence, not a psychometric or source-code defect.

    RED extension: malformed/truncated initial verdict after a successful orchestrator/free gateway preflight and long reasoning/model call, with exact repository/PR/head binding otherwise intact.

    GREEN: preserve fail-closed review readiness, raw-output non-disclosure, and exact-head binding, but classify this deterministically as typed MODEL_OUTPUT_INVALID / review-unavailable evidence and run only the bounded structured-repair contract. A model-format failure must not become a synthetic source finding or trigger a leaf source edit. After the central repair reaches protected truth, rerun unchanged fast-mlsirm#1710@2372f44856d9955b6e390840ae75069a62e24841 and require either a schema-valid verdict or the typed unavailable envelope.

  8. seonghobae commented on Sep 2, 2026

    @seonghobae
    ContributorAuthor

    Additional downstream exact-head evidence from ContextualWisdomLab/fast-mlsirm#1697 shows the same central Noema model-output failure class with a different repair terminal boundary.

    • leaf exact head: 00f5c7b1308f89aba5ffbe1663a4a715abde6a6a; protected base observed by the run: main@45627700c26c29bca150896a9519a9b7426acb56.
    • central workflow source: .github@0774e29acd7d4688fa2224f1c6fe6c56a03bbbb6.
    • run/job: 33397538166 attempt 3 / 99989735554 (noema-review). Runner acquisition, exact-head validation, repository-scoped App credential minting, contextual-orchestrator sidecar provisioning, and orchestrator/free preflight all succeeded.
    • first causal model boundary: Noema reviewed line 3 is not an exact changed-side line.
    • bounded repair then consumed the configured absolute repair budget and failed as NoemaRepairDeadlineExceeded: Noema repair exceeded 900-second absolute wall-clock deadline; the combined step eventually failed after ~57 minutes. No verdict was published. Repository-owned CI, CodeQL, Semgrep and Security Scan are terminal success on the same exact head.

    This confirms invalid changed-line evidence can terminate either via malformed repair output (the original issue evidence) or repair-deadline exhaustion. The central RED matrix should include the latter explicitly. GREEN should preserve fail-closed merge readiness while classifying both outcomes as typed model/review-unavailable evidence, keep exact head/path/side/line binding, and prevent a bad citation from consuming the long outer review budget after the bounded repair path is already exhausted. No leaf source change or gate weakening is appropriate.

  9. seonghobae commented on Sep 2, 2026

    @seonghobae
    ContributorAuthor

    Fresh same-class Wardnet evidence should be added to this central repair rather than patched in the leaf.

    • affected leaf: ContextualWisdomLab/wardnet#93
    • exact current head: e8a2b722c3c7dcbe0569130ccf937c6f2f755c12
    • protected live base: main@cc15cc2c34daf8c104eeb83d52a6a66f3cd6e128
    • required Noema run/job: 33513271348 / 99873911395
    • trusted central workflow source: .github@c70b081dd93cf9ca53c2277ba95eab0e200cbe5c

    The runner was acquired (runner_id=1001602227, Ubuntu 24.04), exact PR/head validation succeeded, the repository-scoped reviewer token was minted, and the contextual-orchestrator sidecar/preflight completed. The first terminal boundary was the Noema model/verdict phase after ~19m33s:

    Noema reviewed line 1 is not an exact changed-side line

    This is not a Wardnet source finding. On the same exact head, repository CI, Fuzz, Security Scan, and SAST Semgrep are terminal success, CodeQL/Semgrep/OSV/dependency/Scorecard evidence is success, and every current inline review thread is resolved. The required coverage-evidence lane remains queued independently, so #93 is still non-passing and is not a bypass candidate.

    Please include this Wardnet reproduction in #1637 RED/GREEN coverage: invalid changed-line model output must become typed model-output/review-unavailable evidence, not a synthetic line-1 source defect. Preserve fail-closed merge readiness, exact (repository, PR, expected_head) binding, raw-output non-disclosure, contextual-orchestrator-only routing, and bounded repair attempts. After the central fix reaches protected truth, unchanged Wardnet #93 should be rerun against the same exact head before any merge classification.

  10. seonghobae commented on Sep 2, 2026

    @seonghobae
    ContributorAuthor

    Fresh Wardnet current-PR reproduction (2026-09-03 KST) extends the same central malformed-model-output class without any leaf source finding.

    • consumer: ContextualWisdomLab/wardnet#155
    • exact current head: e6f05d77858e91c176cff25c4b11e790bc5dcdd1
    • live protected base: main@cc15cc2c34daf8c104eeb83d52a6a66f3cd6e128
    • Noema run/job: 33590351206 / 100122902109
    • trusted workflow source: ContextualWisdomLab/.github@29b931e139ba12319de98f629cdae58479574bc5
    • contextual-orchestrator sidecar pin: 045d17da5e2aea56a97e241ee158ab1628d78660

    Exact-head validation, repository-scoped cwl-noema-review App token, target visibility, sidecar health and gateway preflight all succeeded. The sidecar admitted orchestrator/free; 3 routes were ready and the gateway preflight returned finish_reason=stop. The semantic phase then failed after bounded repair because both model envelopes were malformed JSON: initial Expecting ',' delimiter: line 1 column 4165 (char 4164), length 4164, sha256 prefix e49eca9089a93c63; repair Expecting ',' delimiter: line 1 column 4072 (char 4071), length 4071, sha256 prefix f32c49d2aa43395e. Raw model output was correctly withheld from public logs.

    This same Wardnet exact head has terminal-GREEN repository CI/Fuzz/Security/SAST evidence; the Noema failure occurs after runner acquisition and sidecar/provider preflight, so it is central review-control evidence rather than a Wardnet source defect.

    RED addition: valid exact repo/PR/head binding + healthy orchestrator/free preflight followed by malformed initial JSON and malformed repair JSON. GREEN: preserve raw-output non-disclosure and exact-head binding, but represent exhausted repair as typed MODEL_OUTPUT_INVALID / review-unavailable evidence; do not synthesize a source finding or semantic success. After the central repair reaches protected .github truth, rerun the unchanged then-current Wardnet PR/head and require schema-valid Noema evidence or the typed unavailable contract.

  11. seonghobae commented on Sep 3, 2026

    @seonghobae
    ContributorAuthor

    Fresh Naruon consumer evidence reproduces this central failure class on a product head whose repository-local product gates are already materially GREEN.

    • consumer: ContextualWisdomLab/naruon#1502
    • exact head: c6ed2e6f9d4f8667d9003bdd2df1085f8a3f9aa5
    • required Noema job/check: 100251945105 (run 33587140022)
    • trusted central workflow source used by that run: ContextualWisdomLab/.github@9330d41c92b1e6ab35261f3f5189936ea1ad8bff
    • contextual-orchestrator sidecar pin used by the run: 045d17da5e2aea56a97e241ee158ab1628d78660

    Runner acquisition, repository-scoped GitHub App token minting, exact live-head validation, sidecar health, provider-route preflight, and the orchestrator/free chat/completions preflight all succeeded. The terminal failure was model-envelope/control-plane only:

    initial failure: Noema LLM response was not valid JSON (Expecting ',' delimiter: line 1 column 4822 (char 4821)); response length=4821, sha256=40e401074013161d; repair failure: NoemaRepairDeadlineExceeded: Noema repair exceeded 900-second absolute wall-clock deadline.

    Raw model output was correctly not logged. On this same Naruon exact head, repository-local Application CI, SAST Semgrep, Dependency Review, Bandit, Security Scan, and Docker validation are terminal-success. coverage-evidence is still queued separately, so this is not a merge-ready claim.

    Important chronology: this failed Noema run used central workflow source 9330d41…, which is an ancestor of the now-merged .github#1672 repair (a28fc2f4…). Protected current .github/main contains the permanent contract that forbids NOEMA_REPAIR_DEADLINE_SECONDS, _repair_wall_clock_deadline, NoemaRepairDeadlineExceeded, and signal.setitimer. Therefore do not modify Naruon product source to satisfy this historical caller-owned repair-deadline failure. The next useful owner-side action is to dispatch/revalidate the unchanged naruon#1502@c6ed2e6… against current protected central Noema truth; require either a schema-valid exact-head verdict or the current fail-closed structured-output-unavailable behavior, with no synthetic product finding and no raw-output disclosure.

  12. seonghobae commented on Sep 3, 2026

    @seonghobae
    ContributorAuthor

    Wardnet supplies a fresh same-class exact sample from required Noema review, with the leaf source ruled out before the model-output boundary. ContextualWisdomLab/wardnet#155 exact leaf head e6f05d77858e91c176cff25c4b11e790bc5dcdd1; Noema run 33590351206, job 100122902109; central workflow source ContextualWisdomLab/.github@29b931e139ba12319de98f629cdae58479574bc5; contextual-orchestrator sidecar source 045d17da5e2aea56a97e241ee158ab1628d78660.

    The job acquired ubuntu-24.04, repository-scoped cwl-noema-review[bot] credentials were minted, the sidecar was provisioned, health/provider-route preflight completed, and gateway chat/completions preflight succeeded. Policy evidence used pool=orchestrator/free with priced_selected_count=0; three probed routes were ready and the gateway itself reported ready. The first failing boundary was verdict-envelope parsing/repair: initial output failed JSON parsing at char 4164 (length=4164, logged digest prefix e49eca9089a93c63); bounded repair also failed JSON parsing at char 4071 (length=4071, logged digest prefix f32c49d2aa43395e). Raw model output was correctly withheld from public logs. Job then exited 1.

    This is not a Wardnet source finding, provider-key absence, runner-admission failure, or paid-fallback issue. It confirms #1637 must cover the malformed-initial + malformed-repair path even when CO/gateway preflight is healthy. GREEN for this sample: preserve exact (repository, PR, expected_head_sha) and model-output hash/length evidence; represent exhausted schema repair as typed review-unavailable/model-output evidence distinct from semantic findings; never synthesize source defects/pass from malformed output; keep raw model output out of public logs; and rerun Wardnet #155 only on its unchanged live head after the central repair.

  13. seonghobae commented on Sep 3, 2026

    @seonghobae
    ContributorAuthor

    Fresh leaf evidence from ContextualWisdomLab/naruon#1450 should be treated as predecessor-control evidence, not a Naruon source finding.

    • leaf exact head: 3f8b0cdafcba27d595fd13c125457bbf4b8df484
    • central Noema run/job: 33607709102 / 100175403140
    • central workflow source captured at trigger: .github@5c561a65cca3b925d533e4b40c5c3ac00f16524e
    • exact-head validation, repository-scoped reviewer token, contextual-orchestrator sidecar, orchestrator/free route and chat/completions preflight all succeeded
    • model verdict then failed JSON parsing (Expecting ',' delimiter, response length 3265); the old caller performed its second repair call and terminated with NoemaRepairDeadlineExceeded at 900 seconds

    .github#1672 is now merged (a28fc2f4e185df7847e2f2f5f6ec561d1e84805) and removes that duplicate repair call/deadline. Therefore rerunning the old run/job would not be valid current-owner verification because GitHub reruns preserve the original required-workflow source captured for the trigger. The leaf PR has been lowered to Draft rather than creating a no-op source commit or transferring predecessor evidence.

    GREEN for this leaf class requires a fresh PR event on the unchanged leaf delta after the current canonical owner source is eligible, with typed model-output/unavailable behavior and no resurrection of a leaf-owned model deadline.

  14. seonghobae commented on Sep 3, 2026

    @seonghobae
    ContributorAuthor

    Naruon reproduces the same canonical Noema model-envelope failure on an unchanged exact product head, so no consumer source edit is warranted.

    • consumer: ContextualWisdomLab/naruon#1497
    • exact head: 152d1998c4e8024be9dc7026c8789d343c884fd0
    • protected consumer base: develop@042b0c70531b229af3acbd0421a2f23098d848b3
    • required Noema run/job/check: 33462446295 / 99928013216
    • contextual-orchestrator sidecar pin: 8cd99f139915131ba0239bce12a5d6a5fd85394e

    Runner setup, repository-scoped reviewer token, exact-head validation, contextual-orchestrator sidecar startup and orchestrator/free preflight all succeeded. The model/review phase then terminated with malformed JSON: Expecting ',' delimiter: line 1 column 3541 (char 3540), response length 3540, SHA-256 2ad0e386c189243f. Raw model output was correctly not logged.

    The same Naruon exact head has repository-owned Application CI, Security Scan, Dependency Review, Semgrep, Bandit, Docker and coverage evidence passing, and all current inline review threads are resolved. The Noema result is therefore review-control/model-output evidence, not a provenance-portability source finding.

    The consumer PR has been returned to Draft and records this owner-path blocker. Do not mutate Naruon product source to compensate. GREEN remains a fresh exact-head trigger after the central protected repair, yielding either a schema-valid Noema verdict or the typed fail-closed review-unavailable contract required by #1637.

  15. seonghobae commented on Sep 11, 2026

    @seonghobae
    ContributorAuthor

    Fresh protected-consumer canary from ContextualWisdomLab/Orgmetra#298@161468e3d17b3e1964f570ec9bf286ec4045bc8d: Required Noema run 34550250457, job 103112709717 passed exact-head admission, trusted-source materialization, repository-scoped credential, live-head validation, target visibility and CO sidecar provisioning, then failed at Prepare Noema model verdict and skipped verdict publication. Artifact noema-sidecar-evidence id 10181255682, digest sha256:7f329934d44312f6bcd4e7715649e8a45b71cee847f1b0eb7b933d52320b19e6, terminates with request_failed status=502 code=invalid_structured_output.

    This is another exact case where malformed/invalid structured reviewer output must remain typed infrastructure/model-output evidence, not become a leaf source finding or synthetic pass. The artifact also proves the gateway itself was ready (gateway.status=ready) and two routes passed preflight, so leaf recovery code should not churn to satisfy this failure. Please include this canary in #1637 GREEN acceptance: fresh exact-head Noema should either publish a schema-valid changed-line-bound verdict or fail closed with the typed structured-output outcome after bounded repair, with no raw model output in public logs.

  16. seonghobae commented on Sep 12, 2026

    @seonghobae
    ContributorAuthor

    Current recurrence for fast-mlsirm #1836, exact head c7cb1238fbbd0abb7417e31665ccd4f23588532d: Noema job 103581932596 terminated at 2026-09-12 16:34:01 UTC with:

    Noema model output failed local validation: Noema request_changes requires a confirmed probe on a published finding; caller attempts=1, duration=294.7s, phase=validating, served_model=deepseek-ai/deepseek-v4-pro-0813.

    The consumer diff adds one assertion to an existing Rust test and renames its local result. The runner log publishes no candidate finding body/probe payload. We therefore cannot turn this validation failure into a source defect or claim that an unseen candidate was disproved. Existing exact-head product/security checks pass separately; no synthetic review approval was submitted.

    Please include the missing-confirmed-probe verdict envelope in the existing #1637 validation/recovery acceptance. Keep raw model output private, preserve current-head/finding/probe binding, and expose a typed unavailable outcome when valid evidence cannot be produced. Investigation target: 2026-09-13 UTC; next review: 2026-09-14 UTC, because the review-control failure affects a current consumer repair. This is a recurrence record, not a deployed fix or permission to weaken the review contract.

  17. seonghobae commented on Sep 13, 2026

    @seonghobae
    ContributorAuthor

    Recurrence: fast-mlsirm #1834, 2026-09-13

    Exact head: 05d17c8124b4be36c16efc23160e88a8dd18e485.
    Noema job 103667978670 failed at 04:14:50 UTC during local output validation:

    Noema approve cannot contain a confirmed adversarial probe

    Telemetry reports caller attempts=1, duration=127.6s, phase=validating. This is an inconsistent verdict envelope, not a qualifying approval. The unpublished candidate payload has not been inspected; this report does not assert that an underlying source finding is absent or confirmed. No leaf source change or duplicate dispatch was made to mask the failure.

    Please include approve-plus-confirmed-probe in the existing invalid-envelope regression and preserve exact-head binding and fail-closed validation. The consumer remains awaiting valid review evidence. Internal follow-up: 2026-09-14, to assess the canonical repair path and available validation evidence; this is a review date, not a promised resolution date.

    Consumer: ContextualWisdomLab/fast-mlsirm#1834

  18. seonghobae commented on Sep 13, 2026

    @seonghobae
    ContributorAuthor

    Additional exact-head invalid-envelope case, fast-mlsirm #1684 at 105b100822a0dd1ee46a4a2ef4945714d7aad9a1:

    Raw Noema job 100261580342, run 33634421483, failed on 2026-09-03T00:24:23Z with:

    Noema model-output repair remained invalid; initial failure: Noema LLM response findings must be a list of objects; repair failure: Noema approve requires adversarial_validation.status=passed

    This adds the missing/not-passed adversarial status after a repair of a malformed findings field to the existing envelope cases. It is not approval and does not establish that an unpublished finding is absent. Earlier provider timeout/HTTP errors in the same log are separate from the final validation failure. No cancelled review was restarted and no leaf head was changed.

    The consumer's entire final diff updates two packaging locks to26.3. Official PyPI26.3 metadata matches both proposed SHA256 digests and requiresPython>=3.9; current-head Python3.12/3.14/package/security successes remain separate evidence. No bypass decision is made from those results.

    Owner next action: preserve the initial and repair validation categories with sanitized exact-head evidence; verify malformed findings plus missing adversarial status through the existing Noema envelope regression. Internal next review:2026-09-14 KST, aligned with the consumer's pending release review; not a provider fix deadline. Reuse this issue rather than opening a duplicate.

  19. seonghobae commented on Sep 13, 2026

    @seonghobae
    ContributorAuthor

    Current-head recurrence, 2026-09-13: contextual-orchestrator PR #1145 at 3f1ac5b89644888e4b0b5bcc99bd1bdfb939c0ed, Noema run 34750904362, job 103707389778.

    The sidecar reached model execution. Validation then failed: Noema request_changes requires a confirmed probe on a published finding; caller attempts=1, phase=validating, duration=280.5s. This is a missing review-evidence contract, not a demonstrated source finding or bootstrap outage. Provider attempt errors also occurred, but do not replace this terminal cause. The existing validation guard must remain fail-closed; no model payload is reproduced here.

    The associated warning reports a failed gateway attempt with gateway-owned repair/failover. The evidence artifact uploaded successfully (10315831957; SHA-256 732e7d3b6594e1ccbc716701c2e44de2ca8a9aab867cfa4404f4d8224d5bb719). No rerun or synthetic approval was requested.

    Next owner action: central Noema review-control and CO gateway owners should reproduce missing-confirmed-probe envelopes, compare existing repair/typed-unavailable handling, and verify exact-head publication cannot convert invalid evidence into either a finding or approval. Investigation checkpoint: 2026-09-14, because this actively prevents review acceptance. Completion requires behavioral regression plus a valid exact-head hosted review; issue tracking alone is not resolution.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingpriority: highHigh-priority or P1 work

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions