Skip to content

Product gaps: review depth, full-scope Strix, evidence-based fast routing #1542

Description

@seonghobae

Product/control-plane gaps preserved from superseded traceability work

Protected main@9b57e4bb95b1a6efe9976a208fe7ca2c0d36dfec has already absorbed the #1507 multi-hour Noema/review-lifecycle changes, so the older #1535 document branch can no longer be merged as written. Preserve its still-valid buyer/control-plane goals here as focused implementation ownership rather than keeping a stale PR alive.

G-Review-Depth

OpenCode and Noema do not yet have a measured review-quality parity contract against strong independent reviewers such as CodeRabbit/Devin. Define a reproducible evaluation set and measure finding recall/precision, severity calibration, false-positive rate, exact-location validity, and repair usefulness on the same PR corpus. Audit opencode.jsonc, review prompts, scripts/ci/noema_review_gate.py, and contextual-orchestrator reasoning/test-time-compute settings. Do not optimize to finding count alone or invent a heuristic threshold; select acceptance criteria from measured data and research, then record the design and ablation evidence in ADR/doctoring and docs/product-technical-gap-baseline.md.

G-Strix-FullScope

Verify the actual trusted Strix scan scope on protected main. The product directive is that the required security review must be capable of whole-codebase analysis, not merely a changed-file slice that can permanently miss pre-existing exploitable paths. Design the full-scope path without weakening exact-head binding, artifact evidence, provider-unavailable fail-closed behavior, or source/credential isolation. Measure runtime and resource demand with real repositories; multi-hour scans are acceptable. Add current authoritative security/operations references and exact-head canaries.

G-FastModelRouting

Review/repair traffic through contextual-orchestrator should prefer models that are both capable and empirically fast for the task, rather than using provider/model-name lists or hand-written rules. Use structured runtime evidence keyed by the real deployment/model-group identity, including latency, terminal availability, capability, cost/privacy constraints, and later-success recovery. Route by measured evidence with ablation, not provider-name/model-id heuristics. The implementation owner may be ContextualWisdomLab/contextual-orchestrator; central .github should consume the versioned policy rather than duplicate it.

Acceptance

  • current authoritative research/standards and provider docs recorded in APA 7 style;
  • reproducible evaluation/telemetry schemas rather than rule-of-thumb constants;
  • current-main implementation PRs in the repository that truly owns each boundary;
  • complete statement/branch/docstring/edge-case coverage for owned code;
  • exact-head required security/review evidence and no routine bypass;
  • docs/product-technical-gap-baseline.md updated with current implementation and measured status.

Queue/capacity work remains separately owned by #1531.

Activity

  1. seonghobae commented on Sep 2, 2026

    @seonghobae
    ContributorAuthor

    Fresh owner-path evidence (2026-09-02): .github#1681 exact head e80fcc3daa72f93adb55116e155031c499dae9b0 makes categorical finding confidence = high|medium|low mandatory in the required Noema gate and derives it from adversarial_validation probe strength. That conflicts with the current Noema owner contract in noema#535@af55fad6fe0eddb1f90795509471bb96efb0576a, which explicitly rejects uncontracted Finding/ReviewVerdict fields, no longer publishes categorical confidence, and prohibits Noema from creating a second confidence policy outside contextual-orchestrator.

    Treat this as G-Review-Depth RED, not a cosmetic schema choice: the same organization review path currently has two incompatible authorities for confidence semantics, and #1681 introduces an uncalibrated ordinal label into a required gate without the measured evaluation/ablation evidence this issue requires.

    Acceptance before #1681 can be admitted:

    1. RED: current contract test must fail if central .github requires a categorical confidence field that the live Noema owner contract rejects/does not publish.
    2. Establish one owner for any confidence/calibration semantics; do not duplicate it in .github and Noema. If confidence is retained, it must be a versioned contract consumed from the owning boundary and backed by the reproducible evaluation set required here (precision/recall, calibration/error, location validity, repair usefulness), not by hand-mapped probe-strength labels.
    3. GREEN: required Noema review succeeds with one compatible finding schema end-to-end; malformed/unknown fields still fail closed; severity admission remains independent; no model/bot confidence is promoted to human approval or security/compliance truth.
    4. Re-run exact-head central tests/coverage/docstrings plus the Noema contract canary on unchanged heads. Queued/stale/predecessor evidence does not transfer.

    No source/docs/ref/PR-state mutation was made to either active writer lane.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestpriority: highHigh-priority or P1 workstatus: triagedOpen issue has an organization taxonomy assignment

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions