Skip to content

Evaluate an organizational control plane for autonomous coding harnesses #574

Description

@shiny-code-bot

Objective

Evaluate whether a human and an external supervisory AI can productively oversee complete autonomous coding-harness sessions as organizational units, while each worker harness retains its own agents, providers, validation, and review systems.

Finish Line

Offline evidence determines whether an external read-only supervisory AI adds useful early drift detection without excessive intervention noise, attention cost, or token overhead.

Current Status

State: Intentionally parked research direction with no automatic activation trigger. It is not part of the Codex Lab Private Dogfood milestone and must not delay release.

The concept remains preserved here and related to #121, #298, and #150. Closing the private-dogfood milestone or accumulating sessions does not start this work by itself. After release, the user must explicitly promote this issue before any experiment, implementation, or new child issue begins.

If promoted, the first step is still an offline/read-only replay with no live authority. Until then, no current roadmap, milestone, gate, handoff, or implementation prompt should pull this concept into active scope.

Next action: none until explicit user promotion after private dogfood evidence exists.

Blocked by: no native issue blocker; intentionally parked by product decision.

Last verified: August 6, 2026.

Research Shape

  • Treat each worker harness as an autonomous organizational unit rather than a subagent call.
  • Keep the supervisor outside the worker's context and write plane.
  • Consume structured events and bounded evidence, never an unbounded transcript tail.
  • Begin offline or read-only; grant no steering, pause, file-write, merge, or release authority.
  • Measure intervention quality and human attention before adding product UI or authority levels.

Acceptance Criteria

  • The private dogfood milestone is closed before the user may consider promotion.
  • The user explicitly promotes this issue; milestone closure or session count alone never activates it.
  • Enough completed long-running dogfood sessions provide known outcomes and lifecycle evidence for a meaningful replay corpus.
  • A bounded event schema is defined through Expose a generic Work Pulse activity stream #298 or an explicitly temporary replay adapter.
  • An offline comparison evaluates event-only supervision against existing Background Review and, if useful, transcript-aware supervision.
  • Metrics include precision, false interventions, lead time, human attention required, token overhead, and correlation with worker errors.
  • Kill criteria are chosen before inference runs; poor precision or no material lead over existing controls ends the experiment.
  • Any follow-up preserves a strict no-write supervisor boundary and requires a separate durable plan before live authority is added.

Relationships

Next Action

Do nothing until the private dogfood milestone closes and enough real sessions exist. Then design one offline replay experiment before considering a live shadow supervisor.

Blocked by: no native issue blocker; intentionally waiting for released-product evidence.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    planDurable planning issueplan:donePlan completed or superseded

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions