Skip to content

Explore a thin-fork path: stock Codex engine with Lab features moved to plugin, MCP bridge, and a minimal patch series #926

Description

@cbusillo

Objective

Decide whether Codex Lab should stop being a full fork of openai/codex and instead run a stock (or near-stock) Codex engine with its differentiating features moved to the extension layers upstream now supports. This issue records the 2026-09-16 research comparing four paths — continue the fork, return to Every Code, pivot to T3 Code, or a thin fork — and proposes bounded spikes. It authorizes no implementation, cutover, or release change.

Finish Line

An owner decision on the engine strategy, backed by two spike results (hook-based review dogfood; account-rotation outside the engine), with a successor plan that either supersedes the convergence/release milestones under #498/#230 or records why the full fork stays.

Current Status

State: exploratory, waiting for owner review. Research performed 2026-09-16 against origin/main f6fd999 (lab.7), openai/main 2d90e05, just-every/code v0.6.189, and pingdotgg/t3code v0.0.42. No code changed.

Motivation: the full-fork cost is structural

Existing Substrate (upstream openai/main, 2026-09-16)

Now upstream, so no longer fork-justifying:

  • Hooks: Claude-compatible engine (codex-rs/hooks), events SessionStart/End, UserPromptSubmit, Pre/PostToolUse, PermissionRequest, Pre/PostCompact, Stop, SubagentStart/Stop, Interrupt. Stop can block and inject continuation fragments; async hooks (hook_config.rs async, 8 concurrent); loaded from hooks.json, [hooks] TOML, and plugin hook sources.
  • Plugins (codex-rs/plugin, core-plugins): contribute skills, MCP servers, apps, hooks; local/git/npm marketplaces.
  • Multi-agent: spawn_agent/wait_agent/send_input/resume_agent/close_agent/list_agents with roles (codex-rs/agent-roles).
  • Local models: --oss, --local-provider lmstudio|ollama, models-manager.
  • Remote: app-server --listen ws://, --remote-control, app-server-daemon (desktop/mobile), tcp-tunnel.
  • Config: chatgpt_base_url is supported.

Still fork-only: multi-account and rate-limit-aware switching (upstream is single-account; client_tests.rs "rejects_account_switches"); per-role model provider or external-command agents (child_config.rs forces child provider = parent); background/auto code review (Guardian is approval review, not code review).

Options Considered

A. Continue the full fork

Rejected on the evidence above unless a concrete, measured reduction in owned upstream files is named first (per #778's own rule).

B. Return to Every Code (just-every/code)

Alive (daily releases, 100 commits/30d) and has every feature natively: switch_active_account_on_rate_limit, Auto Review in a worktree, patch-harness validation, claude/agy/copilot subagents, browser. But code-rs is a hard fork that does not take upstream crates: no hooks, no plugins, no multi-agent v2, no gpt-6-astra, models top out at claude-opus-4.8; effectively one maintainer; all 20 of our PRs (#489–#578) were closed unmerged. This is the trade #28 already rejected (desktop/iOS compatibility, upstream parity). Not the base; remains a rollback and idea source.

C. Pivot to T3 Code over stock Codex, Claude Code, and Antigravity

Confirmed strengths: MIT; thin client over codex app-server with per-instance Binary path (packages/contracts/src/settings.ts), shadow-home multi-account, pooled limits UI with reset credits, native rendering of Codex collab-agent events, mobile/remote (t3 serve, T3 Connect, LAN, Tailscale), providers added quickly via ACP.

Confirmed blockers for our feature set:

Verdict: right architecture, not ready, and it structurally refuses two of our four features. Treat as an optional client to evaluate, not a dependency. Re-evaluate when openai#5024, openai#8259, openai#8499/openai#11996 are closed and the pin passes Codex 0.154.

D. Thin fork (proposed): stock engine, features in supported layers

Feature Current footprint (vs merge-base a62e98d) Proposed home Loss
Background Review 15 files, +11.3k (with validation) Lab plugin: async/blocking Stop hook runs a reviewer via codex exec/review/start; findings return as hook feedback; persistence in plugin data. just-every/plugin-auto-review is this shape and is already installed in ~/.codex. TUI/desktop status cells; review/* app-server methods (#231, #787, #784)
Automatic Validation (above) PostToolUse/Stop hooks (shellcheck, cargo check) in the same plugin; also works in Claude Code validation/completed event (#785)
Claude / Antigravity subagents 39 files, +12.9k agent-bridge MCP server (spawn/wait/send/cancel over claude -p, agy, codex exec), ported from core/src/agent/external_command.rs and external-agent-sessions. Native spawn_agent cannot host this (provider forced = parent); MCP is the supported seam and works from Codex, Claude Code, and T3 Tool-call granularity instead of native agent items; cancel/usage plumbing must be reimplemented (#85, #577, #906, #910)
Multi-account auto-switch 47 files, +15.8k (login, secrets, core, TUI) Spike a local auth proxy on chatgpt_base_url that rotates tokens on 429 (CLIProxyAPI-style). Fallback: rebased patch series carrying only auth_accounts.rs, account_switching.rs, and the manager.rs/server.rs hunks Proxy: TUI account views. Patch: ~5k lines to rebase per release instead of ~300k
Daemon behind ChatGPT desktop/iOS (#385) supervisor + Lab engine Same supervisor over stock codex app-server + app-server-daemon None
owner-control-, browser crate, code-bridge-, core-skills, TUI extras, version ~60k lines Drop, or keep as separate services (code-bridge already is one); upstream ships browser/computer-use plugins TUI polish

What this retires: bi-weekly convergence merges, upstream-convergence.yml, the 14 Lab-only workflows, self-hosted release/signing runners, provenance/Bazel/V8 caching work (#853–#857, #875, #876), and most of milestones 3, 6, 7, 8. New cadence: install the upstream standalone package + codex plugin marketplace add, or rebase a ≤5k-line patch per release.

Risks

Open Questions

  • Does blocking Stop-hook feedback render acceptably in the ChatGPT desktop app and iOS client, or only in the TUI?
  • Can the chatgpt_base_url proxy handle ChatGPT token refresh per account without engine changes? If not, is the login-crate patch really ≤5k lines after dropping the TUI views?
  • Which of the ~60k lines of TUI/app-server/protocol changes does anyone actually use daily? (Session-level evidence, not issue count.)
  • Should the orchestrator stay Codex, or does the MCP bridge make Claude Code an equal peer for some workflows?

Proposed Spikes (each separately authorized)

  1. Zero-code dogfood (1 week): enable plugin-auto-review and a PostToolUse shellcheck hook on stock Codex 0.154 in ~/.codex; use alongside Lab. Optionally point a T3 Codex instance at the same binary for a side-by-side.
  2. Auth-proxy spike (1–2 weeks): prototype rotation via chatgpt_base_url; decide proxy vs minimal patch.
  3. Bridge spike: stand up agent-bridge MCP with spawn/wait/cancel for claude and agy from the existing external-command code; verify from Codex and Claude Code.

Related

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    planDurable planning issueplan:donePlan completed or superseded

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions