Skip to content

Explore local-model fallback when cloud rate limits are exhausted #925

Description

@cbusillo

Objective

Explore whether Codex Lab should fall back to trusted local models, or offer them explicitly, when every cloud account is rate-limited, instead of stopping work.

This issue is an exploratory product and architecture record, not an implementation commitment.

Finish Line

The exploration is complete when Codex Lab has a documented recommendation covering:

  • which triggers qualify: all pooled accounts exhausted, a specific provider exhausted, or a user-selected budget mode
  • opt-in versus prompted versus automatic behavior, and how the downgrade is made visible in the TUI, app, and rollout provenance
  • which work may continue locally (for example scouting, review, and bounded edits) and which must pause for a cloud model
  • how the session returns to the cloud model when limits reset, including context handoff and compaction differences
  • local model selection, qualification, and health checks: loaded state, context length, tool-surface compatibility
  • privacy and locality policy for trusted localhost or LAN endpoints
  • the smallest reviewable implementation slices, if promoted

Current Status

State: exploratory and waiting as of September 15, 2026.

The motivating dogfood event was a user running out of OpenAI rate limit, then manually standing up an isolated lab.7 home against LM Studio. On an Apple M4 Max with 128 GB, Qwen3.8-27B (MLX) completed a multi-step edit–test–fix task in about 2m40s at about 20 tok/s.

Next action: none until the local-provider blockers are fixed and dogfood shows local models are good enough for the fallback use case.

Motivation

  • Rate-limit exhaustion currently ends productive work even when capable local models are available on the same machine.
  • auto_switch_accounts_on_rate_limit already switches among cloud accounts. A local tier is the natural next rung, but a silent downgrade to a weaker model could produce bad work unless it is clearly visible and scoped.
  • Local model quality changes weekly, so the design should not hard-code model names.

Existing Substrate

  • Account pooling and rate-limit switching (auto_switch_accounts_on_rate_limit, execution-account leases).
  • Custom Responses providers plus model_catalog_json, which are enough to run root sessions on LM Studio today.
  • Multi-agent V2 per-spawn model and agents.default_subagent_model, which could let a cloud root delegate to local subagents.
  • codex-skills local-llm: private endpoint and role config, inventory, benchmarking, and the isolated local_codex_agent.py runner.

Blockers Found

Candidate Shapes

  1. Manual: a /model or profile switch to a curated local profile, surfaced when rate limits hit.
  2. Prompted: on exhaustion, offer "continue on local model X", with a clear banner and provenance.
  3. Delegation-first: keep the cloud root, route bounded subagent work to local models while limits are low.
  4. Automatic: policy-driven fallback for allowlisted task kinds, with automatic return to cloud.

Risks

  • Quality regression hidden from the user.
  • Context-window and compaction mismatch when switching mid-session.
  • Tool-surface differences between providers.
  • Memory pressure on the host.

Related

#922, #924, #923, #809, #778, #906, #115.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    planDurable planning issueplan:waitingDurable plan parked pending a decision, event, or non-issue condition; not for PR QA

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions