Skip to content

feat(run): guarantee a plan artifact and add the two plan run shapes - #325

Merged
slowdini merged 1 commit into
devfrom
feat/324-plan-mode-handling
Sep 12, 2026
Merged

slowdini merged 1 commit into
devfrom
feat/324-plan-mode-handling

Conversation

@slowdini

Copy link
Copy Markdown
Owner

Closes #324.

The problem

A plan-mode dispatch relied entirely on the harness's own planning prompt to decide what the agent produced. Nothing eval-magic wrote ever asked for a plan, so what came back varied by harness, and a planning phase that presented nothing produced no artifact at all — the run just stopped with plan_not_presented.

The issue's suggested mechanic ("write your final plan to plan.md") cannot work. Claude Code's --permission-mode plan refuses every write except into ~/.claude/plans, and OpenCode's plan agent refuses edits by permission rule. Refusing writes is plan mode. The one channel every harness has is the agent's final message, which eval-magic already captures as TranscriptSummary::final_text, so that is what the prompt now asks for.

The dispatch mechanics themselves were checked and are correct: claude --help on 2.1.269 lists plan among --permission-mode's values and offers no plan-file location flag, so plan_args = " --permission-mode plan" with a bypassPermissions resume is the documented way in and out.

What changed

A plan artifact always exists. New src/cli/run/plan_prompt.rs adds harness-neutral planning instructions to the dispatch prompt — read but do not edit, close the turn with the complete plan, and what follows it. PlanSignal::FinalMessage is a third signal that takes that message as the plan when neither a plan file nor a responder is available. Every planning phase now writes outputs/plan.md.

The signal ladder, recorded in conversation.json as plan.signal:

Signal When
plan_file The harness wrote its declared plan file.
responder The eval's responder returned done.
final_message Neither — the planning round's closing message is the plan.

Because the last always fires, no eval needs a responder to reach a plan: the preflight that required one on a harness without [plan_mode.plan_file] is removed, so OpenCode plan-mode evals work bare. A responder still buys a planning phase of more than one round.

Two plan run shapes. plan_mode is tri-state:

"plan_mode": true              // == "plan_then_act": plan, approve, implement
"plan_mode": "plan_only"       // stop at the plan; outputs/plan.md is the output
"plan_mode": false             // or omitted

false and true serialize exactly as before, so no existing evals.json or dispatch.json shifts a byte. A plan-only run records no approved_in_round and dispatches no act round — for a skill that only shapes how a plan is written, where running the implementation spends tokens on work the eval does not measure.

plan_source is the mirror: it names a plan written beforehand under the skill's evals/ directory (honoring files_root) and splices its text into the prompt as already approved. The session is ordinary act mode, so this reaches Codex and Cline too, neither of which can declare [plan_mode] at all. It is mutually exclusive with plan_mode.

Chaining a plan-only campaign's output into an executing one stays manual — copy outputs/plan.md into the executing skill's evals/ and name it in plan_source. Which plan to carry across is a judgement about the comparison being made.

Before / after

A planning round on OpenCode, which writes no plan file, with no responder declared:

before  $ eval-magic run --harness opencode
        error: --harness opencode needs a responder on plan-mode evals (plan-first): its
        descriptor declares no [plan_mode.plan_file], so only a responder can tell when
        the plan is ready for approval

after   $ eval-magic run --harness opencode
          plan mode: 1 plan-then-act eval(s) start in the harness's native plan mode and
          continue in act mode once the plan is approved
        $ eval-magic dispatch --harness opencode
        [1/1] plan-first:with_skill: completed with 1 follow-up turn(s), plan approved in round 2
        $ cat .../outputs/plan.md
        1. Add an in-memory LRU in pricing.py
        2. Cover eviction with a test

A plan-only run:

$ eval-magic dispatch --harness claude-code
[1/1] plan-only:with_skill: completed — plan presented in round 1, saved as outputs/plan.md

Notes for review

  • The final_message fallback can record a non-plan as a plan. Without a responder, an agent that spent its planning turn asking a question has that question saved as plan.md. plan.signal distinguishes it and the guide says so plainly, but it is a real behavior change from the old plan_not_presented stop.
  • ConversationStopReason::PlanNotPresented is retained but no longer produced — kept so older conversation.json artifacts still deserialize.
  • PlanRecord::approved_in_round is now optional. schema/conversation.schema.json and schema/run-record.schema.json drop it from required; the judge evidence bundle renders "the session ended after planning" in its place.
  • Not verified against a live harness. The empirical check — that headless claude 2.1.259+ still writes ~/.claude/plans/*.md, so plan.signal reads plan_file rather than falling back — needs real model spend and is being run separately.

Documentation

eval-magic docs conversations gains the three-signal table, both plan-mode shapes with worked examples, a plan_source section covering the field and the files-overlay alternative, and the manual chaining recipe. Also updated: eval-magic docs byoh, the run/dispatch/init help, docs/progressive-enhancements.md, harnesses/template.toml, both harness notes, and the README.

Verification

cargo fmt --check                            exit 0
cargo clippy --all-targets -- -D warnings    exit 0
cargo test                                   exit 0 — 1509 passed, 0 failed

New coverage: PlanMode serde across every spelling, the final_message fallback and its precedence, the plan-mode and supplied-plan prompt bodies, plan-source resolution and containment, the three plan-declaration validation rules, and end-to-end tests for a plan-only stop, a multi-round plan-only run with a responder, and a supplied plan on Codex.

🤖 Generated with Claude Code

https://claude.ai/code/session_01ERLJRvxDfHdKHdXMyBLmse

A plan-mode dispatch relied entirely on the harness's own planning prompt to
decide what the agent wrote, so what came back varied by harness and a planning
phase that presented nothing produced no artifact at all. Instructing the agent
to write `plan.md` cannot fix that: plan mode refuses writes into the task
environment by design. The agent's final message is the one channel every
harness has, so eval-magic now asks for the plan there.

`src/cli/run/plan_prompt.rs` adds harness-neutral planning instructions to the
dispatch prompt — read but do not edit, close the turn with the complete plan —
and `PlanSignal::FinalMessage` takes that message as the plan when no plan file
was written and no responder was declared. Every planning phase therefore
produces `outputs/plan.md`, and no eval needs a responder to reach one: the
preflight that required one on a harness without `[plan_mode.plan_file]` is
gone.

`plan_mode` becomes tri-state. `true` and `"plan_then_act"` keep the existing
plan-approve-implement shape; `"plan_only"` stops at the plan, recording no
`approved_in_round` and dispatching no act round, for a skill that only shapes
how a plan is written. `false` and `true` serialize exactly as before, so no
existing evals.json or dispatch.json changes.

`plan_source` is the mirror: it names a plan written beforehand under the
skill's `evals/` directory and splices it into the prompt as already approved.
The session is ordinary act mode, so it works on Codex and Cline too, neither of
which can declare `[plan_mode]`.

`ConversationStopReason::PlanNotPresented` is retained for reading older
artifacts but is no longer produced.

Verification: cargo fmt --check, cargo clippy --all-targets -- -D warnings, and
cargo test all pass (1509 tests).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ERLJRvxDfHdKHdXMyBLmse
@slowdini
slowdini merged commit 83f80d9 into dev Sep 12, 2026
7 checks passed
@slowdini
slowdini deleted the feat/324-plan-mode-handling branch September 12, 2026 06:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Plan mode handling improvements

1 participant