feat(run): guarantee a plan artifact and add the two plan run shapes - #325
Merged
Merged
Conversation
A plan-mode dispatch relied entirely on the harness's own planning prompt to decide what the agent wrote, so what came back varied by harness and a planning phase that presented nothing produced no artifact at all. Instructing the agent to write `plan.md` cannot fix that: plan mode refuses writes into the task environment by design. The agent's final message is the one channel every harness has, so eval-magic now asks for the plan there. `src/cli/run/plan_prompt.rs` adds harness-neutral planning instructions to the dispatch prompt — read but do not edit, close the turn with the complete plan — and `PlanSignal::FinalMessage` takes that message as the plan when no plan file was written and no responder was declared. Every planning phase therefore produces `outputs/plan.md`, and no eval needs a responder to reach one: the preflight that required one on a harness without `[plan_mode.plan_file]` is gone. `plan_mode` becomes tri-state. `true` and `"plan_then_act"` keep the existing plan-approve-implement shape; `"plan_only"` stops at the plan, recording no `approved_in_round` and dispatching no act round, for a skill that only shapes how a plan is written. `false` and `true` serialize exactly as before, so no existing evals.json or dispatch.json changes. `plan_source` is the mirror: it names a plan written beforehand under the skill's `evals/` directory and splices it into the prompt as already approved. The session is ordinary act mode, so it works on Codex and Cline too, neither of which can declare `[plan_mode]`. `ConversationStopReason::PlanNotPresented` is retained for reading older artifacts but is no longer produced. Verification: cargo fmt --check, cargo clippy --all-targets -- -D warnings, and cargo test all pass (1509 tests). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ERLJRvxDfHdKHdXMyBLmse
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #324.
The problem
A plan-mode dispatch relied entirely on the harness's own planning prompt to decide what the agent produced. Nothing eval-magic wrote ever asked for a plan, so what came back varied by harness, and a planning phase that presented nothing produced no artifact at all — the run just stopped with
plan_not_presented.The issue's suggested mechanic ("write your final plan to plan.md") cannot work. Claude Code's
--permission-mode planrefuses every write except into~/.claude/plans, and OpenCode'splanagent refuses edits by permission rule. Refusing writes is plan mode. The one channel every harness has is the agent's final message, which eval-magic already captures asTranscriptSummary::final_text, so that is what the prompt now asks for.The dispatch mechanics themselves were checked and are correct:
claude --helpon 2.1.269 listsplanamong--permission-mode's values and offers no plan-file location flag, soplan_args = " --permission-mode plan"with abypassPermissionsresume is the documented way in and out.What changed
A plan artifact always exists. New
src/cli/run/plan_prompt.rsadds harness-neutral planning instructions to the dispatch prompt — read but do not edit, close the turn with the complete plan, and what follows it.PlanSignal::FinalMessageis a third signal that takes that message as the plan when neither a plan file nor a responder is available. Every planning phase now writesoutputs/plan.md.The signal ladder, recorded in
conversation.jsonasplan.signal:plan_fileresponderdone.final_messageBecause the last always fires, no eval needs a responder to reach a plan: the preflight that required one on a harness without
[plan_mode.plan_file]is removed, so OpenCode plan-mode evals work bare. A responder still buys a planning phase of more than one round.Two plan run shapes.
plan_modeis tri-state:falseandtrueserialize exactly as before, so no existingevals.jsonordispatch.jsonshifts a byte. A plan-only run records noapproved_in_roundand dispatches no act round — for a skill that only shapes how a plan is written, where running the implementation spends tokens on work the eval does not measure.plan_sourceis the mirror: it names a plan written beforehand under the skill'sevals/directory (honoringfiles_root) and splices its text into the prompt as already approved. The session is ordinary act mode, so this reaches Codex and Cline too, neither of which can declare[plan_mode]at all. It is mutually exclusive withplan_mode.Chaining a plan-only campaign's output into an executing one stays manual — copy
outputs/plan.mdinto the executing skill'sevals/and name it inplan_source. Which plan to carry across is a judgement about the comparison being made.Before / after
A planning round on OpenCode, which writes no plan file, with no responder declared:
A plan-only run:
Notes for review
final_messagefallback can record a non-plan as a plan. Without a responder, an agent that spent its planning turn asking a question has that question saved asplan.md.plan.signaldistinguishes it and the guide says so plainly, but it is a real behavior change from the oldplan_not_presentedstop.ConversationStopReason::PlanNotPresentedis retained but no longer produced — kept so olderconversation.jsonartifacts still deserialize.PlanRecord::approved_in_roundis now optional.schema/conversation.schema.jsonandschema/run-record.schema.jsondrop it fromrequired; the judge evidence bundle renders "the session ended after planning" in its place.claude2.1.259+ still writes~/.claude/plans/*.md, soplan.signalreadsplan_filerather than falling back — needs real model spend and is being run separately.Documentation
eval-magic docs conversationsgains the three-signal table, both plan-mode shapes with worked examples, aplan_sourcesection covering the field and thefiles-overlay alternative, and the manual chaining recipe. Also updated:eval-magic docs byoh, therun/dispatch/inithelp,docs/progressive-enhancements.md,harnesses/template.toml, both harness notes, and the README.Verification
New coverage:
PlanModeserde across every spelling, thefinal_messagefallback and its precedence, the plan-mode and supplied-plan prompt bodies, plan-source resolution and containment, the three plan-declaration validation rules, and end-to-end tests for a plan-only stop, a multi-round plan-only run with a responder, and a supplied plan on Codex.🤖 Generated with Claude Code
https://claude.ai/code/session_01ERLJRvxDfHdKHdXMyBLmse