Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion Cargo.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

2 changes: 1 addition & 1 deletion Cargo.toml
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
[package]
name = "eval-magic"
version = "0.10.0"
version = "0.11.0"
edition = "2024"
description = "One-stop CLI for running skill evals — measure whether an agent skill actually shifts behavior."
license = "MIT"
Expand Down
6 changes: 3 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -123,9 +123,9 @@ Each eval case runs once per condition and repetition in its own clean Git repos
receive the same codebase, task, and overlays; only the condition under test changes. Assertions can
combine LLM judgment with runner-owned command checks, transcript checks, and final diff limits.
Multi-turn evals resume one native harness session so follow-up answers remain part of the same
conversation, whether the turns are scripted or derived by a responder. An eval can start that
session in the harness's native plan mode and continue in act mode once the plan is approved
(`eval-magic docs conversations`).
conversation, whether the turns are scripted or derived by a responder. An eval can start in the
harness's native plan mode — implementing the plan once approved, or stopping with the plan as its
output — or be handed a plan written beforehand to carry out (`eval-magic docs conversations`).

Most harness features are declared in TOML descriptors. See the current registry and resolved data
instead of relying on a static compatibility table:
Expand Down
16 changes: 13 additions & 3 deletions docs/claude-notes.md
Original file line number Diff line number Diff line change
Expand Up @@ -89,9 +89,19 @@ scratch repository with one bug:
`permissionMode: bypassPermissions` in `init`, keeps the same `session_id`, and edits succeed:
the session leaves plan mode. The Claude Code documentation states the converse — a `-p --resume`
stays in plan mode only when `--permission-prompt-tool` is passed and no `--permission-mode` is.
- The plan file lands in the operator's `~/.claude/plans`, as in any session. The write guard and
the stray-write audit allow that root; eval-magic keeps the copy the judge reads as
`outputs/plan.md`.
- The plan file lands in the operator's `~/.claude/plans`, as in any session — one directory
shared by every concurrent task, and by the operator's own sessions. That is not a collision:
the plan text is read from the write's `content` argument in the round's own transcript, not
from the file, so tasks cannot read each other's plans. The write guard and the stray-write
audit allow that root; eval-magic keeps the copy the judge reads as `outputs/plan.md`.
- **Plan mode refuses writes into the task environment**, which is why the plan artifact cannot be
a file eval-magic asks the agent to write there. The harness-neutral half of the contract is the
agent's final message: `src/cli/run/plan_prompt.rs` asks every planning round to close with its
complete plan, and that is the `final_message` signal a harness without a plan file falls back
to. On Claude Code the plan file wins, so the fallback is a safety net rather than the path.
- **`--append-system-prompt` is deliberately not used** to steer where the plan is written. There
is no flag that relocates `~/.claude/plans`, and a harness-specific instruction would give
Claude Code a contract no other harness could honor.

Relaxing the default closes the common case, not the class — a deny rule, a managed setting, or an
operator-overridden mode still refuses calls — so refusals are detected and reported rather than
Expand Down
10 changes: 6 additions & 4 deletions docs/guides/byoh.md
Original file line number Diff line number Diff line change
Expand Up @@ -199,10 +199,12 @@ template you already declared, so it works on any harness that can resume. See

A harness with a native read-only planning mode declares `[plan_mode]`: `plan_args` and
`act_args` fill a `{mode_args}` slot in both dispatch templates, and `[plan_mode.plan_file]` names
the file the harness writes its plan to, when it writes one. Verify each value the way you verify a
resume: one headless turn dispatched with the planning arguments must refuse an edit, and the same
session resumed with the act arguments must make it. `eval-magic harness show claude-code` is a
worked descriptor; the eval side is in `eval-magic docs conversations`.
the file the harness writes its plan to, when it writes one. Omitting `plan_file` costs nothing —
the planning round's final message is then the plan, because eval-magic's own dispatch prompt asks
every planning round to close with it. Verify each value the way you verify a resume: one headless
turn dispatched with the planning arguments must refuse an edit, and the same session resumed with
the act arguments must make it. `eval-magic harness show claude-code` is a worked descriptor; the
eval side is in `eval-magic docs conversations`.

When a shadow preflight reports a live copy, isolate every initial and resumed eval-agent dispatch
before setting `isolates_live_sources = true`. The per-harness remedies and verification procedure
Expand Down
112 changes: 96 additions & 16 deletions docs/guides/conversations.md
Original file line number Diff line number Diff line change
Expand Up @@ -110,7 +110,6 @@ dispatch call for different fixes.
| `completed` | The responder judged the agent finished. The run stops rather than burning its remaining turns. |
| `stopped`, `responder_cannot_answer` | The responder produced no usable reply. `responder_outcome.cause` says why. |
| `stopped`, `max_turns_reached` | The agent was still asking at the bound. |
| `stopped`, `plan_not_presented` | A plan-mode session ended its planning phase with no plan to approve. |
| `timed_out` | The task outran `dispatch --timeout`. |

A `stopped` conversation is recorded, not failed: `dispatch` exits zero and
Expand Down Expand Up @@ -192,16 +191,24 @@ approves it. An eval declares that shape with `plan_mode`:
"id": "add-request-caching",
"prompt": "Requests to the pricing API are slow. Can you add caching?",
"expected_output": "A working cache with the pricing endpoint under 100ms.",
"plan_mode": true,
"responder": { "type": "llm" }
"plan_mode": true
}
```

`plan_mode` takes one of three values:

| Value | Shape |
| --- | --- |
| `true`, or `"plan_then_act"` | Plan, approve, implement — the two phases below. |
| `"plan_only"` | Plan and stop. The plan is the run's whole output. |
| `false`, or omitted | Not a plan-mode eval. |

The session runs in two phases, in one native session:

1. **Planning.** The opening prompt is dispatched in the harness's native plan
mode, so the agent explores read-only. If it asks a question, the responder
answers it and the session stays in plan mode.
mode, so the agent explores read-only. If it asks a question and the eval
declares a `responder`, the responder answers and the session stays in plan
mode.
2. **Implementation.** Once the agent has presented its plan, the runner
approves it with one fixed message and resumes the same session in act
mode. The message is always `The plan is approved. Implement it now.` From
Expand All @@ -212,37 +219,66 @@ The approval is fixed rather than judged, so the transition is identical in
every run and both arms: what varies between conditions is the plan and the
implementation, never how the runner reacted to them.

A `"plan_only"` eval runs the first phase and stops there. Nothing is approved
and no act round is dispatched, so the working tree stays clean and
`outputs/plan.md` is what the judge grades. Use it for a skill that only shapes
how the plan is written — running the implementation would spend tokens on work
the eval does not measure. Scripted `turns` are rejected on a plan-only eval,
since the session ends before any of them could be delivered; a `responder` is
still allowed, and answers the agent's questions while it plans.

### What the agent is told

A planning round's dispatch prompt carries instructions eval-magic writes
itself, identical in both arms:

- the session starts in planning mode, so read and explore but do not edit;
- end the turn with the complete plan as the final message — the whole plan
text, not a summary of it and not a pointer to where it lives;
- what follows the plan, which is the one thing the two shapes disagree on.

Harnesses word their own planning modes differently and some say nothing about
how a plan should be presented. These lines are the part that reads the same
everywhere, and they are what makes the final message a reliable place to
recover a plan from.

### How the runner knows the plan is ready

Two signals, tried in this order:
Three signals, tried in this order:

| Signal | When it applies |
| --- | --- |
| `plan_file` | The harness writes the plan it presents to a file, and its descriptor declares where (`[plan_mode.plan_file]`). A planning round that wrote one has presented its plan, and the file's content is the plan. Claude Code writes to `~/.claude/plans`. |
| `responder` | The eval declares a `responder`, which is told the agent is planning. Its `done` verdict means the plan is ready, and the agent's last message is the plan. |
| `final_message` | Neither of the above was available. The planning round's final message is the plan, which is what the dispatch prompt asked the agent to close with. |

A harness that writes no plan file needs the responder to decide, so `run`
rejects a plan-mode eval there unless it declares one. With a plan file the
responder is optional; if the agent never writes one and there is no responder
to ask, the run stops with `plan_not_presented`.
The last signal always fires, so a plan-mode eval needs no responder on any
harness and every planning phase produces a `plan.md`. What the responder buys
is a planning phase that can run for more than one round: without one, the
agent's first closing message ends the phase, so an agent that spent its turn
asking a question has that question recorded as its plan. `conversation.json`'s
`plan.signal` says which signal fired, so a run resting on `final_message` is
distinguishable from one that wrote a real plan file.

`eval-magic harness list` shows `plan-mode` for every harness that can start a
session in plan mode. `run` rejects a plan-mode eval on one that cannot, before
any environment is built.

### What is recorded

- `outputs/plan.md` holds the approved plan, and the judge evidence bundle
renders it in a section of its own.
- `outputs/plan.md` holds the plan, and the judge evidence bundle renders it in
a section of its own.
- Every user message in `conversation.json` carries `mode` (`plan` or `act`),
and the approval turn carries `origin.runner: plan_approval`.
- `conversation.json`'s `plan` names the round the plan was presented in, the
round the approval opened, and which signal fired.
- `conversation.json`'s `plan` names the round the plan was presented in, which
signal fired, and the round the approval opened. A plan-only run has no
approval round, so it records no `approved_in_round`.
- A responder's `max_turns` is one budget across both phases: planning-phase
answers count toward it.
- A plan-mode run whose session never left the planning phase — whatever
- A plan-then-act run whose session never left the planning phase — whatever
stopped it — is counted per condition in `benchmark.json`'s
`validity_warnings`, because it never attempted the task.
`validity_warnings`, because it never attempted the task. A plan-only run
presents a plan, so it is not one of those.
- An agent that tries to edit while planning is refused by the mode itself.
That refusal is recorded in `permission-denials.json` as behavioral evidence
and marked `plan_mode_attributed`, but it raises no validity warning: the
Expand All @@ -251,3 +287,47 @@ any environment is built.
Where the harness writes its plan file, the write guard and the stray-write
audit allow that root beside the task environment. The plan file lands where
the harness puts it in any session; `plan.md` is the copy the judge reads.

## Starting from a plan someone already wrote

The mirror of a plan-only eval: `plan_source` hands the agent a finished plan
and asks it to carry the plan out. The path is relative to the skill's `evals/`
directory, under `files_root` when the eval sets one:

```json
{
"id": "execute-cache-plan",
"prompt": "Implement the approved plan.",
"expected_output": "A working cache matching the plan.",
"plan_source": "plans/add-request-caching.md"
}
```

The file's text is spliced into the dispatch prompt ahead of the request,
framed as already approved. The session is an ordinary act-mode dispatch, so
this needs nothing from the harness: it works on every harness, including one
whose descriptor declares no `[plan_mode]`. `dispatch.json` records the path
each task's plan came from; the text itself is in `dispatch-prompt.txt`.

`plan_source` and `plan_mode` are mutually exclusive — a run either writes its
own plan or is handed one.

To carry a plan-only campaign's output into an executing campaign, copy the
run's `outputs/plan.md` into the executing skill's `evals/` directory and name
it in `plan_source`. Which plan to carry across is a judgement about the
comparison you are making, so the copy is deliberate rather than automatic.

The alternative, when the plan should be a file the agent reads rather than
prompt context, is the `files` overlay:

```json
{
"id": "execute-cache-plan",
"files": ["PLAN.md"],
"prompt": "Implement the plan in PLAN.md.",
"expected_output": "A working cache matching the plan."
}
```

That stages `PLAN.md` in the task repository, where it is part of the baseline
and so does not count against a `diff_scope` assertion.
6 changes: 4 additions & 2 deletions docs/opencode-notes.md
Original file line number Diff line number Diff line change
Expand Up @@ -77,8 +77,10 @@ bug (free model `opencode/nemotron-3.5-lightning-free`):
agent makes, so it belongs to `act_args` only.
- `opencode run --session <id> --agent build --auto` resumed the same session and edited the file.
`--agent build` is explicit so a resumed session does not inherit the plan agent.
- OpenCode writes no plan file, so the descriptor declares no `[plan_mode.plan_file]`; a plan-mode
eval on OpenCode needs a `responder`, whose `done` verdict in the planning phase approves the plan.
- OpenCode writes no plan file, so the descriptor declares no `[plan_mode.plan_file]`. The
planning round's final message is the plan there — the dispatch prompt asks every planning round
to close with it — so a plan-mode eval on OpenCode needs no `responder`. Declaring one still
buys a planning phase of more than one round, with its `done` verdict approving the plan.

## Write guard

Expand Down
24 changes: 18 additions & 6 deletions docs/progressive-enhancements.md
Original file line number Diff line number Diff line change
Expand Up @@ -426,15 +426,27 @@ scan config-declared `skills.paths`/`skills.urls` sources.
*Why harness-specific:* the read-only planning mode is the harness's own — a permission mode for
Claude Code, a built-in agent for OpenCode — and so is the way the agent presents its plan.

*What it unlocks:* evals that declare `plan_mode: true`. The driver dispatches the opening round
with the planning arguments, lets the agent present a plan, approves it with one fixed message, and
resumes the same session with the act arguments; the eval's `turns` or `responder` then proceed as
usual. The approved plan is saved as `outputs/plan.md` and rendered in the judge evidence bundle.
`plan_file` is the deterministic signal that the plan was presented; without one the eval's
responder decides, which is why `run` requires a responder on a harness without a plan file.
*What it unlocks:* evals that declare `plan_mode`. The driver dispatches the opening round with the
planning arguments, lets the agent present a plan, and saves it as `outputs/plan.md`, rendered in
the judge evidence bundle. `plan_mode: true` then approves the plan with one fixed message and
resumes the same session with the act arguments, where the eval's `turns` or `responder` proceed as
usual; `plan_mode: "plan_only"` stops at the plan, making it the run's whole output.

The descriptor supplies only half the capability. The other half is harness-neutral and lives in
the dispatch prompt (`src/cli/run/plan_prompt.rs`): a planning round is told it cannot edit and
that its final message must carry the complete plan. That is what makes the third signal below
work, and it is the part a new harness inherits without declaring anything.

Three signals mark a plan as presented, tried in order: the `plan_file` write, a responder's `done`
verdict, and the planning round's final message. The last always fires, so no eval needs a
responder to reach a plan and every planning phase produces an artifact. `conversation.json`'s
`plan.signal` records which one did, so a plan resting on the final message stays distinguishable
from one read out of a native plan file.

*Fallback:* none. `run` rejects a plan-mode eval for a harness without `[plan_mode]`, before any
environment is built, the way it rejects multi-turn evals for a harness without `[conversation]`.
The eval-side `plan_source` field is not this capability: it splices a pre-written plan into an
ordinary act-mode dispatch and needs no descriptor support at all.

*Descriptor fields:* the `[plan_mode]` table — `plan_args` and `act_args` fill the `{mode_args}`
slot that both `dispatch.exec_template` and `conversation.resume_exec_template` must carry (the act
Expand Down
4 changes: 3 additions & 1 deletion harnesses/template.toml
Original file line number Diff line number Diff line change
Expand Up @@ -190,7 +190,9 @@ label = "{label}"
## [plan_mode.plan_file] names the file the harness writes its plan to: writes under root are
## allowed by the write guard and the stray-write audit, a plan-phase round that wrote one has
## presented its plan, and that write's content_field is the plan artifact. Omit plan_file when
## the harness writes no plan file; the eval's responder then decides when the plan is ready.
## the harness writes no plan file: the eval's responder then decides when the plan is ready, and
## without one the planning round's final message is the plan — eval-magic's own prompt asks every
## planning round to close with it, so a plan artifact is produced either way.
## VERIFY: run one headless turn in the planning mode and confirm edits are refused; resume it
## with the act arguments and confirm that same session edits. See `eval-magic harness show
## claude-code` for a plan_file example and `eval-magic docs conversations` for the eval side.
Expand Down
10 changes: 5 additions & 5 deletions schema/conversation.schema.json
Original file line number Diff line number Diff line change
Expand Up @@ -115,9 +115,9 @@
},
"planRecord": {
"type": "object",
"required": ["presented_in_round", "approved_in_round", "signal"],
"required": ["presented_in_round", "signal"],
"additionalProperties": false,
"description": "How a plan-mode session moved from planning to implementation. Absent unless the eval declared plan_mode and a plan was approved.",
"description": "How a plan-mode session moved from planning to implementation. Absent unless the eval declared plan_mode and a plan was presented.",
"properties": {
"presented_in_round": {
"type": "integer",
Expand All @@ -127,12 +127,12 @@
"approved_in_round": {
"type": "integer",
"minimum": 2,
"description": "The round the runner's fixed approval opened in act mode."
"description": "The round the runner's fixed approval opened in act mode. Absent on a plan_only run, which presents its plan and stops without one."
},
"signal": {
"type": "string",
"enum": ["plan_file", "responder"],
"description": "What marked the plan as presented: the harness wrote its declared plan file (plan_file), or the responder judged the agent finished planning (responder)."
"enum": ["plan_file", "responder", "final_message"],
"description": "What marked the plan as presented, in the order they are tried: the harness wrote its declared plan file (plan_file), the responder judged the agent finished planning (responder), or neither was available and the planning round's final message was taken as the plan (final_message)."
},
"artifact_path": {
"type": "string",
Expand Down
Loading
Loading