Skip to content

Release v0.10.0 - #322

Merged
slowdini merged 99 commits into
mainfrom
dev
Sep 5, 2026
Merged

slowdini merged 99 commits into
mainfrom
dev

Conversation

@slowdini

@slowdini slowdini commented Sep 5, 2026 •

Copy link
Copy Markdown
Owner

Release notes

Highlights

  • Evaluate skills against real projects. Source a codebase from a Git URL and ref or a local
    directory, and run each condition in its own working tree. Codebase caching avoids repeated
    full clones; copied skill inputs and recorded revisions make the treatment traceable through
    runs, benchmarks, and promoted baselines. (#278,
    #279, #280)
  • Let the runner drive the campaign. eval-magic dispatch runs tasks and judges with bounded
    parallelism, task timeouts, and selective retries. An LLM responder can answer follow-up
    questions without seeing grading criteria. Native planning sessions can proceed through a
    recorded approval into implementation on Claude Code and OpenCode.
    (#282, #285,
    #316)
  • Measure combinations of skills. An ordered skill_name list defines one treatment, with
    invocation or access checks for each member and provenance for the complete set. Existing
    scalar configurations retain their single-skill artifact format.
    (#291)
  • Inspect evidence before defining assertions. Git-based diff capture and bounded
    judge-evidence.md bundles preserve the task, conversation, and code changes behind a verdict.
    eval-magic compare presents paired evidence even without authored assertions; promoted
    baselines retain their evidence. Comparison reports are exploratory, not grades.
    (#281, #288,
    #290)
  • Measure judge consistency. --judge-samples N and per-assertion samples request repeated
    verdicts over the same persisted evidence. Reports retain individual votes, pass proportions,
    and pass^k consistency scores. Sampling does not rerun the evaluated agent; one sample keeps
    the existing binary grading format. (#289)
  • Configure development commands explicitly. Eval-authored guard policies combine tool or
    command allowances with Rust, JavaScript, Python, and Next.js profiles. Containment checks
    remain non-overridable. The expanded policy and effective guard state are recorded for
    enforcement, auditing, and baseline review.
    (#286, #313)
  • Compare execution time across harnesses. Runner-measured duration sums eval-agent subprocess
    time across rounds, excluding responder consultations, judges, and queue time. Timing artifacts
    identify token and duration sources separately; missing historical measurements remain missing.
    (#319)

Behavior changes to know about

  • Every eval requires a codebase. Add a top-level codebase default or an eval-level override,
    and remove the retired isolation field. files and files_root supply overlays on that
    codebase. init defaults to the pinned Weeknight fixture; choose another source with
    --codebase-url plus --codebase-ref, --codebase-path, or --codebase-cwd.
    See eval-magic docs codebase.
    (#292, #293)
  • The default workspace moves outside the skill repository. It uses a namespaced directory
    under $XDG_DATA_HOME/eval-magic or ~/.local/share/eval-magic. Use --workspace-dir or
    EVAL_MAGIC_WORKSPACE_DIR to select another location or continue an existing campaign.
  • Windows use requires WSL. eval-magic supports Linux and macOS. On Windows, run eval-magic
    inside WSL; native Windows is unsupported. Git and a POSIX shell are required. Set EVAL_MAGIC_SH
    to select a specific sh. Runner dispatch no longer requires jq.
    (#283)
  • Update dispatch integrations. Replace dispatch-task with dispatch --task-index.
    fill-transcripts and fallback ingestion from final-message or capture-prefix files are
    retired; completion comes from runner-owned conversations and native transcripts. Custom
    harnesses must declare valid [dispatch], [transcript], and [tools] sections; remove
    parallel_command_template and judge_command_template. See eval-magic docs byoh.
  • Plan mode is declared per eval. Replace the simulated run --plan-mode option with
    "plan_mode": true in evals.json. Unsupported harnesses fail preflight; Codex and Cline do
    not support this mode. See eval-magic docs conversations.
  • Sourced project configuration is preserved. Project instructions, settings, plugins, and
    skills remain visible by default. Set codebase.exclude_skill_sources: true to exclude the
    selected harness's project skill directories from both arms while retaining other configuration.
    Matching treatment skills are reported as contamination risks.
    (#287)

Fixes

  • Assertions edited after a run are used when regrading: grade reads live assertions and
    skill_should_trigger, while preserving the captured task definition and recording the
    assertion source. (#300)
  • Partial ingestion leaves unfinished task environments untouched by command checks, including
    held-out setup files. Cached checks require matching assertion and run-record digests.
    (#312)
  • Equivalent tool calls no longer miss transcript assertions solely because harnesses use
    different names; descriptor roles provide portable aliases. Codex skill-access checks require
    a successful read of the exact staged SKILL.md path.
    (#314, #315)
  • Generated follow-up commands retain --harness-file, and descriptor drift produces a warning.
    (#299)
  • Framework files stop polluting supported project tools' checks, and untracked task scratch
    files no longer inflate measured diffs. Guard teardown reaches the selected iteration's task
    environments. (#302,
    #304)
  • Guard denials identify every blocking layer. Containment checks cover touch, mkdir, rm,
    cp, mv, and install, including destinations outside the task environment.
    (#303, #320)
  • Live skill-source reads through symlink aliases are detected, including macOS /var and
    /private/var spellings. The isolation guide also explains nested Codex dispatch failures and
    their remedies. (#318,
    #321)

slowdini and others added 30 commits August 15, 2026 20:35
Merge pull request #276 from slowdini/dev
An eval environment could only be assembled from individual fixture files
copied out of `<skill>/evals/`, so there was no way to say "run this task
against *this project*". Add a `codebase` block: a git `url` + required
`ref`, or a local `path`. It is declarable at the config level as a default
and overridable per eval, mirroring how `runs` already works.

`ref` is required on a git source because the runner records the resolved
SHA. An eval tracking a moving branch could not be re-run against the tree
it actually measured, which is the whole point of recording provenance.

The schema owns the structural contract, but `oneOf` cannot explain itself:
a git source missing its `ref` reports only that the block matched neither
branch, never naming `ref`. A small check ahead of the schema names the
mistakes worth a sentence, and covers the whitespace-only case that
`minLength: 1` admits.

Refs #252

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The codebase block needs turning into a real tree, and #253 needs the same
machinery for skills, so this lands as a shared module rather than inline in
the codebase path. Nothing in it knows what a codebase is.

Two phases, deliberately split. `resolve` is read-only, so a run fails on an
unreachable repository or a ref that does not exist before it has built any
part of a workspace. `materialize` then clones a source that has history, or
copies and initializes one that does not — either way the destination is a
Git repository with no remote, since a task environment must not be able to
reach the source it came from.

Two details worth naming, both pinned by tests:

`ls-remote` runs unfiltered. Passing a ref pattern suppresses the
`ref: refs/heads/<x>\tHEAD` line, and that line is the only way to learn the
remote's default branch — which is where a tag or a bare SHA has to land,
having no branch of its own. One unfiltered call answers both questions.

An annotated tag resolves through `refs/tags/<x>^{}` to the commit it peels
to. The tag object itself is not a commit and cannot be checked out as one.

A local path is materialized as a clean checkout of its committed state, so
uncommitted work in the source is not carried; resolution warns when the
source is dirty rather than letting that pass unnoticed.

Refs #252

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The task-repository assertion pinned `%aI` as `2000-01-01T00:00:00Z`, but
git renders a zero UTC offset as `+00:00` on 2.43 and `Z` only on newer
versions. The test therefore passed on CI and failed on any host with the
older git, for a difference in spelling rather than in behavior.

Normalize the offset before comparing, so the assertion stays exact about
the instant it cares about without pinning a git version.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Resolution now happens before anything is created, so an unreachable
repository or a ref that does not exist fails while the run has still built
nothing. Each distinct codebase is materialized once per iteration and every
`(group, condition, run)` environment is provisioned from that one tree;
`files` is copied on top, making it an overlay on a real project rather than
the whole of the environment.

The task-repository lifecycle had two invariants a real codebase breaks by
construction: it `git init`ed every environment from nothing, and it rejected
any remote. A sourced environment now keeps the `.git` its clone brought,
has its remotes stripped rather than asserted absent, and stays on the
branch the codebase itself was on. A fixture-only environment still starts
from `git init` on `work`, so evals that declare no codebase are untouched.

Both kinds now mark their start state with `refs/eval-magic/baseline`,
which #255 measures against. It sits outside `refs/heads/`, so it adds
nothing to what the agent under test sees.

The baseline `git add` drops `--force`. Forcing made sense when every file
in the environment was one the runner had placed; against a real repository
it would sweep `target/` or `node_modules/` into the state every run starts
from. The add now respects the codebase's `.gitignore`, and the paths the
runner placed — harness config directories and the fixture overlay — are
forced in on top, so a codebase that ignores `.claude/` cannot hide the
staged skill from the baseline and put the condition under test outside
every later diff.

Refs #252

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…base

Sourcing runs git against a URL from an eval config, on a host whose git
configuration belongs to someone else. Inherited, that configuration decides
things the runner has to decide itself.

`url.<base>.insteadOf` is the sharp one: it rewrites the URL, so the tree
sourced is not the tree the report cites — a silent wrong answer rather than
a failure. `init.templateDir` is the quiet one: it seeds hooks into a
repository the write guard assumes has none.

Every invocation in the module now runs with system and global configuration
switched off, the `GIT_CONFIG_COUNT` environment mechanism cleared, and an
empty template directory passed to `clone` and `init`.

Tested at the run boundary rather than in a unit test: the injection
mechanism is process-global environment variables, which a unit test cannot
set without racing every other test in the binary.

Refs #252

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A report that cites a codebase has to say which tree it measured, and the
declared ref cannot say it — a branch moves. `conditions.json` now carries
each distinct resolved codebase with the commit it resolved to and the evals
built from it, and every dispatch task carries the same record, which is the
route it takes to each run record.

One shape, `CodebaseRecord`, is shared by every surface so a reader never has
to reconcile two spellings of one resolution.

A `path` source is flagged `host_local`. Another machine has that directory
somewhere else, or nowhere, so a run citing it is not reproducible from the
config alone. Nothing can fix that, so the artifact states it instead of
implying a reproducibility it does not have — and where the directory is a
repository, its `origin` is recorded too, since `origin_url` + `revision`
does resolve anywhere.

Fixture-only iterations serialize unchanged: the field is omitted when empty.

Refs #252

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ercise

`mod.rs` reached 759 lines with 422 of them tests — the test module had grown
larger than the implementation it covers. CLAUDE.md's rule is a size trigger,
and this crossed it.

No behavior change; the same twelve tests run from a sibling file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…seline

Provenance stopped at `conditions.json` and `dispatch.json`, which is short of
where it is read. Grading consumes `run.json` and nothing else, so without the
record there a result cannot be tied to a tree at the granularity that
matters — the individual run. `benchmark.json` is the artifact a published
comparison is read from. `BASELINE.md` is what someone reads when deciding
whether to believe the claim.

All three now carry it, and both schemas gain the property (each is
`additionalProperties: false`, so the artifacts would otherwise fail their own
validation).

The `BASELINE.md` row names the resolved commit rather than the ref, since a
branch has moved by the time the baseline is read. A host-local path says so
in the cell and shows its origin URL, which is the part a reader elsewhere can
actually resolve.

Absent-when-empty throughout, so fixture-only artifacts are byte-identical.

Refs #252

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The codebase block has no CLI flag, so `--help` cannot carry its rules and a
config author has nowhere to discover them. It gets its own shipped topic,
which `build.rs` picks up from `docs/guides/`.

The guide leads with the parts that cannot be inferred from the schema: that a
git ref is mandatory and why, that `files` layers over the checkout rather
than replacing it, what the resulting repository looks like, and that a local
path is not reproducible by anyone reading the published results.

The isolation guide gains a short section separating the two boundaries it
would otherwise be read as covering: what a dispatch can load is not what a
dispatch can reach.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adding the property programmatically rewrote both files: every compact
one-line object was expanded, and the em dash in the run-record title was
escaped to `—` — a content change to a shipped description, buried in
450 lines of formatting churn.

Hand-written now, in the surrounding style. Both diffs are additions only.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The origin-citation test registered the remote with the host's native path
separators but asserted against the forward-slash form. On Linux those are
one string, so the mismatch was invisible; on Windows the assertion compared
a backslash path against a slash path and failed.

Git stores a remote URL byte-for-byte and eval-magic cites it unchanged, so
the fix is to hold both ends to the host's own spelling. That also makes the
assertion load-bearing on Windows: any separator normalization between the
source repo and conditions.json now shows up here rather than passing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
feat: source a real codebase into each task environment
The module's own docs say it knows nothing about what it is sourcing, but
every message it emits said "codebase". A skill resolved through it would
have reported `codebase path '...' is not a directory`, naming the one thing
the operator cannot act on. The caller now supplies the noun, and the
resolution carries it so materialization reads it back.

Also records `dirty` alongside the existing uncommitted-changes warning. A
warning is advice the operator may miss; the flag is evidence, and a subject
copied as it sits on disk cannot be cited without it. The probe is scoped
with `-- .` so a skill that is one directory among many in a repository is
not called dirty the moment some other skill is edited.

Part of #253.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`CodebaseRecord` is about to carry the skill under test as well as the
codebase a task environment is built from. Leaving it named for one of its
two subjects would misdescribe every use of the other.

Pure rename: `CodebaseUse` flattens the record, so no JSON key moves and
every artifact serializes byte-identically.

Part of #253.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Artifacts defaulted to `<cwd>/.eval-magic`, so running an eval from a skills
repository dropped iteration trees, envs, and benchmarks inside the very
repository under measurement.

The eval home now derives from the skill directory instead of the cwd:
`EVAL_MAGIC_WORKSPACE_DIR`, else `$XDG_DATA_HOME/eval-magic`, else
`~/.local/share/eval-magic` — mirroring the `EVAL_MAGIC_CONFIG_DIR` ladder
already used for descriptor layers. `--workspace-dir` still wins over both.

The derived default is namespaced by `<skill-dir-name>-<digest>`. Without it
a single global root would interleave the iterations of two skills that share
a name and come from different repositories, where `--iteration N` could
reach the wrong one. The digest is a hand-rolled FNV-1a rather than
`DefaultHasher`, which has no cross-release stability guarantee: this names a
directory operators re-type and generated commands embed, so a toolchain
upgrade must not silently relocate it.

An operator upgrading mid-campaign gets a notice naming the old directory and
the `--workspace-dir` value that keeps it reachable — suppressed when the
resolved root already is that directory, and when the only thing there is the
`harnesses/` descriptor layer, which is an unrelated use of the same name and
does not move.

Part of #253.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The codebase was a sourced, copied, SHA-recorded input while the skill was
read in place from wherever the operator's cwd happened to be — two mental
models in one command, and a report could pin the codebase commit while the
skill side was "whatever was on disk at the time".

The skill now resolves through the same resolver and is copied into
`iteration-N/.skills/`, a sibling of `.codebase/`. Every condition stages
from that copy, and the resolved source plus its revision reach
`conditions.json`, each `dispatch.json` task, each `run.json`,
`benchmark.json`, and `BASELINE.md`.

The copy is the working tree as it sits, not a checkout: Mode B's new arm
*is* the uncommitted edit under test, and Mode A's ordinary loop is
edit-then-run, so a committed-state copy would measure the wrong bytes.
`dirty` records when that happened and the warning says the run measured the
uncommitted work — the opposite of the codebase warning, where a clean
checkout leaves it behind. The resolver now reports only the fact; each
caller phrases the consequence it owns.

The sibling roster is recorded at resolution and staging copies exactly those
names, so what the artifacts claim and what the environments hold cannot
drift apart.

`promote-baseline` follows the recorded pointer rather than the operator's
current selection, so it still writes `<skill>/evals/baseline/` — and refuses
loudly when that skill has moved, rather than writing one skill's baseline
into another.

Also drops the dead `stage_root` override in `command_run`: it pointed at an
`env/` directory this layout stopped producing, and nothing inside `run` read
it back.

Part of #253.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`detect_live_source_reads` scanned agent shell commands for a bare relative
path to the live skill, computed from the operator's cwd. That only ever made
sense while the eval home sat inside the skill's own tree: with it outside,
a bare relative token in an agent command resolves against the environment
and cannot name the live skill.

The absolute-path branch stays. The live source still exists on disk and an
agent can still name it outright, so that remains a real finding. `repo_root`
stays with it — read-tool arguments may be relative and are resolved against
it.

The three "staged copy under .claude/skills is not flagged" tests go with the
branch: they existed to pin the config-dir lookbehind inside
`references_bare_rel`, and would have passed vacuously once it was gone,
which is worse than no coverage.

Also drops teardown's cwd sweep of staged skills. Staging is env-scoped —
`run` places nothing at the invocation cwd — so the sweep was residue from
when it did. The lifecycle test's `.claude` assertion moves to just after
`run`, where it can still fail if cwd staging ever comes back; after teardown
it could only pass. The cwd guard disarm stays, since `teardown-guard` is a
documented cwd-only command.

Part of #253.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`grade` read eval definitions and held-out command-check setup files from the
live skill tree. An edit between `run` and `grade` therefore changed what a
finished run was measured against, with nothing recording that it had — the
same provenance hole this ticket closes, one phase later.

Both now come from `iteration-N/.skills/<skill>`, falling back to the live
tree for iterations prepared before skills were sourced. Live-source
detection keeps the live path, which is the one thing it is looking for.

Part of #253.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`--mode revision` stages a snapshot in one arm and the live skill in the
other, so it is the mode where a half-applied change would hide. Pins that
both arms name something the runner placed inside the eval home, and that
revision runs record the skill source like new-skill runs do.

Part of #253.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The isolation guide explained what a dispatch can *load* but not what the
runner *places*, which is now the larger half of the story: everything the
agent can see is a copy, and the eval home lives outside the skill's own
repository. Adds a section covering the copy, what `skill_source` records,
why `dirty` matters before publishing, and why an absolute-path read of the
live directory is still a finding.

The codebase guide's verification snippet names the skill alongside the
codebase, since the two are recorded the same way; `--skill-dir` help says the
roster is captured at resolution; and `--help` says where artifacts land.

Part of #253.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Review follow-ups to the copied-input change.

`promote-baseline` writes into the skill the run recorded, but read "Promoted
from commit" from the operator's current selection, so the table could label a
baseline with a commit from a different repository. `git_cwd` existed to allow
exactly that difference and no caller ever wanted it; the commit now comes
from the tree the baseline lands in.

`--no-stage` populates no skills directory, so recording a roster of siblings
"staged alongside" the skill described an environment that never existed.

Extracts the `context` and `promote` test modules into sibling files, the
convention `adapters/guard` already uses: both had grown past the point where
the module fits the file it exercises.

Part of #253.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Grading reads `run.json` and nothing else, so the carrier from `dispatch.json`
into the record is the link that ties a graded result to a skill revision. It
had no test of its own; the codebase equivalent beside it did.

Part of #253.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A capability probe, not an assumption: two directories a run owns can sit on different filesystems, and link(2) is what says so. Any failure reads as unavailable so callers fall back to copying rather than provisioning wrong. The probe cleans up after itself — it runs inside the per-iteration codebase cache, where a leftover would ship into the next environment.
A run materializes each distinct codebase once per iteration; every (group, condition, run) environment is then provisioned from that single checkout. git clone --local is the fast path — Git hard-links the object store instead of copying it — and the origin remote a local clone adds is removed so no environment retains a path back to the cache. A commitless cache (an empty repository clones to an empty working tree) or a host that refuses the hard link takes a plain materialized copy instead.
…e cache

Replaces the per-environment byte copy of the cached checkout with source::provision_env, so --runs 10 against a real repository pays for one checkout instead of twenty full copies. The run plan now names each codebase and its resolved commit, the same shape as the skill-source line. Integration tests pin the contract: multi-run envs hard-link the shared object store, both revision-mode arms provision from one cache, a historyless codebase falls back to the copy, and no env retains a remote.
One section: one cached checkout per codebase per iteration, environments as local clones with hard-linked object stores and independent working trees, and the plain-copy fallback. The docs test pins the phrases a config author cannot infer.
slowdini and others added 28 commits September 1, 2026 19:48
The packaged profile allowed install, CI, test, build, lint, and typecheck,
but not npm run dev — and framework/nextjs only activates on a next
dependency, so a plain Vite + React project (the pinned Weeknight fixture)
had no packaged way to start its own dev server during a guarded run. dev
and start are generic lifecycle script names, not Next.js-specific, so they
move into the language profile for npm, pnpm, Yarn, and Bun.
`teardown-guard` swept only the invocation cwd, but the write guard now
arms inside each per-(group, condition) task env. Run mid-campaign it
printed "No write guard was installed — nothing to remove" while both
envs held live guards — the most costly moment for a false all-clear,
since the command exists for hand-editing files the guard would block.

It already accepted `CommonArgs`; the dispatch discarded them. Thread
them through and walk the iteration's staged envs the way `teardown` and
the `finalize` reminder already do, still without touching the staged
skill set or the workspace. Where those flags resolve no run, the sweep
now names the scopes it actually checked and warns that the env guards
were not among them.

Before:
    $ eval-magic teardown-guard --iteration 1
    No write guard was installed — nothing to remove.

After:
    $ eval-magic teardown-guard --skill demo --workspace-dir ws --iteration 1
    🛡 Write guard removed: 2 task envs in iteration 1.

    $ eval-magic teardown-guard          # from an unrelated cwd
    No write guard was installed — nothing to remove (checked the invocation cwd).
    ⚠ Task env guards were not checked, so any that were armed still are: …
       Add the run's target flags, or run `eval-magic teardown`.

Verified: cargo test, cargo clippy --all-targets -- -D warnings,
cargo fmt --check, plus a real run → teardown-guard against a scaffolded
skill.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NJjf8XMm1e1XtRZaGnkmvr
Every dispatch prompt designates `<eval-root>/tmp` for temporary work and
tells the agent to use it. An agent that complied then had everything it
put there measured as part of its change: in one prerelease run `tmp/`
accounted for 3 of 11 files touched and 228 of 296 lines added. That
trips `diff_scope` budgets on throwaway notes, puts scratch files in the
`diff.patch` a judge reads as the deliverable, and does so asymmetrically
— only in the arm that happened to use the directory it was told to use.

`.eval-magic-outputs/` already never counts, for the same reason. The
scratch directory now shares that treatment, on both surfaces: the env's
`.git/info/exclude` and each harness's `framework_ignore_paths`.

Both were spelling the outputs entry separately, and the diff-scope test
fixture spelled it a third time — so the fixture could not fail when the
rule changed. `sandbox::framework_owned_entries` is now the one
definition all three read.

Only files created under `tmp/` are affected: gitignore rules never apply
to tracked paths, so a codebase that genuinely tracks a `tmp/` directory
keeps its files measured.

Verified: cargo test, cargo clippy --all-targets -- -D warnings,
cargo fmt --check, plus a real `run` — a scratch file under `tmp/` is
invisible to `git status` in the staged env while the agent's own file
still shows, and both arms get the same `.prettierignore` block.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NJjf8XMm1e1XtRZaGnkmvr
Teardown's kept-iteration warning offered `.eval-magic/<skill>/` as the
directory to delete. That was right before the eval home moved out of the
skill repo; since then no such path exists, and the hint printed it one
line below a `promote-baseline` command carrying the correct absolute
`--workspace-dir`.

`ctx.workspace_root` was already in scope and already rendered correctly
by `command_target_args` in the same message.

Before:
    eval-magic promote-baseline … --workspace-dir /home/u/.local/share/eval-magic/skills-c61a1930 …
    or delete .eval-magic/working-with-tdd/ manually to discard.

After:
    or delete /home/u/.local/share/eval-magic/skills-c61a1930/working-with-tdd/ manually to discard.

The existing teardown test could not catch this: it runs with
EVAL_MAGIC_WORKSPACE_DIR=.eval-magic, where the two spellings coincide.
The new case puts the workspace elsewhere.

Verified: cargo test, cargo clippy --all-targets -- -D warnings,
cargo fmt --check.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NJjf8XMm1e1XtRZaGnkmvr
Review follow-ups on the #298 fixes.

`envs_phrase` read the iteration out of `checked` with `unwrap_or_default`,
which no call site can reach — but a future one would silently print
"iteration 0". Take the iteration as an argument so the invariant is in
the signature.

The task repository's exclude file now covers more than the outputs dir,
so its failure message says "framework path exclusion".

Verified: cargo test, cargo clippy --all-targets -- -D warnings,
cargo fmt --check.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NJjf8XMm1e1XtRZaGnkmvr
Gate held-out command checks on the runner-owned run record so partial ingest leaves undispatched environments untouched. Bind reusable results to both the assertion and run record, and warn when legacy artifacts imply possible contamination.
…and-checks

fix(grade): skip incomplete command checks
Record the post-preflight guard state in conditions and carry it through dispatch, aggregate benchmarks, and promoted baselines.

Preserve historical absence as unknown and document the artifact contract.
Record effective guard state in campaign provenance
`tool_invocation_matches` applied its regex only to the native invocation
rendering, so a frozen `Bash|Read` pattern scored zero across a Codex fleet
whose transcripts record `command_execution` — the same behavioral assertion
reporting different results because a harness picked a different tool name.

Match tool names portably instead, in two stages: the regex runs against the
native rendering, and on a miss the run's own descriptor supplies the role its
tool name belongs to while the registry-wide vocabulary union supplies every
portable spelling of that role. Only the name is substituted, so argument
regexes keep their behavior; a tool declared in no role gets no aliases; and
`assistant_message_matches` plus `must_precede` are unchanged. Nothing names a
harness, so a BYOH descriptor opts in through its `[tools]` table alone.

Evidence distinguishes the two: an alias match reports the invocation the
harness actually recorded and names the alias and role that matched it, and a
miss names the roles whose aliases were tried.

Closes #308.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018CGVCheSsLsP2jdHhtrtER
…script-checks

fix(grade): match transcript_check tool patterns by descriptor role
Persist each task's staged SKILL.md path and require a successful transcript read command against that exact path. Preserve native skill-tool checks and label no-signal judging as behavioral influence rather than invocation proof.\n\nCloses #309
)

The simulated plan mode injected `profiles/shared/plan-mode.md` as a
`<system-reminder>` into every dispatch prompt: run-wide, identical in
both arms, and unenforced. It predates runner-driven dispatch and native
session resume, both of which make a real two-phase session expressible.

An eval now declares `plan_mode: true`. The runner dispatches its opening
round with the harness's planning arguments, approves the presented plan
with one fixed message, and resumes the same session in act mode, where
`turns` and `responder` proceed unchanged. The approval is fixed rather
than judged so the transition is identical in every run and both arms.

Descriptors declare `[plan_mode]`: `plan_args` and `act_args` fill a
`{mode_args}` slot in both command templates, and `[plan_mode.plan_file]`
names the file the harness writes its plan to. Claude Code and OpenCode
declare it; `codex exec` has no plan flag and Cline cannot resume, so the
`run` preflight rejects plan-mode evals for both.

A plan is presented when the declared plan file is written, else when the
responder judges it ready, so a harness without a plan file requires a
responder. The approved plan is saved as `outputs/plan.md` and rendered
in the judge evidence bundle. `conversation.json` records each round's
mode, the runner-authored approval turn, and how the plan was approved.

Verified against claude 2.1.259 and opencode 1.18.10: `ExitPlanMode` is
disabled headless, so the plan-file write is the signal; the write guard
and the stray-write audit allow that root; and a write refused while
planning is recorded but attributed to plan mode, so it raises no
validity warning.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Four `tests/run` assertions compared a path built from `TempDir` against
the same path as the runner recorded it. The runner resolves paths before
recording them, so on a host whose temp directory sits behind a symlink —
macOS, where `/var` links to `/private/var` — the two spellings differ and
the assertion fails. eval-magic supports macOS, so the suite has to pass
there.

Each expected path now names the resolved spelling. The multi-skill case
failed as a detection miss rather than a path mismatch: it fabricates a
tool invocation for live-source detection to find, and detection compares
paths lexically, so an unresolved alias of the same file never matched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
feat(run): start plan-mode evals in the harness's native plan mode (#301)
The two sides of the live-source comparison are spelled by different
authors. The runner records the live skill directory through `real_path`,
so it is resolved; the transcript records whatever the agent typed, so it
is not. Both branches of `detect_live_source_reads` compared them
lexically — `is_under` for read tools, a substring scan for shell tools —
and a symlinked route to the same file matched neither.

macOS makes that route the ordinary one rather than an evasion: `$TMPDIR`
sits under `/var`, itself a link to `/private/var`, so an agent deriving
an absolute path from its own environment spells the directory a
different way than the runner recorded it. The miss is a false negative
in a contamination check, so `aggregate` drops the validity warning and a
tainted arm reports clean.

Resolve symlinks on both sides, as a union with the lexical answer so
nothing detected before stops being detected:

- `resolve_existing_ancestor` splits out of `real_path`. A caller
  comparing agent-recorded paths must not take that function's leading
  `std::path::absolute`, which grafts the process's drive onto a
  rooted-but-prefixless path — the case `lexically_absolute` exists to
  prevent.
- `is_under_through_links` joins `is_under` in the boundary policy and
  serves the read branch.
- The shell branch keeps its raw scan, which reaches spellings no path
  resolution can, and gains a per-word test over the existing `lex_shell`.
  Only words carrying a separator are resolved: every word resolves
  against the runner's cwd, so testing bare ones would make each
  `cargo test` a finding whenever that cwd sits inside the live directory.

The write boundary and the live guard keep the lexical rule. Both sides
there are already resolved — env roots descend from the `real_path`-ed
workspace root, and the harness's cwd is that resolved directory — so
resolving links would change runtime denial decisions for no known defect.

Closes #317.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H1iqzMVs15eSHAd9PTu6Wu
The module and its inline tests together crossed a thousand lines, which
is past what anyone holds in their head at once. Extraction is the size
decision the repository guidelines describe, and this module already has
a directory for the sibling its `realistic_development_tests` lives in.

A verbatim move: the module file keeps the code half, the sibling keeps
the tests, and neither gains or loses a case.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H1iqzMVs15eSHAd9PTu6Wu
…liases

fix(pipeline): match aliased spellings in live-source-read detection (#317)
Measure eval-agent subprocess time at the universal runner boundary so every harness records a comparable duration. Preserve independent token and duration provenance while keeping historical timing records readable.
feat(timing): persist runner-measured duration
The target-aware classifier recognized package installs, pip installs,
Cargo build output, and `sed -i`. Every one of those names its destination
implicitly, through the invocation cwd or a path option. The commands that
name a destination outright — `touch`, `mkdir`, `rm`, `cp`, `mv`,
`install` — were not read at all, so `touch /private/tmp/probe` was
neither blocked by the live guard nor reported by `detect-stray-writes`.
Both surfaces consume the same classifier, so this was one blind spot
rather than a divergence between them.

Coreutils spell their destinations in only a handful of shapes, so the six
families become a `MutatorSpec` table in a new `mutation_targets::filesystem`
module and one shared walker reads them all: which positionals are written,
which options carry a required argument that is an operand rather than a
target, the `-t`/`--target-directory` destination, `install -d`, and the
sources a `mv` unlinks along the way. Teaching the guard another mutator is
a table row, not another hand-written classifier.

Two things the walker has to get right, both driven by what the shell would
actually do:

- The executable is resolved rather than searched for. `command_position`
  matches a name anywhere in a segment, which is how `sudo npm install` is
  caught for free — but `install` is also the action word of every package
  manager, and `mkdir` is an ordinary English word. A parallel `WrapperSpec`
  table steps past leading assignments and wrappers to the word being run,
  so `npm install` stays a package install and `echo mkdir /outside` stays
  an echo.

- Bundled short clusters follow getopt's own rule: a short option whose
  argument is required takes the rest of its cluster (`-m755`, `-tr`), or
  the following word when the cluster ends at it (`-rt DIR`, `-Dm 755`).
  Reading them any other way lets a mode or a suffix swallow the
  destination. Options with an *optional* argument are deliberately absent
  from the table for the same reason: getopt reads those only from the
  `--long=value` spelling, so treating `--backup` as consuming a word would
  hide the destination behind it.

Denying every expansion would have cost more than it bought — `rm -rf
target/*` and `mkdir -p build/{a,b}` are ordinary in a development shell.
A glob draws its result from a directory listing, so it cannot land outside
the directory its literal prefix names, provided the pattern stays within
one path component and does not begin with the `.` that would let it match
`..`. The lexer now records where a word stopped being literal and what
took over there, which is what makes that judgement possible; a variable or
command expansion names no directory and still fails closed with no
evidence, since a half-expanded spelling is not evidence.

These families are recognized for containment only. They are commands an
agent legitimately needs inside its own environment, so they stay out of
`segment_is_recognized` and an in-root `mkdir` or `cp` needs no allowance
the way `npm install` and `sed -i` do. A denial now earns the same
scratch-directory hint a blocked redirect gets, because it is the same
mistake with the same fix.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CwrS1WvcU7EiopAC4xvfsf
feat(sandbox): contain the shell commands that write paths outright
docs(codex): explain nested dispatch sandbox constraints
@slowdini
slowdini merged commit ec27205 into main Sep 5, 2026
9 checks passed
slowdini added a commit that referenced this pull request Sep 5, 2026
Merge pull request #322 from slowdini/dev
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant