Conversation
Merge pull request #276 from slowdini/dev
An eval environment could only be assembled from individual fixture files copied out of `<skill>/evals/`, so there was no way to say "run this task against *this project*". Add a `codebase` block: a git `url` + required `ref`, or a local `path`. It is declarable at the config level as a default and overridable per eval, mirroring how `runs` already works. `ref` is required on a git source because the runner records the resolved SHA. An eval tracking a moving branch could not be re-run against the tree it actually measured, which is the whole point of recording provenance. The schema owns the structural contract, but `oneOf` cannot explain itself: a git source missing its `ref` reports only that the block matched neither branch, never naming `ref`. A small check ahead of the schema names the mistakes worth a sentence, and covers the whitespace-only case that `minLength: 1` admits. Refs #252 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The codebase block needs turning into a real tree, and #253 needs the same machinery for skills, so this lands as a shared module rather than inline in the codebase path. Nothing in it knows what a codebase is. Two phases, deliberately split. `resolve` is read-only, so a run fails on an unreachable repository or a ref that does not exist before it has built any part of a workspace. `materialize` then clones a source that has history, or copies and initializes one that does not — either way the destination is a Git repository with no remote, since a task environment must not be able to reach the source it came from. Two details worth naming, both pinned by tests: `ls-remote` runs unfiltered. Passing a ref pattern suppresses the `ref: refs/heads/<x>\tHEAD` line, and that line is the only way to learn the remote's default branch — which is where a tag or a bare SHA has to land, having no branch of its own. One unfiltered call answers both questions. An annotated tag resolves through `refs/tags/<x>^{}` to the commit it peels to. The tag object itself is not a commit and cannot be checked out as one. A local path is materialized as a clean checkout of its committed state, so uncommitted work in the source is not carried; resolution warns when the source is dirty rather than letting that pass unnoticed. Refs #252 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The task-repository assertion pinned `%aI` as `2000-01-01T00:00:00Z`, but git renders a zero UTC offset as `+00:00` on 2.43 and `Z` only on newer versions. The test therefore passed on CI and failed on any host with the older git, for a difference in spelling rather than in behavior. Normalize the offset before comparing, so the assertion stays exact about the instant it cares about without pinning a git version. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Resolution now happens before anything is created, so an unreachable repository or a ref that does not exist fails while the run has still built nothing. Each distinct codebase is materialized once per iteration and every `(group, condition, run)` environment is provisioned from that one tree; `files` is copied on top, making it an overlay on a real project rather than the whole of the environment. The task-repository lifecycle had two invariants a real codebase breaks by construction: it `git init`ed every environment from nothing, and it rejected any remote. A sourced environment now keeps the `.git` its clone brought, has its remotes stripped rather than asserted absent, and stays on the branch the codebase itself was on. A fixture-only environment still starts from `git init` on `work`, so evals that declare no codebase are untouched. Both kinds now mark their start state with `refs/eval-magic/baseline`, which #255 measures against. It sits outside `refs/heads/`, so it adds nothing to what the agent under test sees. The baseline `git add` drops `--force`. Forcing made sense when every file in the environment was one the runner had placed; against a real repository it would sweep `target/` or `node_modules/` into the state every run starts from. The add now respects the codebase's `.gitignore`, and the paths the runner placed — harness config directories and the fixture overlay — are forced in on top, so a codebase that ignores `.claude/` cannot hide the staged skill from the baseline and put the condition under test outside every later diff. Refs #252 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…base Sourcing runs git against a URL from an eval config, on a host whose git configuration belongs to someone else. Inherited, that configuration decides things the runner has to decide itself. `url.<base>.insteadOf` is the sharp one: it rewrites the URL, so the tree sourced is not the tree the report cites — a silent wrong answer rather than a failure. `init.templateDir` is the quiet one: it seeds hooks into a repository the write guard assumes has none. Every invocation in the module now runs with system and global configuration switched off, the `GIT_CONFIG_COUNT` environment mechanism cleared, and an empty template directory passed to `clone` and `init`. Tested at the run boundary rather than in a unit test: the injection mechanism is process-global environment variables, which a unit test cannot set without racing every other test in the binary. Refs #252 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A report that cites a codebase has to say which tree it measured, and the declared ref cannot say it — a branch moves. `conditions.json` now carries each distinct resolved codebase with the commit it resolved to and the evals built from it, and every dispatch task carries the same record, which is the route it takes to each run record. One shape, `CodebaseRecord`, is shared by every surface so a reader never has to reconcile two spellings of one resolution. A `path` source is flagged `host_local`. Another machine has that directory somewhere else, or nowhere, so a run citing it is not reproducible from the config alone. Nothing can fix that, so the artifact states it instead of implying a reproducibility it does not have — and where the directory is a repository, its `origin` is recorded too, since `origin_url` + `revision` does resolve anywhere. Fixture-only iterations serialize unchanged: the field is omitted when empty. Refs #252 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ercise `mod.rs` reached 759 lines with 422 of them tests — the test module had grown larger than the implementation it covers. CLAUDE.md's rule is a size trigger, and this crossed it. No behavior change; the same twelve tests run from a sibling file. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…seline Provenance stopped at `conditions.json` and `dispatch.json`, which is short of where it is read. Grading consumes `run.json` and nothing else, so without the record there a result cannot be tied to a tree at the granularity that matters — the individual run. `benchmark.json` is the artifact a published comparison is read from. `BASELINE.md` is what someone reads when deciding whether to believe the claim. All three now carry it, and both schemas gain the property (each is `additionalProperties: false`, so the artifacts would otherwise fail their own validation). The `BASELINE.md` row names the resolved commit rather than the ref, since a branch has moved by the time the baseline is read. A host-local path says so in the cell and shows its origin URL, which is the part a reader elsewhere can actually resolve. Absent-when-empty throughout, so fixture-only artifacts are byte-identical. Refs #252 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The codebase block has no CLI flag, so `--help` cannot carry its rules and a config author has nowhere to discover them. It gets its own shipped topic, which `build.rs` picks up from `docs/guides/`. The guide leads with the parts that cannot be inferred from the schema: that a git ref is mandatory and why, that `files` layers over the checkout rather than replacing it, what the resulting repository looks like, and that a local path is not reproducible by anyone reading the published results. The isolation guide gains a short section separating the two boundaries it would otherwise be read as covering: what a dispatch can load is not what a dispatch can reach. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adding the property programmatically rewrote both files: every compact one-line object was expanded, and the em dash in the run-record title was escaped to `—` — a content change to a shipped description, buried in 450 lines of formatting churn. Hand-written now, in the surrounding style. Both diffs are additions only. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The origin-citation test registered the remote with the host's native path separators but asserted against the forward-slash form. On Linux those are one string, so the mismatch was invisible; on Windows the assertion compared a backslash path against a slash path and failed. Git stores a remote URL byte-for-byte and eval-magic cites it unchanged, so the fix is to hold both ends to the host's own spelling. That also makes the assertion load-bearing on Windows: any separator normalization between the source repo and conditions.json now shows up here rather than passing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
feat: source a real codebase into each task environment
The module's own docs say it knows nothing about what it is sourcing, but every message it emits said "codebase". A skill resolved through it would have reported `codebase path '...' is not a directory`, naming the one thing the operator cannot act on. The caller now supplies the noun, and the resolution carries it so materialization reads it back. Also records `dirty` alongside the existing uncommitted-changes warning. A warning is advice the operator may miss; the flag is evidence, and a subject copied as it sits on disk cannot be cited without it. The probe is scoped with `-- .` so a skill that is one directory among many in a repository is not called dirty the moment some other skill is edited. Part of #253. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`CodebaseRecord` is about to carry the skill under test as well as the codebase a task environment is built from. Leaving it named for one of its two subjects would misdescribe every use of the other. Pure rename: `CodebaseUse` flattens the record, so no JSON key moves and every artifact serializes byte-identically. Part of #253. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Artifacts defaulted to `<cwd>/.eval-magic`, so running an eval from a skills repository dropped iteration trees, envs, and benchmarks inside the very repository under measurement. The eval home now derives from the skill directory instead of the cwd: `EVAL_MAGIC_WORKSPACE_DIR`, else `$XDG_DATA_HOME/eval-magic`, else `~/.local/share/eval-magic` — mirroring the `EVAL_MAGIC_CONFIG_DIR` ladder already used for descriptor layers. `--workspace-dir` still wins over both. The derived default is namespaced by `<skill-dir-name>-<digest>`. Without it a single global root would interleave the iterations of two skills that share a name and come from different repositories, where `--iteration N` could reach the wrong one. The digest is a hand-rolled FNV-1a rather than `DefaultHasher`, which has no cross-release stability guarantee: this names a directory operators re-type and generated commands embed, so a toolchain upgrade must not silently relocate it. An operator upgrading mid-campaign gets a notice naming the old directory and the `--workspace-dir` value that keeps it reachable — suppressed when the resolved root already is that directory, and when the only thing there is the `harnesses/` descriptor layer, which is an unrelated use of the same name and does not move. Part of #253. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The codebase was a sourced, copied, SHA-recorded input while the skill was read in place from wherever the operator's cwd happened to be — two mental models in one command, and a report could pin the codebase commit while the skill side was "whatever was on disk at the time". The skill now resolves through the same resolver and is copied into `iteration-N/.skills/`, a sibling of `.codebase/`. Every condition stages from that copy, and the resolved source plus its revision reach `conditions.json`, each `dispatch.json` task, each `run.json`, `benchmark.json`, and `BASELINE.md`. The copy is the working tree as it sits, not a checkout: Mode B's new arm *is* the uncommitted edit under test, and Mode A's ordinary loop is edit-then-run, so a committed-state copy would measure the wrong bytes. `dirty` records when that happened and the warning says the run measured the uncommitted work — the opposite of the codebase warning, where a clean checkout leaves it behind. The resolver now reports only the fact; each caller phrases the consequence it owns. The sibling roster is recorded at resolution and staging copies exactly those names, so what the artifacts claim and what the environments hold cannot drift apart. `promote-baseline` follows the recorded pointer rather than the operator's current selection, so it still writes `<skill>/evals/baseline/` — and refuses loudly when that skill has moved, rather than writing one skill's baseline into another. Also drops the dead `stage_root` override in `command_run`: it pointed at an `env/` directory this layout stopped producing, and nothing inside `run` read it back. Part of #253. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`detect_live_source_reads` scanned agent shell commands for a bare relative path to the live skill, computed from the operator's cwd. That only ever made sense while the eval home sat inside the skill's own tree: with it outside, a bare relative token in an agent command resolves against the environment and cannot name the live skill. The absolute-path branch stays. The live source still exists on disk and an agent can still name it outright, so that remains a real finding. `repo_root` stays with it — read-tool arguments may be relative and are resolved against it. The three "staged copy under .claude/skills is not flagged" tests go with the branch: they existed to pin the config-dir lookbehind inside `references_bare_rel`, and would have passed vacuously once it was gone, which is worse than no coverage. Also drops teardown's cwd sweep of staged skills. Staging is env-scoped — `run` places nothing at the invocation cwd — so the sweep was residue from when it did. The lifecycle test's `.claude` assertion moves to just after `run`, where it can still fail if cwd staging ever comes back; after teardown it could only pass. The cwd guard disarm stays, since `teardown-guard` is a documented cwd-only command. Part of #253. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`grade` read eval definitions and held-out command-check setup files from the live skill tree. An edit between `run` and `grade` therefore changed what a finished run was measured against, with nothing recording that it had — the same provenance hole this ticket closes, one phase later. Both now come from `iteration-N/.skills/<skill>`, falling back to the live tree for iterations prepared before skills were sourced. Live-source detection keeps the live path, which is the one thing it is looking for. Part of #253. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`--mode revision` stages a snapshot in one arm and the live skill in the other, so it is the mode where a half-applied change would hide. Pins that both arms name something the runner placed inside the eval home, and that revision runs record the skill source like new-skill runs do. Part of #253. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The isolation guide explained what a dispatch can *load* but not what the runner *places*, which is now the larger half of the story: everything the agent can see is a copy, and the eval home lives outside the skill's own repository. Adds a section covering the copy, what `skill_source` records, why `dirty` matters before publishing, and why an absolute-path read of the live directory is still a finding. The codebase guide's verification snippet names the skill alongside the codebase, since the two are recorded the same way; `--skill-dir` help says the roster is captured at resolution; and `--help` says where artifacts land. Part of #253. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Review follow-ups to the copied-input change. `promote-baseline` writes into the skill the run recorded, but read "Promoted from commit" from the operator's current selection, so the table could label a baseline with a commit from a different repository. `git_cwd` existed to allow exactly that difference and no caller ever wanted it; the commit now comes from the tree the baseline lands in. `--no-stage` populates no skills directory, so recording a roster of siblings "staged alongside" the skill described an environment that never existed. Extracts the `context` and `promote` test modules into sibling files, the convention `adapters/guard` already uses: both had grown past the point where the module fits the file it exercises. Part of #253. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Grading reads `run.json` and nothing else, so the carrier from `dispatch.json` into the record is the link that ties a graded result to a skill revision. It had no test of its own; the codebase equivalent beside it did. Part of #253. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Feat/skill source
A capability probe, not an assumption: two directories a run owns can sit on different filesystems, and link(2) is what says so. Any failure reads as unavailable so callers fall back to copying rather than provisioning wrong. The probe cleans up after itself — it runs inside the per-iteration codebase cache, where a leftover would ship into the next environment.
A run materializes each distinct codebase once per iteration; every (group, condition, run) environment is then provisioned from that single checkout. git clone --local is the fast path — Git hard-links the object store instead of copying it — and the origin remote a local clone adds is removed so no environment retains a path back to the cache. A commitless cache (an empty repository clones to an empty working tree) or a host that refuses the hard link takes a plain materialized copy instead.
…e cache Replaces the per-environment byte copy of the cached checkout with source::provision_env, so --runs 10 against a real repository pays for one checkout instead of twenty full copies. The run plan now names each codebase and its resolved commit, the same shape as the skill-source line. Integration tests pin the contract: multi-run envs hard-link the shared object store, both revision-mode arms provision from one cache, a historyless codebase falls back to the copy, and no env retains a remote.
One section: one cached checkout per codebase per iteration, environments as local clones with hard-linked object stores and independent working trees, and the plain-copy fallback. The docs test pins the phrases a config author cannot infer.
The packaged profile allowed install, CI, test, build, lint, and typecheck, but not npm run dev — and framework/nextjs only activates on a next dependency, so a plain Vite + React project (the pinned Weeknight fixture) had no packaged way to start its own dev server during a guarded run. dev and start are generic lifecycle script names, not Next.js-specific, so they move into the language profile for npm, pnpm, Yarn, and Bun.
Issue 297 guard dev server verdict
`teardown-guard` swept only the invocation cwd, but the write guard now
arms inside each per-(group, condition) task env. Run mid-campaign it
printed "No write guard was installed — nothing to remove" while both
envs held live guards — the most costly moment for a false all-clear,
since the command exists for hand-editing files the guard would block.
It already accepted `CommonArgs`; the dispatch discarded them. Thread
them through and walk the iteration's staged envs the way `teardown` and
the `finalize` reminder already do, still without touching the staged
skill set or the workspace. Where those flags resolve no run, the sweep
now names the scopes it actually checked and warns that the env guards
were not among them.
Before:
$ eval-magic teardown-guard --iteration 1
No write guard was installed — nothing to remove.
After:
$ eval-magic teardown-guard --skill demo --workspace-dir ws --iteration 1
🛡 Write guard removed: 2 task envs in iteration 1.
$ eval-magic teardown-guard # from an unrelated cwd
No write guard was installed — nothing to remove (checked the invocation cwd).
⚠ Task env guards were not checked, so any that were armed still are: …
Add the run's target flags, or run `eval-magic teardown`.
Verified: cargo test, cargo clippy --all-targets -- -D warnings,
cargo fmt --check, plus a real run → teardown-guard against a scaffolded
skill.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NJjf8XMm1e1XtRZaGnkmvr
Every dispatch prompt designates `<eval-root>/tmp` for temporary work and tells the agent to use it. An agent that complied then had everything it put there measured as part of its change: in one prerelease run `tmp/` accounted for 3 of 11 files touched and 228 of 296 lines added. That trips `diff_scope` budgets on throwaway notes, puts scratch files in the `diff.patch` a judge reads as the deliverable, and does so asymmetrically — only in the arm that happened to use the directory it was told to use. `.eval-magic-outputs/` already never counts, for the same reason. The scratch directory now shares that treatment, on both surfaces: the env's `.git/info/exclude` and each harness's `framework_ignore_paths`. Both were spelling the outputs entry separately, and the diff-scope test fixture spelled it a third time — so the fixture could not fail when the rule changed. `sandbox::framework_owned_entries` is now the one definition all three read. Only files created under `tmp/` are affected: gitignore rules never apply to tracked paths, so a codebase that genuinely tracks a `tmp/` directory keeps its files measured. Verified: cargo test, cargo clippy --all-targets -- -D warnings, cargo fmt --check, plus a real `run` — a scratch file under `tmp/` is invisible to `git status` in the staged env while the agent's own file still shows, and both arms get the same `.prettierignore` block. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NJjf8XMm1e1XtRZaGnkmvr
Teardown's kept-iteration warning offered `.eval-magic/<skill>/` as the
directory to delete. That was right before the eval home moved out of the
skill repo; since then no such path exists, and the hint printed it one
line below a `promote-baseline` command carrying the correct absolute
`--workspace-dir`.
`ctx.workspace_root` was already in scope and already rendered correctly
by `command_target_args` in the same message.
Before:
eval-magic promote-baseline … --workspace-dir /home/u/.local/share/eval-magic/skills-c61a1930 …
or delete .eval-magic/working-with-tdd/ manually to discard.
After:
or delete /home/u/.local/share/eval-magic/skills-c61a1930/working-with-tdd/ manually to discard.
The existing teardown test could not catch this: it runs with
EVAL_MAGIC_WORKSPACE_DIR=.eval-magic, where the two spellings coincide.
The new case puts the workspace elsewhere.
Verified: cargo test, cargo clippy --all-targets -- -D warnings,
cargo fmt --check.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NJjf8XMm1e1XtRZaGnkmvr
Review follow-ups on the #298 fixes. `envs_phrase` read the iteration out of `checked` with `unwrap_or_default`, which no call site can reach — but a future one would silently print "iteration 0". Take the iteration as an argument so the invariant is in the signature. The task repository's exclude file now covers more than the outputs dir, so its failure message says "framework path exclusion". Verified: cargo test, cargo clippy --all-targets -- -D warnings, cargo fmt --check. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NJjf8XMm1e1XtRZaGnkmvr
Issue 298 prerelease fixes
Gate held-out command checks on the runner-owned run record so partial ingest leaves undispatched environments untouched. Bind reusable results to both the assertion and run record, and warn when legacy artifacts imply possible contamination.
…and-checks fix(grade): skip incomplete command checks
Record the post-preflight guard state in conditions and carry it through dispatch, aggregate benchmarks, and promoted baselines. Preserve historical absence as unknown and document the artifact contract.
Record effective guard state in campaign provenance
`tool_invocation_matches` applied its regex only to the native invocation rendering, so a frozen `Bash|Read` pattern scored zero across a Codex fleet whose transcripts record `command_execution` — the same behavioral assertion reporting different results because a harness picked a different tool name. Match tool names portably instead, in two stages: the regex runs against the native rendering, and on a miss the run's own descriptor supplies the role its tool name belongs to while the registry-wide vocabulary union supplies every portable spelling of that role. Only the name is substituted, so argument regexes keep their behavior; a tool declared in no role gets no aliases; and `assistant_message_matches` plus `must_precede` are unchanged. Nothing names a harness, so a BYOH descriptor opts in through its `[tools]` table alone. Evidence distinguishes the two: an alias match reports the invocation the harness actually recorded and names the alias and role that matched it, and a miss names the roles whose aliases were tried. Closes #308. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018CGVCheSsLsP2jdHhtrtER
…script-checks fix(grade): match transcript_check tool patterns by descriptor role
Persist each task's staged SKILL.md path and require a successful transcript read command against that exact path. Preserve native skill-tool checks and label no-signal judging as behavioral influence rather than invocation proof.\n\nCloses #309
fix(codex): verify staged skill access
) The simulated plan mode injected `profiles/shared/plan-mode.md` as a `<system-reminder>` into every dispatch prompt: run-wide, identical in both arms, and unenforced. It predates runner-driven dispatch and native session resume, both of which make a real two-phase session expressible. An eval now declares `plan_mode: true`. The runner dispatches its opening round with the harness's planning arguments, approves the presented plan with one fixed message, and resumes the same session in act mode, where `turns` and `responder` proceed unchanged. The approval is fixed rather than judged so the transition is identical in every run and both arms. Descriptors declare `[plan_mode]`: `plan_args` and `act_args` fill a `{mode_args}` slot in both command templates, and `[plan_mode.plan_file]` names the file the harness writes its plan to. Claude Code and OpenCode declare it; `codex exec` has no plan flag and Cline cannot resume, so the `run` preflight rejects plan-mode evals for both. A plan is presented when the declared plan file is written, else when the responder judges it ready, so a harness without a plan file requires a responder. The approved plan is saved as `outputs/plan.md` and rendered in the judge evidence bundle. `conversation.json` records each round's mode, the runner-authored approval turn, and how the plan was approved. Verified against claude 2.1.259 and opencode 1.18.10: `ExitPlanMode` is disabled headless, so the plan-file write is the signal; the write guard and the stray-write audit allow that root; and a write refused while planning is recorded but attributed to plan mode, so it raises no validity warning. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Four `tests/run` assertions compared a path built from `TempDir` against the same path as the runner recorded it. The runner resolves paths before recording them, so on a host whose temp directory sits behind a symlink — macOS, where `/var` links to `/private/var` — the two spellings differ and the assertion fails. eval-magic supports macOS, so the suite has to pass there. Each expected path now names the resolved spelling. The multi-skill case failed as a detection miss rather than a path mismatch: it fabricates a tool invocation for live-source detection to find, and detection compares paths lexically, so an unresolved alias of the same file never matched. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
feat(run): start plan-mode evals in the harness's native plan mode (#301)
The two sides of the live-source comparison are spelled by different authors. The runner records the live skill directory through `real_path`, so it is resolved; the transcript records whatever the agent typed, so it is not. Both branches of `detect_live_source_reads` compared them lexically — `is_under` for read tools, a substring scan for shell tools — and a symlinked route to the same file matched neither. macOS makes that route the ordinary one rather than an evasion: `$TMPDIR` sits under `/var`, itself a link to `/private/var`, so an agent deriving an absolute path from its own environment spells the directory a different way than the runner recorded it. The miss is a false negative in a contamination check, so `aggregate` drops the validity warning and a tainted arm reports clean. Resolve symlinks on both sides, as a union with the lexical answer so nothing detected before stops being detected: - `resolve_existing_ancestor` splits out of `real_path`. A caller comparing agent-recorded paths must not take that function's leading `std::path::absolute`, which grafts the process's drive onto a rooted-but-prefixless path — the case `lexically_absolute` exists to prevent. - `is_under_through_links` joins `is_under` in the boundary policy and serves the read branch. - The shell branch keeps its raw scan, which reaches spellings no path resolution can, and gains a per-word test over the existing `lex_shell`. Only words carrying a separator are resolved: every word resolves against the runner's cwd, so testing bare ones would make each `cargo test` a finding whenever that cwd sits inside the live directory. The write boundary and the live guard keep the lexical rule. Both sides there are already resolved — env roots descend from the `real_path`-ed workspace root, and the harness's cwd is that resolved directory — so resolving links would change runtime denial decisions for no known defect. Closes #317. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H1iqzMVs15eSHAd9PTu6Wu
The module and its inline tests together crossed a thousand lines, which is past what anyone holds in their head at once. Extraction is the size decision the repository guidelines describe, and this module already has a directory for the sibling its `realistic_development_tests` lives in. A verbatim move: the module file keeps the code half, the sibling keeps the tests, and neither gains or loses a case. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H1iqzMVs15eSHAd9PTu6Wu
…liases fix(pipeline): match aliased spellings in live-source-read detection (#317)
Measure eval-agent subprocess time at the universal runner boundary so every harness records a comparable duration. Preserve independent token and duration provenance while keeping historical timing records readable.
feat(timing): persist runner-measured duration
The target-aware classifier recognized package installs, pip installs,
Cargo build output, and `sed -i`. Every one of those names its destination
implicitly, through the invocation cwd or a path option. The commands that
name a destination outright — `touch`, `mkdir`, `rm`, `cp`, `mv`,
`install` — were not read at all, so `touch /private/tmp/probe` was
neither blocked by the live guard nor reported by `detect-stray-writes`.
Both surfaces consume the same classifier, so this was one blind spot
rather than a divergence between them.
Coreutils spell their destinations in only a handful of shapes, so the six
families become a `MutatorSpec` table in a new `mutation_targets::filesystem`
module and one shared walker reads them all: which positionals are written,
which options carry a required argument that is an operand rather than a
target, the `-t`/`--target-directory` destination, `install -d`, and the
sources a `mv` unlinks along the way. Teaching the guard another mutator is
a table row, not another hand-written classifier.
Two things the walker has to get right, both driven by what the shell would
actually do:
- The executable is resolved rather than searched for. `command_position`
matches a name anywhere in a segment, which is how `sudo npm install` is
caught for free — but `install` is also the action word of every package
manager, and `mkdir` is an ordinary English word. A parallel `WrapperSpec`
table steps past leading assignments and wrappers to the word being run,
so `npm install` stays a package install and `echo mkdir /outside` stays
an echo.
- Bundled short clusters follow getopt's own rule: a short option whose
argument is required takes the rest of its cluster (`-m755`, `-tr`), or
the following word when the cluster ends at it (`-rt DIR`, `-Dm 755`).
Reading them any other way lets a mode or a suffix swallow the
destination. Options with an *optional* argument are deliberately absent
from the table for the same reason: getopt reads those only from the
`--long=value` spelling, so treating `--backup` as consuming a word would
hide the destination behind it.
Denying every expansion would have cost more than it bought — `rm -rf
target/*` and `mkdir -p build/{a,b}` are ordinary in a development shell.
A glob draws its result from a directory listing, so it cannot land outside
the directory its literal prefix names, provided the pattern stays within
one path component and does not begin with the `.` that would let it match
`..`. The lexer now records where a word stopped being literal and what
took over there, which is what makes that judgement possible; a variable or
command expansion names no directory and still fails closed with no
evidence, since a half-expanded spelling is not evidence.
These families are recognized for containment only. They are commands an
agent legitimately needs inside its own environment, so they stay out of
`segment_is_recognized` and an in-root `mkdir` or `cp` needs no allowance
the way `npm install` and `sed -i` do. A denial now earns the same
scratch-directory hint a blocked redirect gets, because it is the same
mistake with the same fix.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CwrS1WvcU7EiopAC4xvfsf
feat(sandbox): contain the shell commands that write paths outright
docs(codex): explain nested dispatch sandbox constraints
slowdini
added a commit
that referenced
this pull request
Sep 5, 2026
Merge pull request #322 from slowdini/dev
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Release notes
Highlights
directory, and run each condition in its own working tree. Codebase caching avoids repeated
full clones; copied skill inputs and recorded revisions make the treatment traceable through
runs, benchmarks, and promoted baselines. (#278,
#279, #280)
eval-magic dispatchruns tasks and judges with boundedparallelism, task timeouts, and selective retries. An LLM responder can answer follow-up
questions without seeing grading criteria. Native planning sessions can proceed through a
recorded approval into implementation on Claude Code and OpenCode.
(#282, #285,
#316)
skill_namelist defines one treatment, withinvocation or access checks for each member and provenance for the complete set. Existing
scalar configurations retain their single-skill artifact format.
(#291)
judge-evidence.mdbundles preserve the task, conversation, and code changes behind a verdict.eval-magic comparepresents paired evidence even without authored assertions; promotedbaselines retain their evidence. Comparison reports are exploratory, not grades.
(#281, #288,
#290)
--judge-samples Nand per-assertionsamplesrequest repeatedverdicts over the same persisted evidence. Reports retain individual votes, pass proportions,
and pass^k consistency scores. Sampling does not rerun the evaluated agent; one sample keeps
the existing binary grading format. (#289)
command allowances with Rust, JavaScript, Python, and Next.js profiles. Containment checks
remain non-overridable. The expanded policy and effective guard state are recorded for
enforcement, auditing, and baseline review.
(#286, #313)
time across rounds, excluding responder consultations, judges, and queue time. Timing artifacts
identify token and duration sources separately; missing historical measurements remain missing.
(#319)
Behavior changes to know about
codebasedefault or an eval-level override,and remove the retired
isolationfield.filesandfiles_rootsupply overlays on thatcodebase.
initdefaults to the pinned Weeknight fixture; choose another source with--codebase-urlplus--codebase-ref,--codebase-path, or--codebase-cwd.See
eval-magic docs codebase.(#292, #293)
under
$XDG_DATA_HOME/eval-magicor~/.local/share/eval-magic. Use--workspace-dirorEVAL_MAGIC_WORKSPACE_DIRto select another location or continue an existing campaign.eval-magicsupports Linux and macOS. On Windows, runeval-magicinside WSL; native Windows is unsupported. Git and a POSIX shell are required. Set EVAL_MAGIC_SH
to select a specific
sh. Runner dispatch no longer requiresjq.(#283)
dispatch-taskwithdispatch --task-index.fill-transcriptsand fallback ingestion from final-message or capture-prefix files areretired; completion comes from runner-owned conversations and native transcripts. Custom
harnesses must declare valid
[dispatch],[transcript], and[tools]sections; removeparallel_command_templateandjudge_command_template. Seeeval-magic docs byoh.run --plan-modeoption with"plan_mode": trueinevals.json. Unsupported harnesses fail preflight; Codex and Cline donot support this mode. See
eval-magic docs conversations.skills remain visible by default. Set
codebase.exclude_skill_sources: trueto exclude theselected harness's project skill directories from both arms while retaining other configuration.
Matching treatment skills are reported as contamination risks.
(#287)
Fixes
gradereads live assertions andskill_should_trigger, while preserving the captured task definition and recording theassertion source. (#300)
held-out setup files. Cached checks require matching assertion and run-record digests.
(#312)
different names; descriptor roles provide portable aliases. Codex skill-access checks require
a successful read of the exact staged
SKILL.mdpath.(#314, #315)
--harness-file, and descriptor drift produces a warning.(#299)
files no longer inflate measured diffs. Guard teardown reaches the selected iteration's task
environments. (#302,
#304)
touch,mkdir,rm,cp,mv, andinstall, including destinations outside the task environment.(#303, #320)
/varand/private/varspellings. The isolation guide also explains nested Codex dispatch failures andtheir remedies. (#318,
#321)