From 081310b6cd17f2cb3eecffd8ba95de1ab09108a5 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Przemys=C5=82aw=20Galarowicz?= Date: Tue, 29 Sep 2026 10:15:53 +0200 Subject: [PATCH] docs(measurements): post-optimization performance audit of the delivery pipeline (apparatus) An analysis-only audit of /pharn-ship and /pharn-loop at 6.35.0, after the four optimization increments (6.32.0-6.35.0). No product-surface byte changes, so no SKILLS_VERSION bump. - .dev/measurements/pipeline-performance-audit-2026-09-29.md: the report. 0 of 69 real cost ledgers are post-optimization (all 6.12.1 / 6.7.0), so measured model usage and wall-clock cannot be ranked yet. From static and controlled evidence: at least 70 pinned orchestrator calls in a green one-iteration loop (57 deterministic), a 68k-130k-token first-request cache write per fresh stage agent, and the project's test gate at least twice per iteration. Recommends a prerequisite (update pharn-starter with CLI >= 0.7.0 and collect real runs) and two contingent increments (C1, C2). - .dev/features/pipeline-performance-audit/audit.mjs: the read-only helper behind the figures (--ledgers, --static, --prefix, --self-test). - The /pharn-dev-ship stage artifacts (PLAN, GRILL, REGRESSION, VERIFY, REVIEW, SHIP) and one [Unreleased] CHANGELOG entry. Co-Authored-By: Claude Opus 5.5 --- .../pipeline-performance-audit/GRILL.md | 225 +++++++ .../pipeline-performance-audit/PLAN.md | 238 +++++++ .../pipeline-performance-audit/REGRESSION.md | 43 ++ .../pipeline-performance-audit/REVIEW.md | 319 +++++++++ .../pipeline-performance-audit/SHIP.md | 87 +++ .../pipeline-performance-audit/VERIFY.md | 28 + .../pipeline-performance-audit/audit.mjs | 546 ++++++++++++++++ .../regression-report.json | 33 + .../verify-report.json | 15 + .../pipeline-performance-audit-2026-09-29.md | 610 ++++++++++++++++++ CHANGELOG.md | 23 + 11 files changed, 2167 insertions(+) create mode 100644 .dev/features/pipeline-performance-audit/GRILL.md create mode 100644 .dev/features/pipeline-performance-audit/PLAN.md create mode 100644 .dev/features/pipeline-performance-audit/REGRESSION.md create mode 100644 .dev/features/pipeline-performance-audit/REVIEW.md create mode 100644 .dev/features/pipeline-performance-audit/SHIP.md create mode 100644 .dev/features/pipeline-performance-audit/VERIFY.md create mode 100644 .dev/features/pipeline-performance-audit/audit.mjs create mode 100644 .dev/features/pipeline-performance-audit/regression-report.json create mode 100644 .dev/features/pipeline-performance-audit/verify-report.json create mode 100644 .dev/measurements/pipeline-performance-audit-2026-09-29.md diff --git a/.dev/features/pipeline-performance-audit/GRILL.md b/.dev/features/pipeline-performance-audit/GRILL.md new file mode 100644 index 00000000..3c803a14 --- /dev/null +++ b/.dev/features/pipeline-performance-audit/GRILL.md @@ -0,0 +1,225 @@ +# GRILL — pipeline-performance-audit + +Plan: `.dev/features/pipeline-performance-audit/PLAN.md`. Spec-hash check: `node .dev/floor/hash-doc.mjs pharn/ARCHITECTURE.md` +→ `d831d30d399a37dc403080072763d13383de6f6f31875e7e8cb4eadeb642f4f4`, equal to the plan's `spec_content_hash` (no +drift). **Step 1b (FLOOR):** `node pharn/floor/check-plan-lessons.mjs .dev/features/pipeline-performance-audit/PLAN.md +.dev/memory-bank/lessons-learned.md` → exit **0**, `GREEN — applied_lessons: L4, L24, L38, L40, L43, L63 … all 6 cited +id(s) resolve in .dev/memory-bank/lessons-learned.md and are referenced in the plan body.` That verdict covers the +DECLARATION only; whether each lesson was genuinely applied is advisory and is interrogated below (see G9). + +**How this grill was run (recorded, advisory).** One read-only pass by the grill stage itself. The plan is +`trust: untrusted`; every `problem` / `evidence` value below inherits that tag and is quoted DATA (P2). No +instruction-shaped content addressed to the grill was found in the plan. Discovery claims were re-read live where they +bear on a finding: the four part-file byte counts, `ROUTE_POLICY` in `pharn/floor/stage-agent-core.mjs`, this repo's +`pharn.config.json`, `stage-agent.mjs`'s `brief` branch (it writes stdout only), `pharn/pharn-contracts/cost-ledger.md`, +the two controlled `MEASUREMENT.md` files, the 6.32.0 CHANGELOG entry, and the pharn-starter ledgers' enum fields +(69 ledgers: 56 `6.12.1` / `run-window/1`, 13 `6.7.0` / schema `/1` with no `membership`; all `coverage: partial`, +none with `executions` or `work` — the plan's counts reproduce). + +**Registered grillers** (`node pharn/floor/count-grillers.mjs .` → `registered: 13`): a11y, architecture, +comprehension, coupling, documentation, error-handling, i18n, migrations, observability, performance, privacy, +security, testability. Scanners: `scan-plan-secrets` `{"found":false}`, `scan-plan-pii` `{"found":false}`, +`scan-plan-i18n` `{"found":false}`, `scan-plan-observability` `mentions:true` (lines 83, 86, 87 — the word "metrics" +meaning model-usage metrics; incidental), `scan-plan-migrations` `mentions:true` (line 237 — `pharn update` migrating a +config block in an option not chosen; incidental). + +## Findings — built-in axes + +### Determinism and the frozen run-selection rule (P5) + +```yaml +# G1 +- type: FINDING + rule_id: "P5" + severity: important + file: ".dev/features/pipeline-performance-audit/PLAN.md:214" + problem: "The claim that eligibility and every exclusion reason are membership tests over enum or version fields is false for two of the rule's own tests: `fixture-not-workload` has no ledger field to test (a cost.json does not record whether its project is a fixture, so the decision is a classification of a directory — a human's, at GATE 1, or a guess), and the version test reads `skills_version`, which cost-ledger.md labels ADVISORY ('records the configured version, never the bytes of the emitter that ran')." + evidence: "PLAN.md:214 'Eligibility and every exclusion reason are membership tests over enum or version fields.' PLAN.md:85 'the run happened in a real project, not a fixture.' PLAN.md:83 '`skills_version` ≥ 6.35.0 for the full profile.'" +# G2 +- type: FINDING + rule_id: "P5" + severity: important + file: ".dev/features/pipeline-performance-audit/PLAN.md:100" + problem: "The 'closed' exclusion set is not closed over the rule's own conditions, so the rule as frozen cannot decide every ledger: the 13 schema-/1 ledgers have NO `membership.method` (they fail 'must be run-window/2' but are not `run-window/1`, and PLAN.md:112 gives them no membership reason); the marker-completeness and open-execution checks of job (b) and the per-metric exclusions of rule 3 (an `unmeasured` execution, a missing work row) map to no reason; `coverage: unknown` is not a member of the contract's coverage enum (`partial | unavailable`); and rule 4 names no sample size above 20 and strata keyed on `mode`, which is not a top-level ledger field." + evidence: "PLAN.md:88-89 '`membership.method` must be `run-window/2`. `coverage` must not be `unavailable` or `unknown`.' PLAN.md:97-98 'Above 20, the sample is stratified by (command, mode, outcome, iteration count), taking every stratum, round-robin' PLAN.md:122 'marker-completeness and open-execution checks the prompt lists'" +``` + +### Guarantee audit (P0) + +```yaml +# G3 +- type: FINDING + rule_id: "P0" + severity: important + file: ".dev/features/pipeline-performance-audit/PLAN.md:74" + problem: "The rule is headed 'frozen before any cost is inspected' and the build 'may not amend it', but the plan's own discovery already inspected a ledger's cost figures before the rule was written, and nothing on the floor pins the rule text between GATE 1 and the build (PLAN.md is not content-hashed by any stage), while the guarantee audit neither reduces the freeze to a floor op nor labels it advisory." + evidence: "PLAN.md:74 'frozen HERE, before any cost is inspected' PLAN.md:61-62 'One sampled ledger has 393 requests and 16,852 output tokens.' PLAN.md:76-77 'The build applies it verbatim and may not amend it after seeing cost figures.'" +# G4 +- type: FINDING + rule_id: "P0" + severity: important + file: ".dev/features/pipeline-performance-audit/PLAN.md:71" + problem: "Static evidence is described as the bytes 'that enter a stage's context', but the planned pass measures only prescribed command, part and brief bytes: it omits the fixed context every stage agent and the orchestrator carry (system prompt, tool schemas, the project's CLAUDE.md — 172,421 B in this repo against 43,164 B for the largest command, pharn-loop.md — memory, MCP instructions), it cannot see what a model actually Reads, and `ceil(bytes / 4)` yields neither a token CLASS (input vs cache-creation vs cache-read) nor the per-request multiplier the 6.32.0 entry says governs cost — so section 4 'Dominant remaining costs' cannot be RANKED from it, and the per-class rule at line 152 has no stated way to hold a static estimate; the divisor 4 is also given no source." + evidence: "PLAN.md:71-72 'the bytes of every command, part, stage-agent brief and prescribed read that enters a stage's context.' PLAN.md:134 'Token counts are derived as `ceil(bytes / 4)` and labelled estimates.' PLAN.md:152 'Token classes are always reported separately.'" +# G5 +- type: FINDING + rule_id: "P0" + severity: important + file: ".dev/features/pipeline-performance-audit/PLAN.md:150" + problem: "One label per number conflates two axes — source (real / controlled / static) and precision (estimate) — so a static byte-derived estimate and the 6.32.0 request-profile estimate share `estimate` and read as comparable; `controlled` carries no fixture identity or floor commit, although the two MEASUREMENT.md files used different fixtures (3 gates with an install vs 4 gates with `--no-install`) and pre-6.35.0 floors (L24's inherited-bound disease); and line 70 itself mislabels the 6.32.0 entry, whose request counts were MEASURED over 37 real orchestrator transcripts and only whose savings are estimates." + evidence: "PLAN.md:150 'Every number carries one label, from `real` / `controlled` / `static` / `estimate`, plus its denominator.' PLAN.md:70 'The 6.32.0 CHANGELOG entry: measured command bytes, plus request and token figures it labels estimates.'" +# G6 +- type: FINDING + rule_id: "P0" + severity: minor + file: ".dev/features/pipeline-performance-audit/PLAN.md:193" + problem: "`validate.mjs` GREEN is cited as part of the floor reduction for 'changed no PHARN runtime behaviour', but validate checks capability and contract shape over the product surface, not that its bytes are unchanged; the real backstops are the writes-scope hook (prevents Write-tool writes outside `## Files`) and reconcile (detects Bash writes), which the same bullet already names." + evidence: "PLAN.md:193-194 'floor, but only for the product surface in the dev chain: `validate.mjs` GREEN, plus the diff limited to the four `## Files` paths.'" +``` + +### Discovery-first (P6) + +```yaml +# G7 +- type: FINDING + rule_id: "P6" + severity: important + file: ".dev/features/pipeline-performance-audit/PLAN.md:47" + problem: "The inventory lists ONE real orchestrator transcript, while the 6.32.0 CHANGELOG entry reports request counts measured in 37 orchestrator transcripts on the maintainer's machine; the rule's universe admits only cost.json files, so transcript-only evidence (per-request usage, served model, request counts per run) is excluded by construction, with no closed reason and without the human having decided it — for the human: is that exclusion intended, and are those 37 transcripts all pre-optimization?" + evidence: "PLAN.md:47-48 'one real `/pharn-loop` transcript in `~/.claude/projects/…pharn-loop-run/`, dated 2026-09-21' PLAN.md:79 'Every `cost.json` under a `pharn/features/*/` directory that the evidence inventory names'" +# G8 +- type: FINDING + rule_id: "P6" + severity: important + file: ".dev/features/pipeline-performance-audit/PLAN.md:35" + problem: "The 'current execution profile' is derived from THIS repo's `models.stages`, but routing is config-conditional and the only project with real runs (pharn-starter) carries the pre-0.7.0 block (top-level `default`, `opus-4-8` / `sonnet-5`), which check-model-config REDs at 6.35.0 so every routed stage runs `inline:config-red` on the session model; the plan does not say which config section 2 describes or that the profile differs per install." + evidence: "PLAN.md:35-36 'The requested models come from `pharn.config.json` `models.stages`. In this repo that is spec, plan, grill, review, memory-promote and ac-test on opus; build, regress, verify, ship and loop on sonnet'" +# G9 +- type: FINDING + rule_id: "P6" + severity: minor + file: ".dev/features/pipeline-performance-audit/PLAN.md:181" + problem: "L38 is applied to run-window/1 contamination, but canon's L38 records concurrent sessions contending for `.pharn/writes-scope.json` (a false-cause scope RED); the ledger contamination is the 6.29.0 measurement, and the one L38-apt application (serial controlled runs) belongs to Q1 option (2), which was not chosen — the declaration is GREEN, the application is misattributed." + evidence: "PLAN.md:181-183 'Concurrent sessions contaminated `run-window/1` membership in the historical corpus. So `membership-run-window-1` is a closed exclusion reason. Any controlled delivery runs made for this audit (Q1 option 2) run one at a time'" +``` + +### Honest scope / no speculation (P7) + +```yaml +# G10 +- type: FINDING + rule_id: "P7" + severity: important + file: ".dev/features/pipeline-performance-audit/PLAN.md:123" + problem: "Under the resolved route (0 eligible runs) helper job (c), per-stage profile extraction over `executions` / `work` / markers, and job (b)'s marker-completeness and open-execution checks have no input to run on; they are justified only by a future re-run, which is the hypothetical P7 forbids and outside the prompt's 'only if necessary to inspect existing evidence' — and (c) re-derives views cost.json already stores and check-cost-ledger.mjs recomputes (`executions` via stage-executions-core.mjs), a second owner of one derivation." + evidence: "PLAN.md:123-124 '(c) extract per-stage profiles, keeping four dimensions separate' PLAN.md:220-222 'The rule above still decides eligibility, and it currently includes 0 real runs.' PLAN.md:232-233 'The helper re-runs the full analysis once real ledgers at 6.35.0 or later exist.' PLAN.md:15-16 'only if necessary to inspect existing evidence'" +``` + +## Findings — registered grillers + +### testability (P1) + +Presence recognized: the plan declares a reproducibility check (the two helper invocations print every cited number). +Layer 2 (adequacy): + +```yaml +# T1 +- type: FINDING + rule_id: "P1" + severity: important + file: ".dev/features/pipeline-performance-audit/PLAN.md:134" + problem: "Reproducibility is declared but correctness is not: the helper gets no test, although the headline '0 of 69' rests on its version comparison (a string compare orders '6.7.0' after '6.35.0' and would admit the 13 6.7.0 ledgers), the static numbers change with every later commit unless the report records the commit they were measured at, and nothing compares the report's figures with the helper's output." + evidence: "PLAN.md:134-136 'The helper must be reproducible: `node .dev/features/pipeline-performance-audit/audit.mjs --static` and `… --ledgers ` print every number the report cites.' PLAN.md:197-198 'nothing on the floor checks the report against the helper's output.'" +``` + +### error-handling (P7) + +```yaml +# E1 +- type: FINDING + rule_id: "P7" + severity: minor + file: ".dev/features/pipeline-performance-audit/PLAN.md:121" + problem: "The helper shells check-cost-ledger.mjs but the plan does not map its exits to reasons: exit 1 is both a RED and node's own crash code (the class pharn/floor/shelled-verdict-core.mjs owns — a RED is exit 1 WITH its `RED — ` line), and exit 2 (unusable input) has no stated reason, so a crashed checker could exclude a run as `ledger-red`." + evidence: "PLAN.md:121 '(b) validate each ledger by shelling `pharn/floor/check-cost-ledger.mjs` as a CLI'" +``` + +### security (P2) + +Scanner clean. Layer 2: + +```yaml +# S1 +- type: FINDING + rule_id: "P2" + severity: minor + file: ".dev/features/pipeline-performance-audit/PLAN.md:121" + problem: "The ledger paths the helper passes to a child process come from another project's directory names and a `--ledgers` argv list, and the plan does not state an argv-array spawn with no shell — the shell-sink class this repo recorded in 6.28.0 and 6.30.0." + evidence: "PLAN.md:121 'shelling `pharn/floor/check-cost-ledger.mjs` as a CLI' PLAN.md:136 '`… --ledgers `'" +``` + +### privacy (P2) + +Scanner clean. Layer 2: + +```yaml +# PR1 +- type: FINDING + rule_id: "P2" + severity: minor + file: ".dev/features/pipeline-performance-audit/PLAN.md:136" + problem: "The report is committed to a public repo and cites helper output over files under the home directory, yet the plan states no path-redaction rule, while the repo's own ledger contract forbids absolute paths in a committed artifact; a printed ledger path or transcript directory key would carry the user's home path." + evidence: "PLAN.md:42 '`~/Projects/pharn-starter/pharn/features/*/cost.json`' PLAN.md:47 '`~/.claude/projects/…pharn-loop-run/`' PLAN.md:136 'print every number the report cites'" +``` + +### comprehension (P7) + +```yaml +# C1 +- type: FINDING + rule_id: "P7" + severity: minor + file: ".dev/features/pipeline-performance-audit/PLAN.md:24" + problem: "The four byte counts are listed in an order that does not follow the brace expansion they annotate (live: loop-quick 11,091, loop-close 31,211, ship-quick 11,223, ship-close 28,179), so a reader or the report can attribute them to the wrong files." + evidence: "PLAN.md:24 '`.claude/commands/pharn-{loop,ship}-{quick,close}.md` exist (11,091 / 11,223 / 31,211 / 28,179 B)'" +``` + +### No finding (reason recorded) + +- **architecture** — fit recognized: helper scripts under `.dev/features//` have precedent + (`regress-base-reuse/measure.mjs`, `verify-head-gate-reuse/measure.mjs`, `run-performance-breakdown/demo.mjs`), and + measurement reports live in `.dev/measurements/`. +- **coupling** — clean seam: the checker is shelled as a CLI, not imported; the duplicated derivation is G10. +- **performance** — no scaling risk: 69 JSON files and a fixed set of command files. +- **documentation** — the helper is apparatus with its invocations stated; no public surface is added. +- **observability** — scanner mentions are incidental ("metrics" = model-usage metrics); an offline analysis script + needs no runtime observability. +- **migrations** — the one mention is `pharn update`'s config migration in an unchosen option; no persisted schema is + touched. +- **a11y**, **i18n** — no UI and no user-facing strings. + +## Summary + +The plan is honest about its central limit (0 of 69 real runs are post-optimization) and keeps the analysis-only +scope, but the frozen rule is weaker than its framing. It is neither closed nor fully decidable: two tests are not +membership tests over enum fields (G1), and several conditions map to no reason, with the sample size above 20 +undefined (G2). Its freeze is asserted with no floor backing, after cost figures had already been read (G3). Under +the chosen route the report rests on static and controlled evidence. The static pass as scoped omits the largest fixed +context component and cannot yield a token class or a per-run multiplier, so it cannot rank "dominant remaining +costs" (G4). The four-label scheme cannot keep estimates of different provenance, or controlled figures from +different fixtures and floors, apart (G5). Two discovery gaps bear on the evidence set and the execution profile: +transcript evidence (G7) and config-conditional routing (G8). The helper is partly built for inputs that do not exist +yet (G10), and its correctness is unverified where the headline number depends on it (T1). The minor findings (G6, G9, +E1, S1, PR1, C1) are corrections to citations, exit handling, spawn form, path hygiene and one byte list. + +For the human (P5 — asked, not guessed): who decides `fixture-not-workload`, and on what record (G1); whether +transcript-only evidence is meant to be excluded (G7); which project's config section 2 describes (G8); and whether +job (c) should be deferred rather than built (G10). + +Finding ids, in order of appearance (each is the `#` comment above its object): G1 (P5 :214), G2 (P5 :100), G3 (P0 +:74), G4 (P0 :71), G5 (P0 :150), G6 (P0 :193), G7 (P6 :47), G8 (P6 :35), G9 (P6 :181), G10 (P7 :123), T1 (P1 :134), +E1 (P7 :121), S1 (P2 :121), PR1 (P2 :136), C1 (P7 :24). + +ADVISORY VERDICT: 15 concerns raised (0 blocking-severity, 15 advisory — 9 important, 6 minor) — for the human to +weigh before /pharn-dev-build. This verdict covers the interrogation only; the Step 1b lessons-declaration verdict +(GREEN, exit 0) is a separate floor result and is not counted here. diff --git a/.dev/features/pipeline-performance-audit/PLAN.md b/.dev/features/pipeline-performance-audit/PLAN.md new file mode 100644 index 00000000..5ce30f81 --- /dev/null +++ b/.dev/features/pipeline-performance-audit/PLAN.md @@ -0,0 +1,238 @@ +# PLAN — pipeline-performance-audit: a measurement-driven audit of the post-optimization delivery pipeline + +- spec_content_hash: d831d30d399a37dc403080072763d13383de6f6f31875e7e8cb4eadeb642f4f4 +- applied_lessons: [L4, L24, L38, L40, L43, L63] +- increment: an ANALYSIS-ONLY audit of the product delivery pipeline (`/pharn-ship`, `/pharn-loop`) as it stands at + 6.35.0. It produces one report and one read-only analysis helper, and changes no PHARN runtime behaviour. +- layer(s): build apparatus only (`.dev/features/`, `.dev/measurements/`) plus a `CHANGELOG.md` `[Unreleased]` entry. + No product-surface byte changes, so no `SKILLS_VERSION` bump. +- constitution_refs: [P0, P2, P5, P6, P7] + +## Why (P7) + +**The trigger is the maintainer's explicit direction**: this run's prompt asks for a "Post-Optimization Performance +Audit". It is recorded that way, not as a failure this plan re-derives. The prompt forbids implementing any +optimization. This increment therefore adds no capability, rule or enforcer. Its only code is an analysis helper, +which the prompt allows "only if necessary to inspect existing evidence". + +## Discovery — live state read this run (P6) + +### The four preceding optimization areas are on `main` + +| area | version / PR | live evidence (files at `c9737b4`) | +| ------------------------------------- | --------------- | ------------------------------------------------------------------------------------------------- | +| 1. orchestrator-context reduction | 6.32.0, PR #294 | `.claude/commands/pharn-{loop,ship}-{quick,close}.md` exist (11,091 / 11,223 / 31,211 / 28,179 B) | +| 2. run-scoped BASE regression reuse | 6.33.0, PR #297 | `pharn/floor/regress-base-reuse{,-core}.mjs` | +| 3. VERIFY reuse of REGRESS/HEAD gates | 6.34.0, PR #298 | `pharn/floor/gate-reuse-core.mjs`, `head-reuse-offer.mjs` | +| 4. per-run performance/cost breakdown | 6.35.0, PR #299 | `pharn/floor/stage-executions-core.mjs`, `stage-work.mjs`; `SKILLS_VERSION` = `6.35.0` | + +### The routing table as it stands (`pharn/floor/stage-agent-core.mjs` `ROUTE_POLICY`) + +- `/pharn-ship` full: spec `interactive` (it is GATE 1); plan, grill, test, build `agent`; regress, verify + `floor-only` (inline thin callers over `stage-regress.mjs` / `stage-verify.mjs`). +- `/pharn-loop` full: spec, plan, grill, test, build `agent`; regress, verify `floor-only`. +- quick columns: grill `floor-only`, regress `skipped`, the rest as in full. +- The requested models come from `pharn.config.json` `models.stages`. In this repo that is spec, plan, grill, review, + memory-promote and ac-test on opus; build, regress, verify, ship and loop on sonnet; effort `high` everywhere. Effort + is not routed to a stage agent: the Agent tool takes no effort parameter (CLAUDE.md, STAGE-MODEL ROUTING). + +### The evidence inventory, with the result of the selection rule below + +- **Real product-pipeline ledgers found: 69.** All of them are in + `~/Projects/pharn-starter/pharn/features/*/cost.json`: 56 with `skills_version` 6.12.1 and 13 with 6.7.0. The + newest was written 2026-09-28 10:17. pharn-starter's `pharn.config.json` still records `skillsVersion` 6.12.1 + (installed 2026-09-23). +- **Other real run records, none of them a ledger:** + - this repo's `pharn/features/loop-decision-integrity/` (a `LOOP.md` and its reports, from the 6.3.0 era); + - one real `/pharn-loop` transcript in `~/.claude/projects/…pharn-loop-run/`, dated 2026-09-21, before the + optimizations. + + The build re-counts every one of these with the helper and does not trust the numbers above. + +- **Post-optimization real runs (installed version ≥ 6.32.0): 0.** Every ledger is 20 or more minor versions older + than the first optimization. + The 6.12.1 pipeline has no `/pharn-test` stage, no stage-agent routing (6.27.0), and no regress/verify stage + scripts (6.23.0 / 6.26.0). It is not the pipeline this audit is about (L24). +- **Known measurement defects of those ledgers.** These are independent of their age and are found by reading them: + - `membership.method` is `run-window/1` on all 56 6.12.1 ledgers, and absent on the 13 6.7.0 ledgers. That + membership is not context-scoped, and 6.29.0 measured it counting concurrent agents' rows (L38). Several ledgers + share one `window_start`: 10 share 2026-09-22T17:06:51Z, 2 share 2026-09-23T08:56:20Z and 2 share + 2026-09-25T20:39:38Z. + - Output was derived from each request's first transcript line before 6.24.1 (L63). One sampled ledger has 393 + requests and 16,852 output tokens. + - There are no `executions` or `work` keys, because those arrived in 6.35.0. + - `coverage` is `partial` on all of them. +- **Controlled evidence (fixtures, never user workload, L4):** + - `.dev/features/regress-base-reuse/MEASUREMENT.md`: process counts and wall-clock on a fixture, 3 repetitions. + - `.dev/features/verify-head-gate-reuse/MEASUREMENT.md`: the same kind of measurement for verify reuse. + - `.dev/features/run-performance-breakdown/DEMO.md`: its token counts are **synthetic**. It may be cited for helper + overhead only, never for model usage. + - The 6.32.0 CHANGELOG entry: measured command bytes, plus request and token figures it labels estimates. +- **Static evidence (the current tree, measurable now):** the bytes of every command, part, stage-agent brief and + prescribed read that enters a stage's context. The token figures derived from these bytes are estimates. + +## The run-selection rule — frozen HERE, before any cost is inspected + +Step 2 of the prompt requires fixing this rule before any stage's cost is looked at. The build applies it verbatim and +may not amend it after seeing cost figures. An amendment goes back to the human (P6). + +1. **Universe.** Every `cost.json` under a `pharn/features/*/` directory that the evidence inventory names, plus any + added at GATE 1. +2. **Eligible as post-optimization real evidence** only if all of these hold: + - `command` ∈ {`/pharn-ship`, `/pharn-loop`}; + - `skills_version` ≥ 6.35.0 for the full profile. A ledger from 6.32.0–6.34.x is eligible for model-usage metrics + only, because it has no `executions` or `work`; + - the run happened in a real project, not a fixture. +3. **Per-metric validity.** A run is used for a metric only if its evidence for that metric passes. Each exclusion is + reported per metric, never silently dropped: + - **Run token totals.** `check-cost-ledger.mjs` must be GREEN. `membership.method` must be `run-window/2`. `coverage` + must not be `unavailable` or `unknown`. A `partial` coverage is used but reported as partial, never as complete. + - **Per-stage requests and tokens.** The same conditions, plus the stage attribution taken from `markers[]`. + `unattributed` is reported beside the stage figures, never folded into them. + - **Elapsed time.** Only `executions` rows with a non-null `elapsed_ms`. `unmeasured` rows are counted by their + reason and never read as zero. + - **Deterministic work.** Only `work[]` rows. A stage with no row is `not recorded`, never zero. + - **Model.** The requested model comes from each marker's `route`. The served model comes from `requests[].model`. + The two are always reported separately. +4. **Sampling.** Every eligible run is used when there are 20 or fewer. Above 20, the sample is stratified by + `(command, mode, outcome, iteration count)`, taking every stratum, round-robin, in `window_start` order. The strata + are structural and are fixed here. +5. **Closed exclusion reasons:** + - `pre-optimization-version`; + - `wrong-command`; + - `fixture-not-workload`; + - `ledger-red`; + - `membership-run-window-1`; + - `coverage-unavailable`; + - `unreadable`. + + A ledger may carry several reasons. The report states all of them. + +**Applying the rule to the live inventory gives 0 included real runs and 69 excluded**, all as +`pre-optimization-version`, and 56 of them also as `membership-run-window-1`. Open question Q1 decides what the audit does about +that. + +## What the build does + +The work depends on the GATE 1 answer to Q1. Every route writes the same files. Only the evidence set differs. + +1. **`audit.mjs` (read-only helper).** It writes to stdout only and changes no runtime state. It has four jobs: + - (a) apply the frozen rule to a list of ledgers; + - (b) validate each ledger by shelling `pharn/floor/check-cost-ledger.mjs` as a CLI, plus the schema, membership, + coverage, marker-completeness and open-execution checks the prompt lists; + - (c) extract per-stage profiles, keeping four dimensions separate: model usage per token class, observed elapsed + time, deterministic work, and retry/rerun counts; + - (d) run a `--static` pass over the current tree. + + The static pass measures: + - the bytes each routed stage agent receives from `stage-agent.mjs brief`, which only writes stdout (read at + `stage-agent.mjs:527`); + - the bytes of the stage command file the brief points at; + - the bytes of the orchestrator command at invocation and of each part at its load point; + - the bytes of each inline stage's thin caller. + + Token counts are derived as `ceil(bytes / 4)` and labelled estimates. The helper must be reproducible: + `node .dev/features/pipeline-performance-audit/audit.mjs --static` and + `… --ledgers ` print every number the report cites. + +2. **The report,** at `.dev/measurements/pipeline-performance-audit-2026-09-29.md`. Measurement reports live in + `.dev/measurements/`, as in `token-cost-2026-08-18.md`. It has the nine sections the prompt names: + 1. Evidence set + 2. Current execution profile + 3. Measured cost profile + 4. Dominant remaining costs + 5. Root-cause analysis + 6. Candidate optimizations + 7. Recommended next 2–4 increments + 8. Do not optimize yet + 9. Measurement gaps + + **Every number carries one label**, from `real` / `controlled` / `static` / `estimate`, plus its denominator. + Figures with different labels or coverage semantics are never compared as percentages of each other. + Token classes are always reported separately. There is no weighted score and no dollar figure. + +3. **Conclusion discipline (P0/P7).** + - A candidate optimization must cite a measured finding earlier in the report. With no real runs, a candidate may + rest on a `static` or `controlled` measurement. It is then labelled "contingent on real-run confirmation". + - A candidate that changes agent independence, model, effort, stage boundaries, what BUILD sees, or the + validation/retry strategy gets a defined counterfactual experiment, never an adoption recommendation. + - If the evidence supports fewer than two implementation increments, the report says so rather than inventing + any. The prompt's "2–4" is an upper frame, not a quota to fill. + - A root cause is stated only when the evidence tells it apart from the alternatives. Otherwise the report names + the rival causes and what would separate them (L40). +4. **`CHANGELOG.md`:** one `[Unreleased]` entry, dated `- 2026-09-29:`. This is apparatus-only, so there is no bump. + +## Files + +- `.dev/features/pipeline-performance-audit/PLAN.md` — this plan — apparatus +- `.dev/features/pipeline-performance-audit/audit.mjs` — NEW. A read-only analysis helper: rule application, ledger + validation, per-stage profile extraction and the static byte pass. Node stdlib only; it writes stdout only. — + apparatus +- `.dev/measurements/pipeline-performance-audit-2026-09-29.md` — NEW. The audit report. — apparatus +- `CHANGELOG.md` — one `[Unreleased]` entry — repo-meta + +## Applied lessons + +- **L4.** Controlled and fixture evidence is kept apart from real workload in every table. A fixture's + reuse-count result shows that the mechanism fires. It never shows how often a real run hits it. +- **L24.** No figure from a 6.7.0 or 6.12.1 ledger, or from a pre-optimization transcript, is carried into a + post-optimization claim. The pipeline those figures describe was replaced, so they are listed as exclusions + only. +- **L38.** Concurrent sessions contaminated `run-window/1` membership in the historical corpus. So + `membership-run-window-1` is a closed exclusion reason. Any controlled delivery runs made for this audit (Q1 + option 2) run one at a time, never in parallel. +- **L40.** Root causes are stated only where the evidence separates them from rival causes. Otherwise the report + names the rivals and the observation that would discriminate between them. +- **L43.** A GREEN `check-cost-ledger.mjs` is treated as internal consistency only. The helper checks membership and + coverage separately, and the report never reads GREEN as "the ledger is accurate". +- **L63.** A pre-6.24.1 ledger's output is derived from a still-growing line, so it under-counts. That is one more + reason those ledgers cannot support an output-token claim. + +## Guarantee audit (P0) + +- "The audit changed no PHARN runtime behaviour" → floor, but only for the product surface in the dev chain: + `validate.mjs` GREEN, plus the diff limited to the four `## Files` paths. `enforce-writes-scope.cjs` pins + Write-tool writes to them. Bash writes are detected by `/pharn-dev-verify`'s `reconcile` gate and are not prevented + (L19). +- "Run X was excluded for reason R" → advisory. The helper applies an enum rule, but nothing on the floor checks the + report against the helper's output. +- "Stage S accounted for N of M requests" → advisory arithmetic over ledger fields. Its inputs are only as honest as + the ledgers, and a GREEN ledger is agreement, not provenance (L43). +- "Byte and token figures of the static pass" → bytes are measured. Every token figure is an estimate and is + labelled as one. +- Every candidate's expected benefit and quality risk → advisory, always. + +## Trust audit (P2) + +The helper reads numeric and enum fields of `cost.json` ledgers and the byte sizes of files. It quotes no free text +from pharn-starter's artifacts. Where the report names a pharn-starter feature slug, the slug is a directory name, +shown as data. No ledger field reaches a proceed/stop decision of this chain. The dev chain's verdicts come from its +own floor gates only. + +## Determinism audit (P5) + +Eligibility and every exclusion reason are membership tests over enum or version fields. Anything the rule cannot +decide goes to the human. + +## Resolved at GATE 1 (2026-09-29, the maintainer) + +- **Plan: approved as written.** +- **Q1: option (1), "Audit now".** The build uses static and controlled evidence. It prepares no fixture delivery + runs and does not wait for real runs. The rule above still decides eligibility, and it currently includes 0 real + runs. +- **Dates.** The work is authored on 2026-09-29, so the report is + `.dev/measurements/pipeline-performance-audit-2026-09-29.md` and the CHANGELOG entry is dated `- 2026-09-29:`. + The path and date above were updated in place from the 2026-09-28 draft. + +For the record, the three options that were offered: + +- **Q1 — the evidence route.** No post-optimization real run exists. The options: + - **(1) Audit now from static and controlled evidence.** The report states "0 of 69 real runs are + post-optimization" in each model-usage and elapsed-time section, and restricts its recommendations to what + static and controlled measurement supports. The helper re-runs the full analysis once real ledgers at 6.35.0 or + later exist. + - **(2) Make controlled delivery runs first.** The helper prepares a fixture: a copy of pharn-starter at a pinned + commit with the 6.35.0 product surface. It also prepares 2–3 pre-chosen small tasks. You start each `/pharn-loop` + in its own session, one at a time (L38). + - **(3) Pause for real runs.** You run `pharn update` in pharn-starter with a CLI at 0.7.0 or later, which migrates + the `models` block. Your normal work then produces the ledgers, and the audit resumes on them. diff --git a/.dev/features/pipeline-performance-audit/REGRESSION.md b/.dev/features/pipeline-performance-audit/REGRESSION.md new file mode 100644 index 00000000..309368db --- /dev/null +++ b/.dev/features/pipeline-performance-audit/REGRESSION.md @@ -0,0 +1,43 @@ +# REGRESSION — pipeline-performance-audit + +This is the second run, after the GATE-2 fix build (the maintainer chose "Fix, then PR"). The first run was also +`no-regressions`. + +- **Base:** `c9737b4486bc47759bd36f43f3430cf86bc068b4` (`HEAD`; a working-tree build, so `git status --porcelain` was + non-empty and the base auto-resolved to `HEAD`). +- **Inside (changed since base, plus untracked):** 11 paths. + - The four files the plan declares: `PLAN.md` and `audit.mjs` in `.dev/features/pipeline-performance-audit/`, + `.dev/measurements/pipeline-performance-audit-2026-09-29.md`, and `CHANGELOG.md`. + - This feature's stage artifacts, which `scope` exempts (7 in `escape_exempt`). + + `scope` found no escape (`escaped: []`). + +- **Outside gates:** + - `tests`: all 130 tracked `*.test.mjs` / `*.test.cjs` files, the same set as the first run, run with the pinned + `xargs node --test` form; + - `validate`; + - `structural:pharn/pharn-review/trust-fence/evals/expected/expected-injection-comment.json`. + + The style gates were skipped: no shared style config is inside. + +- **Baseline: REUSED from the first run, not re-executed.** + - The base map comes from this feature's first regress run. That run used a detached worktree at the same base SHA, + removed after it. + - The base commit, the outside gate set and the install decision (none) are all unchanged, so the base side's input + is identical. This is the BASE-reuse rule 6.33.0 applies in the product pipeline, applied here by hand and stated. + - Only the HEAD side re-ran: 130 test files, `validate` and the structural pair, on the fixed tree. + +| gate | base (reused) | head | +| ------------------------------------------------------------------------------------------ | ------------: | ---: | +| `tests` | 0 | 0 | +| `validate` | 0 | 0 | +| `structural:pharn/pharn-review/trust-fence/evals/expected/expected-injection-comment.json` | 0 | 0 | + +- `regressions[]`: none +- `pre_existing[]`: none + +**REGRESSIONS: none — no deterministically-detectable breakage outside the feature** (`check-regress.mjs verdict`, +exit 0, `"verdict": "no-regressions"`). + +This catches exactly what the suite catches, nothing more: a regression that no test, eval or rule covers is invisible +here. diff --git a/.dev/features/pipeline-performance-audit/REVIEW.md b/.dev/features/pipeline-performance-audit/REVIEW.md new file mode 100644 index 00000000..8d215d2a --- /dev/null +++ b/.dev/features/pipeline-performance-audit/REVIEW.md @@ -0,0 +1,319 @@ +# REVIEW — pipeline-performance-audit + +Reviewed by `/pharn-dev-review` on 2026-09-29. The increment under review is `trust: untrusted`. Every `problem` / +`evidence` value below inherits that tag and is quoted DATA (P2). + +**Scope.** An analysis-only increment: + +- `PLAN.md` and `GRILL.md`; +- `audit.mjs`, a read-only helper; +- `.dev/measurements/pipeline-performance-audit-2026-09-29.md`, the report; +- one `CHANGELOG.md` `[Unreleased]` entry; +- the stage artifacts (`REGRESSION.md`, `regression-report.json` `no-regressions`, `VERIFY.md`, `verify-report.json` + `PASS`). + +No product-surface byte changed (`git status`: only `CHANGELOG.md` modified, plus the two untracked apparatus paths), so +there is correctly no `SKILLS_VERSION` bump. + +## Step 1 — Floor first (P0) + +- `node pharn/floor/validate.mjs .` → **exit 0**, `FLOOR: GREEN — 36 capabilities checked`. +- `npm run check:changelog` → GREEN (`1 dated [Unreleased] entr(ies)`). +- `node .dev/floor/check-changelog-entry.mjs --base-ref HEAD` → GREEN (1 new entry; no merged entry or released heading + changed). +- `node .dev/features/pipeline-performance-audit/audit.mjs --self-test` → `{"ok": true, "checks": 14, "failed": []}`, + exit 0. + +That is the only guaranteed part of this review. **No floor gate reads the report's figures** (VERIFY.md says so), so +everything below is **advisory**. + +## Reproduction of the report's figures (advisory, re-run live) + +| report figure | re-run | result | +| ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------- | +| §1 ledger table (69 / 0 / 0 / 0·0·0, 56 + 13, `partial` ×69, GREEN ×69, 49 multi-context, 3 shared window starts, 2 exclusion reasons) | `audit.mjs --ledgers ~/Projects/pharn-starter/pharn/features/*/cost.json` | **matches** exactly (the 49 are all 6.12.1) | +| byte figures (43,164 / 35,982 / 11,091 / 11,223 / 31,211 / 28,179; 7,958; 19,848 / 17,929; 172,421 / 387,549; briefs 2,498–4,063) | `audit.mjs --static --project ~/Projects/pharn-starter` at `c9737b4` | **matches** | +| route `inline:config-red` on every routed cell for pharn-starter | same run | **matches** (every `project_route` exit 3) | +| pinned blocks: Step 1a 5, Step 3 5, Step 4 14, Step 5 10; close 6b 8, 6c 5, Step 7 1, Final 2 | same run | **matches** as section counts (but see F2 on what they count) | +| 29,581 B (2,498 + 7,958 + 19,125, ship test) … 37,485 B (3,266 + 7,958 + 26,261, quick-loop spec) | arithmetic | **follows** | +| 7.4k–9.4k tokens; 7–9% of 102,546 | 29,581 / 4 = 7,395; 37,485 / 4 = 9,371; 7.2% and 9.1% | **follows** arithmetically (see F5 on whether it may be stated) | +| 66 = 5 + 5 + 14 + 10 + 16 + 5 + (2 × 4) + 3; 19 = 10 + 1 + 8 | arithmetic | the sums **follow**; the inputs do not all hold (F2) | +| "about 60 of the 66" deterministic | 66 − 5 Agent calls − 3 model file operations − 2 slash-command invocations = 56 | **does not follow** (F2) | +| 37,777 B = 19,848 + 17,929; ~9.4k tokens | 37,777 / 4 = 9,444 | **follows** | +| ~0.5M = 5 × 102,546; ~1M = 10 × 10⁵ | 512,730; 1,000,000 | **follow** arithmetically (see F3 on the token class) | +| 2.25× = 387,549 / 172,421 | 2.248 | **follows** | +| CH1: 37 transcripts, min 34,936, median 102,546, max 129,648; 15 with ~31k cache read; opus since 09-27 122,475–129,648 | `find ~/.claude/projects/-Users-…-pharn-oss -path '*/subagents/agent-*.jsonl' -mtime -8 -exec audit.mjs --prefix {} +` → now 38 transcripts, median 102,781, max 130,176 (this review's own agent added one). Recomputed without that transcript: n 37, min 34,936, median 102,546, max 129,648 | **right when taken**; see F4 (population) and F9 (selection not stated) | +| CF1 10.1–10.6 s → 4.4 s; checkouts / installs / base gates 1 / 1 / 3 → 0; all gates 6 → 3 | `.dev/features/regress-base-reuse/MEASUREMENT.md` | **matches** | +| CF2 4.8–4.9 s → 2.5 s; verify 4 → 2; total 12 → 10; `test` 0.3 s | `.dev/features/verify-head-gate-reuse/MEASUREMENT.md` | **matches** | +| code claims: `ROUTE_POLICY` cells; `test` ∈ `NON_REUSABLE_IDS`; effort inherited; brief rule 7; `/pharn-build` Step 4; cap default 3; 570 s budget; `ship-closeout-script`, `gate-process-duration`, `stage-agent-effort` named; regress HEAD runs the outside-scope test subset; pharn-starter's `models` block shape | read live (`stage-agent-core.mjs`, `gate-run-core.mjs:238`, `pharn-loop.md`, `pharn-build.md`, `stage-regress.mjs:504/542`, `CHANGELOG [6.32.0]`, `[6.35.0]`, pharn-starter `pharn.config.json`) | **all correct** | + +- **Home paths.** No committed file of the increment contains a home-directory path. A grep for `/Users/`, the + account name, `/home/` and `/private/tmp` finds nothing, and the helper's printed paths are `~`-redacted. +- **"Do not optimize yet" and "Recommended next increments".** Both follow from the evidence in their structure: every + candidate is gated on real-run confirmation or on E1. One exception, the first "Do not optimize yet" row, rests on + F5. + +## Lenses + +- **L-floor (P0).** The report claims no floor guarantee. Its "What this audit may claim" block labels the estimates + and root causes advisory. But the report states discipline rules of its own and does not keep them: + - "Token classes are always kept apart" (broken: F3); + - "Figures with different sources are never expressed as percentages of each other" (broken: F5); + - "Every number carries a source and a precision" (broken: F9). + + Several headline sentences also state more than static or controlled evidence can carry (F1, F4, F5, F7). All of + these are advisory, because no guaranteed decision rests on the report. + +- **L-eval (P1).** No `role:`-bearing capability was added, so no eval binding is owed. The floor agrees: validate is + GREEN over 36 capabilities, the same set as before. The helper's `--self-test` covers the version compare, + `classify` and the block counter. It does **not** cover `storedProfile`, which is where F6 lives. +- **L-trust (P2).** + - The helper prints no free text from pharn-starter or from the transcripts. It prints only enum and numeric fields, + the checker's exit, WARN count and context count, and slugs as directory names. + - The checker is spawned with an argv array and no shell. + - The report quotes no untrusted artifact text. + - Nothing in the reviewed files addressed this reviewer or changed its behaviour. +- **L-axis (P3).** `audit.mjs` imports `pharn/floor/shelled-verdict-core.mjs` and `stage-agent-core.mjs`. That is the + permitted `.dev/` → `pharn/` direction, and there is no sibling reference. The helper's four modes all serve this + one audit. + +## Floor-gate findings (blocking) + +None. + +## Advisory findings (warn — each rests on reviewer judgment of free text or severity) + +```yaml +# F1 +- type: FINDING + rule_id: "P0" + severity: important + file: ".dev/measurements/pipeline-performance-audit-2026-09-29.md:20" + problem: "The headline 'Most orchestrator requests are deterministic ceremony' (and §4.1's 'Orchestrator requests are mostly ceremony') is a claim about the share of a real run's orchestrator requests, but the evidence is a static count of PINNED calls whose denominator (all orchestrator requests, pinned or not) is unmeasured — the report itself says unpinned verdict reads, resumes, re-runs and extra reads are not counted, and the 6.32.0 entry measured pre-routing /pharn-loop runs at a median of about 200 requests against this count of 66; what the static pass supports is 'most PINNED orchestrator calls are deterministic'." + evidence: "report:20 '**Most orchestrator requests are deterministic ceremony.** In a one-iteration `/pharn-loop`, about 66 orchestrator tool calls are pinned or mandated (estimate).' report:247 'This is a lower bound: unpinned verdict reads, `continue` resumes, re-runs and any extra reads the model makes are not counted.' CHANGELOG [6.32.0]: 'a `/pharn-loop` run made 11–426 requests (median about 200)'" +# F2 +- type: FINDING + rule_id: "P0" + severity: minor + file: ".dev/measurements/pipeline-performance-audit-2026-09-29.md:244" + problem: "The 66-call estimate sums per-section pinned-block counts as if each were a green-path call, so it is neither the green-path count nor a lower bound: (a) Step 4's 14 includes `check-red-run.mjs --preflight`, which runs only when `check-test-stage.mjs` exits non-zero, so a green run executes 13; (b) it omits mandated calls — the model's Write of LOOP.md in Step 6b, and, because Step 5 pins `/pharn-regress --base ` and `/pharn-verify` with no name argument, each thin caller's Step 0 slug Write plus its `feature-name.mjs` line; (c) '(2 × 4)' is never reconciled with the sentence's '3 pinned blocks'; and (d) 'About 60 of the 66' does not follow — 66 − 5 Agent calls − 3 model file operations − 2 invocations = 56." + evidence: "report:241-245 '2 inline stage invocations each carrying 3 pinned blocks of their own (scope set, run, release). Add 3 model file operations … about 5 + 5 + 14 + 10 + 16 + 5 + (2 × 4) + 3 = **66**' report:247-248 'This is a lower bound … About 60 of the 66 run a deterministic script' pharn-loop.md:476-480 'non-zero → decide the row with the checker … node pharn/floor/check-red-run.mjs --preflight' pharn-regress.md Step 0 'A `` this command did not receive as its argument is resolved only through `pharn/floor/feature-name.mjs`: write the slug alone … with the Write tool'" +# F3 +- type: FINDING + rule_id: "P0" + severity: important + file: ".dev/measurements/pipeline-performance-audit-2026-09-29.md:178" + problem: "The report says token classes are always kept apart, yet its CH1 headline figure is a blend of three classes (uncached input + cache write + cache read), the ~0.5M per-run figure multiplies that blend, and C1's '~1M cache-read tokens per run' reuses a subagent's first-request figure, which is mostly cache WRITE, as the orchestrator's per-request cache READ; the class-separated numbers the helper already prints differ materially (cache write alone: median 87,766 over the same 37 transcripts, vs the blended 102,546)." + evidence: "report:48 'Token classes are always kept apart.' report:178-181 'The total of uncached input, cache write and cache read per first request was: min 34,936, median 102,546' report:254-255 'For 5 agents at CH1's median that is about 0.5M first-request tokens per run [e]' report:329-330 'on the order of 10⁵ tokens of cache read per request here, so ~1M cache-read tokens per run [e]'" +# F4 +- type: FINDING + rule_id: "P0" + severity: minor + file: ".dev/measurements/pipeline-performance-audit-2026-09-29.md:23" + problem: "CH1 is described as general-purpose (and, in the CHANGELOG, stage-agent) first requests, but the helper does not filter by agent type: 6 of the 37 are Explore agents (34,936–46,734 tokens — the whole low tail) and 2 are claude-code-guide agents on haiku, and none is a product stage agent; a PHARN stage agent is always spawned as `general-purpose`, and over the 29 general-purpose transcripts the range is 68,225–129,648 (median 103,607), so the '35k–130k' range in the headline, §4.2 and the CHANGELOG describes a different population from the one the claim is about." + evidence: "report:23-24 'A new general-purpose subagent's first request carried 102,546 tokens at the median (35k–130k, 37 transcripts).' CHANGELOG.md:35 'a fresh stage agent's first request carrying 35k–130k tokens of fixed prefix, median 102,546, over 37 subagent transcripts in this repo' subagent meta: agent-a8e7deaa2ae78c444 'Explore', agent-a9348ab2f315d84af 'claude-code-guide' (model claude-haiku-4-5)" +# F5 +- type: FINDING + rule_id: "P0" + severity: important + file: ".dev/measurements/pipeline-performance-audit-2026-09-29.md:190" + problem: "The '7–9%' that grounds the first 'Do not optimize yet' row expresses a static estimate as a percentage of a CH-sourced figure, which the report's own label rule forbids; §8 then calls it a 'measured contribution' although it is an estimate; it calls it a share of 'a stage agent's first-request context' although the brief, CONSTITUTION.md and the stage command all enter AFTER the first request (the agent's prompt is one line); and the byte sum omits the `pharn/ARCHITECTURE.md §6` read that /pharn-plan, /pharn-grill and /pharn-build prescribe (§6 alone is 4,754 B, the file 23,435 B), so the ratio's numerator understates PHARN-owned text for those three stages." + evidence: "report:47-48 'Figures with different sources are never expressed as percentages of each other.' report:190-191 'roughly 7–9% of CH1's median prefix [e]' report:470 'PHARN-owned text is roughly 7–9% of a stage agent's first-request context [e] … Negligible measured contribution next to the prefix' pharn-build.md:49 'Read the `pharn/ARCHITECTURE.md §6` build-stage row'" +# F6 +- type: FINDING + rule_id: "P0" + severity: important + file: ".dev/features/pipeline-performance-audit/audit.mjs:120" + problem: "`by_stage_context` labels a request 'main' only when `sidechain === false`, treating 'the session's main thread' as 'the orchestrator'; the cost-ledger contract says that since 6.29.0 the orchestrator may itself be an agent (membership.context `agent:`), whose own rows are `sidechain: true`, so for any /pharn-loop or /pharn-ship run started inside an agent every orchestrator request is counted as 'agent' — and C1/C2's one pre-registered adoption gate ('`main`-context requests are at least 20% of a run's requests') would read near 0% and stop; the function is not covered by --self-test and has never run on a real ledger (0 eligible), and §5's 'requests[].sidechain separates the orchestrator's rows from the agents' rows' carries the same conflation." + evidence: 'audit.mjs:120 ''const key = `${r.stage ?? "(unattributed)"} · ${r.sidechain === true ? "agent" : "main"}`;'' cost-ledger.md:473-474 ''Since 6.29.0 that holds whether the orchestrator is the session''s own thread or itself an agent'' cost-ledger.md:105 ''"context": "agent:"'' report:344-345 ''In real ledgers, `main`-context requests are at least 20% of a run''s requests. That threshold is pre-registered here.''' +# F7 +- type: FINDING + rule_id: "P6" + severity: important + file: ".dev/measurements/pipeline-performance-audit-2026-09-29.md:98" + problem: "§1 attributes to its sources more than they say: (a) CHANGELOG [6.32.0] does not say its 37 orchestrator transcripts predate 6.27.0 routing, and does not place them in pharn-starter (it says 'on the maintainer's machine'); (b) the checker's WARN reports how many contexts a run-window/1 ledger's rows come from, which is not proof of FOREIGN rows (a run's own spawned agents are contexts that run-window/2 also counts), yet §1 and §8 read the 49 as rows of concurrent or other contexts; (c) the PLAN's inventory item, the 2026-09-21 `pharn-loop-run` transcript (this repo's worktree, SKILLS_VERSION 6.6.0 on that date), is dropped from 'Other real run records' without a word, though the PLAN said the build re-counts each. The exclusion conclusions themselves still hold on independent evidence (pharn-starter records 6.12.1; 6.6.0 < 6.32.0)." + evidence: "report:97-98 'pharn-starter's install record reads 6.12.1 since 2026-09-23, and the 6.32.0 entry says its 37 predate 6.27.0 routing.' CHANGELOG [6.32.0] 'measured in 37 orchestrator transcripts on the maintainer's machine' report:88 '`run-window/1` counts concurrent contexts' rows. The checker names 49 such ledgers.' report:479 '49 carry other contexts' rows (L24)' PLAN.md:47-48 'one real `/pharn-loop` transcript in `~/.claude/projects/…pharn-loop-run/`, dated 2026-09-21'" +# F8 +- type: FINDING + rule_id: "P0" + severity: minor + file: ".dev/features/pipeline-performance-audit/audit.mjs:276" + problem: "The header says the helper 'writes nothing', but `--static --project` shells `stage-agent.mjs route … --name demo-feature` with cwd set to this checkout, and route, on exit 0, calls `clearResult`, which unlinks `.pharn//demo-feature/stage-result.json` if one exists; the recorded run deleted nothing only because every pharn-starter route exited 3, so a project whose config routes would make the helper perform a delete." + evidence: "audit.mjs:3 'It prints one JSON document on stdout and writes nothing: no file, no .pharn/ state, no git write.' stage-agent.mjs:514-515 'if (d.exit === 0) { const why = clearResult(root, o.command, o.name);' stage-agent.mjs:325 'unlinkSync(path);'" +# F9 +- type: FINDING + rule_id: "P0" + severity: minor + file: ".dev/measurements/pipeline-performance-audit-2026-09-29.md:7" + problem: "The report's sourcing rules are stated more strongly than they are kept: 'Every number below is printed by audit.mjs' is false for the CF1/CF2 figures (read from the two MEASUREMENT.md files) and the 2026-09-23 install date; §1's ledger counts and the §2 table's byte figures carry no source/precision label, and the label set has no member for real PRE-optimization evidence; C2's 'about −10 to −13' has no formula although derived figures promise one; and CH1's input set (the `find … -mtime -8` over this project's `subagents/` transcripts) is not stated, so the 37 / 102,546 cannot be reproduced once the window moves (today: 38, median 102,781)." + evidence: "report:7-8 'Every number below is printed by `.dev/features/pipeline-performance-audit/audit.mjs`. The mode is given with each figure, or the figure is arithmetic on printed numbers with its formula shown.' report:36 'Each figure carries a **source** and a **precision**' report:120 '`audit.mjs --prefix` over this repo's subagent transcripts' report:362 'Requests: about −10 to −13 per run [S·e]'" +# F10 +- type: FINDING + rule_id: "P0" + severity: minor + file: ".dev/features/pipeline-performance-audit/audit.mjs:255" + problem: "Smaller helper defects: (a) every build brief is measured at `--iteration 2` only (loop 4,063 B; iteration 1 is 3,598 B), undisclosed, so the report's '2,498–4,063 B' upper end applies only to a rebuild; (b) a crashed or unusable checker blocks the token metric but adds no closed reason, so such a ledger leaves the token denominator without appearing in `excluded_by_reason` (the rule says no exclusion is silent); (c) a ledger that parses to a non-object (e.g. `null`) throws in `classify` and crashes the whole run; (d) `--project` with no value silently drops the project section." + evidence: 'audit.mjs:255 ''if (stage === "pharn-build") argv.push("--iteration", "2");'' audit.mjs:78 ''if (checker === "red") reasons.push("ledger-red");'' audit.mjs:147-153 ''ledger = JSON.parse(readFileSync(p, "utf8")); … const c = classify(ledger, …'' report:168 ''then its brief, 2,498–4,063 B [S·m]''' +``` + +## Verdict + +**GREEN — 0 floor-gate findings.** 10 advisory findings, 0 of blocking severity: + +- **important ×5:** F1, F3, F5, F6, F7; +- **minor ×5:** F2, F4, F8, F9, F10. + +The floor passed, and the stage artifacts agree: regress `no-regressions`, verify `PASS`, both CHANGELOG checks GREEN. +Every figure the brief asked to re-check reproduces, or reproduces as of when it was taken. The report is honest about +its central limit: 0 of 69 real runs are post-optimization, and every candidate is gated on real-run confirmation. + +The advisory findings are about the report stating more than its evidence carries, in its own words: + +- a pinned-call count read as a share of requests (F1); +- token classes blended despite its own rule (F3); +- a cross-source percentage called "measured" (F5); +- a citation the source does not make (F7). + +One helper defect bears on the recommendations: F6 would misread the one pre-registered adoption gate for any run +started inside an agent. Fix it before the helper's first real-ledger pass. + +## Lesson candidate (proposed, not written — P2/P7) + +None proposed. The nearest candidate is "a context split read from `sidechain` conflates the session's main thread +with the orchestrator" (F6). This is its first occurrence, and nothing has consumed the helper's output on real data +yet, so it does not meet P7's real-recurrence bar (L20: the second occurrence is the trigger). If a later +cost-ledger consumer repeats it, propose it then, with provenance `pipeline-performance-audit` / F6. + +## Re-review after the GATE-2 fix build (2026-09-29) + +Re-reviewed by `/pharn-dev-review` after the fix build. The sections above are the first review, kept as written. The +increment is still `trust: untrusted`, and every `problem` / `evidence` value below is quoted DATA (P2). + +### Floor, re-run + +- `node pharn/floor/validate.mjs .` → **exit 0**, `FLOOR: GREEN — 36 capabilities checked`. +- `node .dev/features/pipeline-performance-audit/audit.mjs --self-test` → `{"ok": true, "checks": 25, "failed": []}`, + exit 0. +- `npm run check:changelog` → GREEN (`1 dated [Unreleased] entr(ies)`). `node .dev/floor/check-changelog-entry.mjs +--base-ref HEAD` → GREEN (1 new entry). +- Stage artifacts: `regression-report.json` `no-regressions` (the base map is reused from the first run, and + `REGRESSION.md` says so); `verify-report.json` `PASS`, `failing_gates: []`. + +No floor gate reads the report's figures, so everything below is **advisory**. + +### Figures re-run (advisory) + +- **`audit.mjs --static --project ~/Projects/pharn-starter`** at `c9737b4`. Every byte figure reproduces: + - 43,164 / 35,982 / 11,091 / 11,223 / 31,211 / 28,179; 19,848 / 17,929; 7,958; §6 4,754; 172,421 / 387,549; + - briefs 2,498–4,063 B, with the loop build at 3,598 (iteration 1) and 4,063 (iteration 2). + + Every routed cell for pharn-starter prints `inline:config-red`, exit 3. The block counts per section match (Step 1a + 5, Step 3 5, Step 4 14, Step 5 10; close 6b 8, 6c 5, Step 7 1, Final 2). + +- **`audit.mjs --ledgers ~/Projects/pharn-starter/pharn/features/*/cost.json`.** Universe 69; included 0 / 0; + denominators 0 / 0 / 0; `pre-optimization-version` ×69, `membership-run-window-1` ×56; 6.12.1 ×56, 6.7.0 ×13; + `partial` ×69; GREEN ×69; 49 multi-context (all 6.12.1); shared window starts 10 / 2 / 2; + `eligible_excluded_by_metric` empty. This matches §1. +- **`find … -path '*/subagents/agent-*.jsonl' -exec audit.mjs --since 2026-09-23 --until 2026-09-29T07:00 --prefix +{} +`.** 37 transcripts: + - `general-purpose` 29: cache write 68,223 / 91,759 / 129,646; cache read 0 / 30,927 / 31,428, non-zero in 15; + uncached input 2; + - `Explore` 6: cache write 34,934–46,732; + - `claude-code-guide` 2, on haiku: 72,869–73,357. + + This matches §3 and the CHANGELOG. The fixed window leaves out this re-review's own agent, as intended. + +- **Arithmetic, against the command text.** + - Inputs. `pharn-loop.md`: Step 1a 5; Step 3 5; Step 4 14, of which `check-red-run --preflight` (`pharn-loop.md:479`) + runs only after a non-zero `check-test-stage`, so 13; Step 5 10. `pharn-loop-close.md`: 6b 8, 6c 5, Step 7 1, + Final 2, so 16; 6a's 2 and 6d's 1 are off the green path. `pharn-regress.md` / `pharn-verify.md`: the invocation, + the Step 0 slug Write and `feature-name.mjs` line, the scope set, the script and the release, so 6 each. The loop + pins both without a name (`pharn-loop.md:535`). + - 5 + 5 + 13 + 10 + 16 + (2 × 6) + 5 + 4 = 70. 57 = 49 + 2 × 4. The other 13 = 5 + 2 + 2 + 4. 23 = 10 + 12 + 1, of + which 18 are deterministic. All **follow**. + - 5 × 91,759 = 458,795 (about 459k) and 5 × 30,927 = 154,635 (about 155k): **follow** (see R3 on "up to"). + - C1 −10 = 5 × 2. C2 −10 to −13 = 14 − k for k 1–4, and the tail is 6 + 5 + 1 + 2 = 14 (see R3 on which two blocks + precede the Write). C3 −6 = 2 × 3. 37,777 / 4 = 9,444 and 43,164 / 4 = 10,791. 2,498 + 7,958 + 19,125 = 29,581 + and 3,266 + 7,958 + 26,261 = 37,485. All **follow**. +- **Home paths.** None in any file of the increment. A grep for `/Users/`, the account name, `/private/` and + `/var/folders` hits only the first review's own list of those patterns. + +### F1–F10 status + +- **F1 — fixed.** Report:25-28, "Most PINNED orchestrator calls are deterministic … The share of ALL orchestrator + requests these make in a real run is **not** known"; report:585 strikes "most orchestrator requests are ceremony". +- **F2 — fixed.** (a) report:283-284 counts Step 4 as 13; (b) report:288-294 adds each inline stage's slug Write and + `feature-name.mjs` line, and the `LOOP.md` Write; (c) report:296 uses (2 × 6) and names the six; (d) report:297-299 + derives 57 as 49 + 2 × 4. 70 is now a valid green-path lower bound. Mandated calls it still leaves out are R2. +- **F3 — fixed.** Report:59-60 keeps the classes apart; report:213-218 tabulates CH1 per class; report:320-321 gives + 459k cache write and 155k cache read separately. The blended ~0.5M and ~1M figures are gone. +- **F4 — fixed.** audit.mjs:396-403 and 417-418 group by the `.meta.json` agent type. Report:147-149 and 212-220 use + the 29 `general-purpose` agents only, and the CHANGELOG says "over 29 such subagents in this repo". +- **F5 — fixed.** Report:232-233, "No ratio of the two is stated (different sources)"; report:204 and 230 add §6 + (4,754 B); report:545 no longer calls the share "measured". +- **F6 — fixed.** audit.mjs:141-145 (`requestRole`) compares the row's context with `membership.context`, and the + self-tests at audit.mjs:472-494 cover a run started inside an agent. Report:344-347 and 415-416 use the orchestrator + role. +- **F7 — fixed.** (a) Report:113-114 cites `.dev/features/orchestrator-context/PLAN.md`, whose line 288 does say + "predate 6.27.0 routing"; (b) report:101-103 says the WARN does not prove foreign rows; (c) report:109-110 lists the + `pharn-loop-run` transcript and excludes it by date. +- **F8 — fixed.** audit.mjs:4-6 states the exception; audit.mjs:307, 332 and 339 run the route probe with its cwd in a + temp directory that is then removed. `brief` writes nothing (stage-agent.mjs:526-538). +- **F9 — fixed as raised.** Report:12-14 names the CF1/CF2 and install-date exceptions; report:67 labels §1 `[H·m]` + and report:176 the §2 table; report:48 adds `H`; report:434-435 gives 14 − k; report:140-149 states CH1's input set + over a fixed window. The fix added a quoted figure outside that list (R3). +- **F10 — fixed.** (a) audit.mjs:315 renders the build brief at iterations 1 and 2, and report:200-201 says which is + which; (b) audit.mjs:100 names `checker-` per metric; (c) audit.mjs:190 reads a non-object as `unreadable` + (re-run: `null` and `[1]` both read `unreadable`, no crash); (d) audit.mjs:533 exits 2 on a bare `--project` + (re-run: exit 2). + +### Lenses, re-applied + +- **L-floor (P0).** The report still claims no floor guarantee. Its own discipline rules now hold, except where R1–R3 + say otherwise. +- **L-eval (P1).** No `role:`-bearing capability was added. validate is GREEN over the same 36 capabilities. +- **L-trust (P2).** The new route probe and the `--prefix` grouping print only closed tokens, numbers, basenames, model + ids and the `.meta.json` agent type. The checker and probes are still spawned with argv arrays and no shell. Nothing + in the reviewed files addressed this reviewer or changed its behaviour. +- **L-axis (P3).** The same two imports, in the permitted `.dev/` → `pharn/` direction. No sibling reference. + +### Floor-gate findings (blocking) + +None. + +### New advisory findings + +```yaml +# R1 +- type: FINDING + rule_id: "P0" + severity: minor + file: ".dev/measurements/pipeline-performance-audit-2026-09-29.md:32" + problem: "The report and the CHANGELOG say unconditionally that the project's `test` gate runs at least three times per iteration, but at regress `test` runs only over the outside-feature test files and is recorded `no-files`, without running, when that subset is empty (a case the regress script proceeds through silently, such as every test file being declared by the PLAN or AC-TESTS.md); the static floor is therefore two (the build agent's gate and verify), with three or four only when an outside-feature test file exists." + evidence: 'report:32 ''**The project''s `test` gate runs at least three times per iteration:**'' CHANGELOG [Unreleased] ''the project''s `test` gate running at least three times per iteration'' run-gates.mjs:964-967 ''if (needsFiles && next.files.length === 0 && rec.stage === "regress") { exit = 0; ran = false; reason = "no-files";'' stage-regress.mjs:636-638 ''An EMPTY outside-scope partition with a NON-empty universe (every discovered test happens to live INSIDE the feature) is legitimate and must proceed silently''' +# R2 +- type: FINDING + rule_id: "P0" + severity: minor + file: ".dev/measurements/pipeline-performance-audit-2026-09-29.md:304" + problem: "70 is called 'the count of pinned or mandated calls', but the green path mandates further orchestrator calls it leaves out: Step 1b's lookup of a prior LOOP.md, Step 6b's Read of the loop-record contract, Step 6c's post-commit `git rev-parse HEAD`, and Step 7's reads of RUN-REPORT.md and the pre-run snapshot; so 70 is a lower bound on that count (up to about 75 with these). Separately, 'every call is a model request' does not hold for independent calls issued in one turn, which the loop command itself says add no request." + evidence: "report:304-305 'So 70 is the count of pinned or mandated calls, not the run's request count.' report:27 'Every call is a model request that carries the orchestrator's whole context.' pharn-loop.md:171 'Look for `pharn/features/-/LOOP.md` with the highest existing ``' pharn-loop-close.md:50-51 'Read the contract and follow its canonical template' pharn-loop-close.md:250 'On success, capture the SHA for the summary (`git rev-parse HEAD`).' pharn-loop-close.md:311 'print `pharn/features//RUN-REPORT.md`'s `## Tokens` table and its `## Files` list' pharn-loop-close.md:308 'any committed path that was already dirty in the pre-run snapshot' pharn-loop.md:321 '(two Reads in one turn add no request)'" +# R3 +- type: FINDING + rule_id: "P0" + severity: minor + file: ".dev/measurements/pipeline-performance-audit-2026-09-29.md:420" + problem: "Small inaccuracies the fix build introduced: (a) C2 names 'the LOOP.md scope set and its amend' as the two close blocks that must precede the LOOP.md Write, but those two lines are one fenced block, and the second pre-Write block is Step 6b's `git rev-parse HEAD` commit capture (the 14-block tail still holds); (b) the sourcing note lists two exceptions to 'figures come from audit.mjs', but the fix added a third quoted figure, CHANGELOG [6.32.0]'s 11–426 requests, median about 200; (c) 'up to about 155k' is five times the median hit (30,927), not the largest (5 × 31,428 = 157,140), and 155k is missing from the claims block's list of estimates; (d) the 459k, 155k and 9.4k figures carry a precision label (`[e]`) but no source, although the label rule says every figure carries both." + evidence: "report:420-421 'Two of them (the `LOOP.md` scope set and its amend) must precede the model's `LOOP.md` Write.' pharn-loop-close.md:43-46 (one bash fence holding set-writes-scope.cjs and reconcile-baseline.mjs --amend-scope) pharn-loop-close.md:62-64 'git rev-parse HEAD 2>/dev/null || echo unknown' report:12-14 'There are two exceptions: CF1/CF2 are quoted from their `MEASUREMENT.md` files, and pharn-starter's install date' report:305-306 'the 6.32.0 entry measured pre-routing loop runs at 11–426 requests, median about 200 `[H·m]`' report:321 'plus up to about 155k cache-read tokens [e] (5 × 30,927)' report:579 'Every estimate: 70 calls, 57, 23 per iteration, about 459k, −10, −10 to −13, −6, about 9.4k, about 10.8k.' report:43 'Each figure carries a **source** and a **precision**'" +``` + +None of R1–R3 changes a recommendation. R1 and R2 are structural statements a little stronger than the code supports. +R3's items are wording and labels. + +### Re-review verdict + +**GREEN — 0 floor-gate findings.** F1–F10: all 10 fixed. **3 open advisory findings** (R1, R2, R3, all minor), none of +blocking severity. The report's own claim that it "applies the review's findings F1–F10", and the CHANGELOG's "the +review's ten advisory findings were fixed", both hold. R1 also applies to the CHANGELOG entry's wording. + +No lesson candidate is proposed (P7): R1–R3 are first occurrences. diff --git a/.dev/features/pipeline-performance-audit/SHIP.md b/.dev/features/pipeline-performance-audit/SHIP.md new file mode 100644 index 00000000..61bd2645 --- /dev/null +++ b/.dev/features/pipeline-performance-audit/SHIP.md @@ -0,0 +1,87 @@ +# SHIP — pipeline-performance-audit + +A `/pharn-dev-ship` roll-up. It records that the chain ran and the floor verdicts it read. It is not an approval. + +## Where the run ended + +**GATE 2, twice.** + +- **The first GATE 2.** The maintainer chose **"Fix, then PR"** (2026-09-29, an interactive form). That is the + maintainer's decision, not the model's. +- **The fix pass.** One fix build inside the plan's `## Files`, then regress, verify and a fresh re-review. +- **Then a PR, as the maintainer asked.** + +### First pass + +1. **`/pharn-dev-plan`.** Written in the orchestrator's context (opus). GATE 1: the maintainer approved the plan as + written, and chose Q1 option (1), "Audit now" (recorded in `PLAN.md`, "Resolved at GATE 1"). +2. **`/pharn-dev-grill`.** A fresh opus subagent. Step 1b, `check-plan-lessons.mjs`: **exit 0** → proceed. It raised + 15 advisory concerns (0 blocking): `GRILL.md`. +3. **`/pharn-dev-build`.** Run inline, in the orchestrator's context (opus). `pharn.config.json` routes `build` to + sonnet; running it inline was the orchestrator's choice, because the build needed the discovery already in that + context. It was not a config routing. The build wrote: + - `audit.mjs`; + - the report `.dev/measurements/pipeline-performance-audit-2026-09-29.md`; + - the `CHANGELOG.md` entry. + + It addressed the grill's findings inside the plan's files, and disclosed the rule gaps without amending the frozen + rule. Floor: **`validate.mjs` exit 0** → proceed. + +4. **`/pharn-dev-regress`.** Inline. `.verdict` **`no-regressions`** → proceed. +5. **`/pharn-dev-verify`.** Inline. `.verdict` **`PASS`** (7 gates, `reconcile` CLEAN) → proceed. +6. **`/pharn-dev-review`.** A fresh opus subagent. 0 floor-gate findings and 10 advisory findings (F1–F10) in + `REVIEW.md`. + +### Fix pass (after the maintainer's "Fix, then PR") + +1. **Fix build.** Inline, under a fresh `--from-plan` scope and a new reconcile epoch. It fixed F1–F10 in `audit.mjs` + (25 self-tests), the report and the `CHANGELOG.md` entry. Floor: **`validate.mjs` exit 0**. +2. **`/pharn-dev-regress`.** `.verdict` **`no-regressions`**. The HEAD side re-ran. The BASE map was reused from the + first run, because the base SHA, the outside gate set and the install decision were all unchanged. This is + disclosed in `REGRESSION.md`. +3. **`/pharn-dev-verify`.** `.verdict` **`PASS`** (7 gates, `reconcile` CLEAN). +4. **Re-review.** A fresh opus subagent. Floor GREEN; F1–F10 **all fixed**; 3 new minor advisory findings (R1–R3, + `REVIEW.md` "Re-review"). +5. **R1–R3 fixed**, as report and CHANGELOG prose, under a fresh `--from-plan` scope and epoch. **Not** re-run after + this fix: `npm test`, `regress` and a third review. Instead the gates the change can move were run: + - `format:check` 0, `lint:md` 0, `lint` 0, `check:changelog` 0, `check:changelog-entry` 0; + - `validate` 0; + - the two CHANGELOG test files 0; + - `audit.mjs --self-test` 0; + - `reconcile` CLEAN (2 paths, no escapes). + + CI runs the full suite on the pull request. + +## Structural verdicts read, verbatim + +| stage | read | first pass | fix pass | +| -------------------- | ----------------------------- | ------------------ | ------------------ | +| `/pharn-dev-grill` | `check-plan-lessons.mjs` exit | `0` | (not re-run) | +| `/pharn-dev-build` | `validate.mjs` exit | `0` | `0` | +| `/pharn-dev-regress` | `.verdict` | `"no-regressions"` | `"no-regressions"` | +| `/pharn-dev-verify` | `.verdict` | `"PASS"` | `"PASS"` | + +## Pointers + +- Review and re-review: `.dev/features/pipeline-performance-audit/REVIEW.md` (advisory). +- Grill: `.dev/features/pipeline-performance-audit/GRILL.md` (advisory). + +## Recorded lines + +changelog-entry: exit 0 + +lesson: none — the nearest candidate (REVIEW F6: an orchestrator/agent split read from `sidechain` misreads an +orchestrator that itself runs inside an agent) is a first occurrence with no consumer yet. The remaining findings are +instances of existing lessons (L6 structured location, L40 attribution, L43 agreement). So nothing clears L20's bar. + +deferred: none + +## Notes for the human + +- **A local-only housekeeping step.** The git-ignored `.pharn/pr-body.md` (the merged PR #298's body, from an earlier + session) was renamed in place to `.pharn/pr-body.md.bak`, so that the whole-repo `lint:md` gate measured this + increment rather than that file (L61). Nothing was deleted. +- **The commit, branch and PR are made at the maintainer's GATE-2 choice.** No merge and no seal. + +Chain ran; the named floor verdicts are as shown. This is NOT a judgment that the increment is good or wise; that is +the human's call at the post-review gate. diff --git a/.dev/features/pipeline-performance-audit/VERIFY.md b/.dev/features/pipeline-performance-audit/VERIFY.md new file mode 100644 index 00000000..21864b4d --- /dev/null +++ b/.dev/features/pipeline-performance-audit/VERIFY.md @@ -0,0 +1,28 @@ +# VERIFY — pipeline-performance-audit + +This is the second run, after the GATE-2 fix build (the maintainer chose "Fix, then PR"). The first run was also +`PASS`. The gates ran at HEAD, on the working tree with the fixed increment present. + +| gate | exit | +| ------------------------------------------------------------------------------------------ | ---: | +| `test` (`npm test`) | 0 | +| `validate` (`pharn/floor/validate.mjs .`) | 0 | +| `lint` | 0 | +| `format:check` | 0 | +| `lint:md` | 0 | +| `structural:pharn/pharn-review/trust-fence/evals/expected/expected-injection-comment.json` | 0 | +| `reconcile` (`check-bash-reconcile.mjs --require-baseline`) | 0 | + +**VERIFIED: floor gates PASS** (`check-verify.mjs`, exit 0, `failing_gates: []`). + +- **`reconcile`:** `CLEAN`. The epoch was re-anchored by the fix build (`pharn-dev-build`) at 2026-09-29T07:37:23Z. + It reconciled 3 paths and found no escapes. +- **`lint:md`, a local note.** Before the first run, a git-ignored `.pharn/pr-body.md` sat in the checkout. It was left + by an earlier session and held the body of already-merged PR #298, and it would have turned the whole-repo `lint:md` + red. It is not this increment's file and CI never sees it (L61). Following L61's precedent, it was moved aside in + place, to `.pharn/pr-body.md.bak`, and not deleted. +- **Verifiers:** none registered (`count-verifiers.mjs` → `{"registered":0}`), so these are the floor gates only. + +"Verified" means the named gates passed. It is not a guarantee of correctness beyond what those gates check. In +particular, no gate checks this increment's report for accuracy: its figures are advisory, and `/pharn-dev-review` +is the stage that reads them. diff --git a/.dev/features/pipeline-performance-audit/audit.mjs b/.dev/features/pipeline-performance-audit/audit.mjs new file mode 100644 index 00000000..7377c582 --- /dev/null +++ b/.dev/features/pipeline-performance-audit/audit.mjs @@ -0,0 +1,546 @@ +#!/usr/bin/env node +// audit.mjs — the READ-ONLY analysis helper of the pipeline-performance-audit (apparatus; never shipped, never +// imported by PHARN). It prints one JSON document on stdout and writes nothing in the repository: no file, no .pharn/ +// state, no git write. ONE bounded exception, stated (REVIEW F8): `--static --project` runs `stage-agent.mjs route`, +// which removes a leftover `.pharn///stage-result.json` under ITS cwd on an `agent:` route, so the +// helper runs that probe with its cwd in an empty temp directory it creates and removes. The figures of +// `.dev/measurements/pipeline-performance-audit-2026-09-29.md` come from these modes, except those the report quotes +// from the two MEASUREMENT.md files and pharn-starter's install record. +// +// node .dev/features/pipeline-performance-audit/audit.mjs --self-test +// node .dev/features/pipeline-performance-audit/audit.mjs --ledgers [--fixture ] +// node .dev/features/pipeline-performance-audit/audit.mjs --static [--project ] +// node .dev/features/pipeline-performance-audit/audit.mjs --prefix [--since ] [--until ] +// +// --ledgers applies the run-selection rule PLAN.md froze at GATE 1 (verbatim; the rule's own gaps are stated in the +// report, never repaired here). Versions compare NUMERICALLY (a string compare orders "6.7.0" after "6.35.0" — GRILL +// T1). `skills_version` is the configured version, which the cost-ledger contract labels advisory (GRILL G1), so the +// structural evidence (schema, membership method, the 6.35.0 keys) is printed beside it. Whether a ledger comes from a +// fixture is the CALLER's declaration (`--fixture`): nothing in a cost.json records it. The checker is shelled with an +// argv array and no shell (GRILL S1), and its exit is read through shelledVerdict, so a crash is never a RED (GRILL E1). +// Rule 3's per-metric exclusions of an ELIGIBLE ledger are listed per metric with why (REVIEW F10). Per-stage figures +// are READ from the views cost.json stores — never re-derived (GRILL G10). The orchestrator's own requests are the rows +// whose CONTEXT is the run's own, `membership.context` — never "the main thread", because a run started inside an +// agent has an `agent:` orchestrator (cost-ledger.md "The context half"; REVIEW F6). +// --static measures bytes in THIS checkout (the commit is printed) and, with --project, that project's CLAUDE.md and +// the route its models block would get. --prefix reads each transcript's FIRST request usage per token class, grouped +// by the agent type its sibling `.meta.json` records: a harness observation, never a PHARN run (REVIEW F3, F4). +// Every printed path has the home directory replaced by `~` (GRILL PR1). +import { readFileSync, statSync, existsSync, mkdtempSync, rmSync } from "node:fs"; +import { spawnSync, execFileSync } from "node:child_process"; +import { homedir, tmpdir } from "node:os"; +import { join, resolve, basename, dirname } from "node:path"; +import { fileURLToPath } from "node:url"; +import { shelledVerdict } from "../../../pharn/floor/shelled-verdict-core.mjs"; +import { ROUTE_POLICY } from "../../../pharn/floor/stage-agent-core.mjs"; + +const ROOT = resolve(dirname(fileURLToPath(import.meta.url)), "../../.."); +const HOME = homedir(); + +/** Replace the home directory with `~` in any string. */ +export function redact(s) { + return typeof s === "string" && HOME && s.startsWith(HOME) ? `~${s.slice(HOME.length)}` : s; +} + +/** Numeric SemVer compare over MAJOR.MINOR.PATCH; null when either side is not that shape. */ +export function cmpVersion(a, b) { + const re = /^(\d+)\.(\d+)\.(\d+)$/; + const x = typeof a === "string" && a.match(re); + const y = typeof b === "string" && b.match(re); + if (!x || !y) return null; + for (let i = 1; i <= 3; i++) { + const d = Number(x[i]) - Number(y[i]); + if (d !== 0) return Math.sign(d); + } + return 0; +} + +/** A parsed ledger the rule can read: a plain object, never null, an array or a scalar. */ +export function isLedgerObject(v) { + return v !== null && typeof v === "object" && !Array.isArray(v); +} + +// ---- the frozen rule (PLAN.md, "The run-selection rule") ---- +export const COMMANDS = Object.freeze(["/pharn-ship", "/pharn-loop"]); +export const FULL_PROFILE_MIN = "6.35.0"; +export const MODEL_USAGE_MIN = "6.32.0"; +export const REASONS = Object.freeze([ + "pre-optimization-version", + "wrong-command", + "fixture-not-workload", + "ledger-red", + "membership-run-window-1", + "coverage-unavailable", + "unreadable", +]); + +/** + * Eligibility of one parsed ledger under the frozen rule. `checker` is "green" | "red" | "crashed" | "unusable". + * Returns {tier, reasons[], metrics{tokens, elapsed, work}, metric_exclusions{tokens[], elapsed[], work[]}}. A version + * that is not MAJOR.MINOR.PATCH is not ≥ anything, so it reads pre-optimization. `metric_exclusions` names, for an + * ELIGIBLE ledger only, which of rule 3's conditions failed for each metric; these are the rule's own conditions + * spelled out, not new reasons (the rule-level `reasons[]` stay the frozen set). + */ +export function classify(ledger, { fixture, checker }) { + const reasons = []; + if (!COMMANDS.includes(ledger.command)) reasons.push("wrong-command"); + const v = ledger.skills_version; + const full = cmpVersion(v, FULL_PROFILE_MIN); + const usage = cmpVersion(v, MODEL_USAGE_MIN); + const tier = full !== null && full >= 0 ? "full" : usage !== null && usage >= 0 ? "model-usage" : null; + if (tier === null) reasons.push("pre-optimization-version"); + if (fixture) reasons.push("fixture-not-workload"); + if (checker === "red") reasons.push("ledger-red"); + const method = ledger.membership?.method ?? null; + if (method === "run-window/1") reasons.push("membership-run-window-1"); + if (ledger.coverage === "unavailable") reasons.push("coverage-unavailable"); + const eligible = tier !== null && !reasons.includes("wrong-command") && !fixture; + const ex = { tokens: [], elapsed: [], work: [] }; + if (eligible) { + if (checker !== "green") ex.tokens.push(`checker-${checker}`); + if (method !== "run-window/2") ex.tokens.push(`membership-${method ?? "absent"}`); + if (ledger.coverage === "unavailable") ex.tokens.push("coverage-unavailable"); + if (tier !== "full") { + ex.elapsed.push("version-below-6.35.0"); + ex.work.push("version-below-6.35.0"); + } else { + if (!Array.isArray(ledger.executions?.rows)) ex.elapsed.push("no-executions-rows"); + if (!Array.isArray(ledger.work)) ex.work.push("no-work-array"); + } + } + return { + tier: eligible ? tier : null, + reasons, + metrics: { + tokens: eligible && ex.tokens.length === 0, + elapsed: eligible && ex.elapsed.length === 0, + work: eligible && ex.work.length === 0, + }, + metric_exclusions: ex, + }; +} + +function runChecker(path) { + const r = spawnSync(process.execPath, [join(ROOT, "pharn/floor/check-cost-ledger.mjs"), path], { + encoding: "utf8", + shell: false, + maxBuffer: 64 * 1024 * 1024, + }); + const v = r.status === 2 ? "unusable" : shelledVerdict(r); + const out = r.stdout ?? ""; + const warns = out.split("\n").filter((l) => l.startsWith("WARN")).length; + const ctx = out.match(/rows come from (\d+) context/); + return { verdict: v, exit: r.status, warns, contexts: ctx ? Number(ctx[1]) : null }; +} + +/** + * The role of one stored request row in its run: `orchestrator` when the row's CONTEXT (`main`, or + * `agent:` for a sidechain row — cost-ledger.md "Field shape") is the run's own `membership.context`; + * `stage-agent` for any other context the run admitted; `unknown-context` when the ledger names no run context. + */ +export function requestRole(row, runContext) { + if (typeof runContext !== "string") return "unknown-context"; + const ctx = row.sidechain === true ? `agent:${row.agent_id}` : "main"; + return ctx === runContext ? "orchestrator" : "stage-agent"; +} + +/** Per-stage figures read from the ledger's STORED views and rows (never recomputed as a view). */ +export function storedProfile(ledger) { + const byStage = {}; + for (const row of ledger.by_stage_iteration_model ?? []) { + const s = (byStage[row.stage ?? "(unattributed)"] ??= { requests: 0, tokens: {} }); + s.requests += row.requests ?? 0; + for (const [k, n] of Object.entries(row.tokens ?? {})) s.tokens[k] = (s.tokens[k] ?? 0) + n; + } + const runContext = ledger.membership?.context ?? null; + const byStageRole = {}; + for (const r of ledger.requests ?? []) { + const key = `${r.stage ?? "(unattributed)"} · ${requestRole(r, runContext)}`; + const s = (byStageRole[key] ??= { requests: 0, tokens: {} }); + s.requests += 1; + for (const [k, n] of Object.entries(r.tokens ?? {})) s.tokens[k] = (s.tokens[k] ?? 0) + n; + } + const routes = (ledger.markers ?? []) + .filter((m) => m.kind === "stage-start") + .map((m) => ({ stage: m.stage, iteration: m.iteration ?? null, route: m.route ?? null })); + return { + run_context: runContext, + totals: ledger.totals ?? null, + unattributed: ledger.unattributed ?? null, + by_stage: byStage, + by_stage_role: byStageRole, + served_models: ledger.by_model ?? null, + requested_routes: routes, + executions: ledger.executions ?? null, + work: ledger.work ?? null, + }; +} + +function ledgersMode(paths, fixtures) { + const rows = []; + const all = [...paths.map((p) => [p, false]), ...fixtures.map((p) => [p, true])]; + for (const [p, fixture] of all) { + const name = basename(dirname(p)); + let ledger; + try { + ledger = JSON.parse(readFileSync(p, "utf8")); + } catch { + ledger = undefined; + } + if (!isLedgerObject(ledger)) { + rows.push({ name, path: redact(p), reasons: ["unreadable"], tier: null }); + continue; + } + const checker = runChecker(p); + const c = classify(ledger, { fixture, checker: checker.verdict }); + rows.push({ + name, + path: redact(p), + command: ledger.command ?? null, + skills_version: ledger.skills_version ?? null, + schema: ledger.schema ?? null, + membership_method: ledger.membership?.method ?? null, + coverage: ledger.coverage ?? null, + has_executions: Object.hasOwn(ledger, "executions"), + has_work: Object.hasOwn(ledger, "work"), + window_start: ledger.window_start ?? null, + checker, + tier: c.tier, + reasons: c.reasons, + metrics: c.metrics, + metric_exclusions: c.tier ? c.metric_exclusions : undefined, + profile: c.tier ? storedProfile(ledger) : undefined, + }); + } + const count = (f) => rows.reduce((m, r) => ((m[f(r)] = (m[f(r)] ?? 0) + 1), m), {}); + const byReason = {}; + for (const r of rows) for (const x of r.reasons) byReason[x] = (byReason[x] ?? 0) + 1; + const byMetricExclusion = {}; + for (const r of rows) + for (const [m, xs] of Object.entries(r.metric_exclusions ?? {})) + for (const x of xs) byMetricExclusion[`${m}:${x}`] = (byMetricExclusion[`${m}:${x}`] ?? 0) + 1; + const sharedWindow = Object.entries(count((r) => r.window_start ?? "(none)")).filter(([, n]) => n > 1); + return { + rule: { commands: COMMANDS, full_profile_min: FULL_PROFILE_MIN, model_usage_min: MODEL_USAGE_MIN, reasons: REASONS }, + universe: rows.length, + included: { full: rows.filter((r) => r.tier === "full").length, model_usage_only: rows.filter((r) => r.tier === "model-usage").length }, + per_metric_denominators: { + tokens: rows.filter((r) => r.metrics?.tokens).length, + elapsed: rows.filter((r) => r.metrics?.elapsed).length, + work: rows.filter((r) => r.metrics?.work).length, + }, + excluded_by_reason: byReason, + eligible_excluded_by_metric: byMetricExclusion, + by_version: count((r) => r.skills_version ?? "(none)"), + by_schema_and_membership: count((r) => `${r.schema} · ${r.membership_method ?? "no membership"}`), + by_coverage: count((r) => r.coverage ?? "(none)"), + with_executions_or_work: rows.filter((r) => r.has_executions || r.has_work).length, + checker_verdicts: count((r) => r.checker?.verdict ?? "(not run)"), + checker_rows_from_two_or_more_contexts: rows.filter((r) => (r.checker?.contexts ?? 0) > 1).length, + shared_window_starts: sharedWindow, + rows: rows.map((r) => ({ ...r, window_start: undefined })), + }; +} + +// ---- --static ---- +function bytes(p) { + return statSync(join(ROOT, p)).size; +} + +/** Fenced ```bash blocks per `##`/`###` section of one command file — a count of PINNED shell lines, not of calls. */ +export function bashBlocksBySection(text) { + let sec = "(top)"; + let inFence = false; + const counts = {}; + for (const l of text.split("\n")) { + const f = l.match(/^\s*```(\w*)/); + if (f) { + if (!inFence && f[1] === "bash") counts[sec] = (counts[sec] ?? 0) + 1; + inFence = !inFence; + continue; + } + const h = !inFence && l.match(/^#{2,3} (.*)/); + if (h) sec = h[1].slice(0, 70); + } + return counts; +} + +/** Bytes of one `## .` section of a markdown file (heading through the line before the next `## `). */ +export function sectionBytes(text, number) { + const lines = text.split("\n"); + const start = lines.findIndex((l) => l.startsWith(`## ${number}. `)); + if (start < 0) return null; + let end = lines.findIndex((l, i) => i > start && l.startsWith("## ")); + if (end < 0) end = lines.length; + return Buffer.byteLength(lines.slice(start, end).join("\n") + "\n"); +} + +function sumBlocks(text) { + return Object.values(bashBlocksBySection(text)).reduce((a, b) => a + b, 0); +} + +function brief(command, stage, mode, iteration) { + const argv = [join(ROOT, "pharn/floor/stage-agent.mjs"), "brief", "--command", command, "--stage", stage, "--name", "demo-feature"]; + if (iteration !== null) argv.push("--iteration", String(iteration)); + if (mode === "quick") argv.push("--mode", "quick"); + const r = spawnSync(process.execPath, argv, { cwd: ROOT, encoding: "utf8", shell: false }); + return { exit: r.status, bytes: Buffer.byteLength(r.stdout ?? "") }; +} + +function staticMode(project) { + const commit = execFileSync("git", ["rev-parse", "HEAD"], { cwd: ROOT, encoding: "utf8" }).trim(); + const families = {}; + for (const cmd of ["pharn-loop", "pharn-ship"]) { + const files = [`${cmd}.md`, `${cmd}-quick.md`, `${cmd}-close.md`]; + families[cmd] = files.map((f) => { + const p = `.claude/commands/${f}`; + const blocks = bashBlocksBySection(readFileSync(join(ROOT, p), "utf8")); + return { file: p, bytes: bytes(p), bash_blocks: Object.values(blocks).reduce((a, b) => a + b, 0), bash_blocks_by_section: blocks }; + }); + } + const stageCommands = {}; + for (const s of ["pharn-spec", "pharn-plan", "pharn-grill", "pharn-test", "pharn-build", "pharn-regress", "pharn-verify"]) { + const p = `.claude/commands/${s}.md`; + const text = readFileSync(join(ROOT, p), "utf8"); + stageCommands[s] = { bytes: bytes(p), bash_blocks: sumBlocks(text), cites_architecture_s6: text.includes("ARCHITECTURE.md §6") }; + } + const probeDir = project ? mkdtempSync(join(tmpdir(), "audit-route-probe-")) : null; + const cells = []; + try { + for (const [command, modes] of Object.entries(ROUTE_POLICY)) + for (const [mode, row] of Object.entries(modes)) + for (const [stage, cell] of Object.entries(row)) { + const entry = { command, mode, stage, cell }; + if (cell === "agent") { + const its = stage === "pharn-build" ? [1, 2] : [null]; + entry.briefs = its.map((it) => ({ iteration: it, ...brief(command, stage, mode, it) })); + entry.command_bytes = stageCommands[stage].bytes; + if (project) { + const ra = [ + join(ROOT, "pharn/floor/stage-agent.mjs"), + "route", + "--command", + command, + "--stage", + stage, + "--name", + "demo-feature", + ]; + if (stage === "pharn-build") ra.push("--iteration", "1"); + if (mode === "quick") ra.push("--mode", "quick"); + ra.push("--config", join(project, "pharn.config.json")); + const rr = spawnSync(process.execPath, ra, { cwd: probeDir, encoding: "utf8", shell: false }); + entry.project_route = { exit: rr.status, token: (rr.stdout ?? "").trim().split("\n").pop() }; + } + } + cells.push(entry); + } + } finally { + if (probeDir) rmSync(probeDir, { recursive: true, force: true }); + } + const trusted = Object.fromEntries( + ["pharn/CONSTITUTION.md", "pharn/ARCHITECTURE.md", "THREAT-MODEL.md", "LIMITS.md"].map((p) => [p, bytes(p)]) + ); + const out = { + commit, + note: "bytes are measured in this checkout; any token figure derived from them is an ESTIMATE with no token class", + orchestrator_families: families, + stage_commands: stageCommands, + route_cells: cells, + trusted_docs: trusted, + architecture_section_6_bytes: sectionBytes(readFileSync(join(ROOT, "pharn/ARCHITECTURE.md"), "utf8"), 6), + this_repo_claude_md_bytes: bytes("CLAUDE.md"), + }; + if (project) { + const cm = join(project, "CLAUDE.md"); + let sv = null; + try { + sv = JSON.parse(readFileSync(join(project, "pharn.config.json"), "utf8")).skillsVersion ?? null; + } catch { + // an unreadable or malformed config leaves `sv` null + } + out.project = { path: redact(project), claude_md_bytes: existsSync(cm) ? statSync(cm).size : null, skills_version: sv }; + } + return out; +} + +// ---- --prefix ---- +function firstRequest(path) { + let text; + try { + text = readFileSync(path, "utf8"); + } catch { + return null; + } + for (const l of text.split("\n")) { + if (!l) continue; + let j; + try { + j = JSON.parse(l); + } catch { + continue; + } + if (j.type !== "assistant" || !j.message?.usage) continue; + const u = j.message.usage; + return { + ts: typeof j.timestamp === "string" ? j.timestamp : null, + model: j.message.model ?? null, + input: u.input_tokens ?? 0, + cache_write: u.cache_creation_input_tokens ?? 0, + cache_read: u.cache_read_input_tokens ?? 0, + }; + } + return null; +} + +function agentType(path) { + try { + const m = JSON.parse(readFileSync(path.replace(/\.jsonl$/, ".meta.json"), "utf8")); + return typeof m.agentType === "string" ? m.agentType : "(no type)"; + } catch { + return "(no meta)"; + } +} + +/** min / median (lower middle) / max of a numeric list; nulls on an empty list. */ +export function stats(values) { + const v = [...values].sort((a, b) => a - b); + if (v.length === 0) return { n: 0, min: null, median: null, max: null }; + return { n: v.length, min: v[0], median: v[Math.floor((v.length - 1) / 2)], max: v.at(-1) }; +} + +function prefixMode(paths, since, until) { + const rows = paths + .map((p) => ({ file: basename(p), agent_type: agentType(p), ...(firstRequest(p) ?? { none: true }) })) + .filter((r) => !r.none && r.ts !== null && (!since || r.ts >= since) && (!until || r.ts < until)); + rows.sort((a, b) => a.ts.localeCompare(b.ts)); + const byType = {}; + for (const r of rows) (byType[r.agent_type] ??= []).push(r); + const summarize = (rs) => ({ + uncached_input: stats(rs.map((r) => r.input)), + cache_write: stats(rs.map((r) => r.cache_write)), + cache_read: stats(rs.map((r) => r.cache_read)), + first_request_total_all_classes: stats(rs.map((r) => r.input + r.cache_write + r.cache_read)), + with_cache_read: rs.filter((r) => r.cache_read > 0).length, + first_ts: rs[0]?.ts ?? null, + last_ts: rs.at(-1)?.ts ?? null, + }); + return { + note: "harness observation: each transcript's FIRST request only (its fixed prefix), per token class, not a PHARN run; the all-classes total is shown for reference and is NOT a token class", + window: { since: since ?? null, until: until ?? null }, + transcripts: rows.length, + by_agent_type: Object.fromEntries(Object.entries(byType).map(([t, rs]) => [t, summarize(rs)])), + rows: rows.map((r) => ({ ...r, ts: r.ts.slice(0, 16) })), + }; +} + +// ---- --self-test ---- +function selfTest() { + const checks = []; + const eq = (name, got, want) => checks.push({ name, ok: JSON.stringify(got) === JSON.stringify(want), got, want }); + eq("6.7.0 < 6.35.0 numerically", cmpVersion("6.7.0", "6.35.0"), -1); + eq("6.12.1 < 6.32.0", cmpVersion("6.12.1", "6.32.0"), -1); + eq("6.35.0 == 6.35.0", cmpVersion("6.35.0", "6.35.0"), 0); + eq("7.0.0 > 6.35.0", cmpVersion("7.0.0", "6.35.0"), 1); + eq("non-semver is null", cmpVersion("6.35", "6.35.0"), null); + const base = { command: "/pharn-loop", membership: { method: "run-window/2" }, coverage: "partial", executions: { rows: [] }, work: [] }; + const ok = (l, checker = "green", fixture = false) => classify(l, { fixture, checker }); + eq("6.7.0 is pre-optimization", ok({ ...base, skills_version: "6.7.0" }).reasons, ["pre-optimization-version"]); + eq("6.33.0 is model-usage tier, no elapsed/work", ok({ ...base, skills_version: "6.33.0" }), { + tier: "model-usage", + reasons: [], + metrics: { tokens: true, elapsed: false, work: false }, + metric_exclusions: { tokens: [], elapsed: ["version-below-6.35.0"], work: ["version-below-6.35.0"] }, + }); + eq("6.35.0 green is full, all metrics", ok({ ...base, skills_version: "6.35.0" }).metrics, { tokens: true, elapsed: true, work: true }); + eq("a fixture is never included", ok({ ...base, skills_version: "6.35.0" }, "green", true).tier, null); + eq("a crashed checker blocks tokens, is no RED, and is named", ok({ ...base, skills_version: "6.35.0" }, "crashed"), { + tier: "full", + reasons: [], + metrics: { tokens: false, elapsed: true, work: true }, + metric_exclusions: { tokens: ["checker-crashed"], elapsed: [], work: [] }, + }); + eq("absent membership is named per metric", ok({ ...base, skills_version: "6.35.0", membership: undefined }).metric_exclusions.tokens, [ + "membership-absent", + ]); + eq("run-window/1 is a rule reason", ok({ ...base, skills_version: "6.35.0", membership: { method: "run-window/1" } }).reasons, [ + "membership-run-window-1", + ]); + eq("wrong command", ok({ ...base, command: "/pharn-review", skills_version: "6.35.0" }).tier, null); + eq("null parses but is no ledger", isLedgerObject(null), false); + eq("an array is no ledger", isLedgerObject([]), false); + eq("orchestrator in the main thread", requestRole({ sidechain: false, agent_id: null }, "main"), "orchestrator"); + eq("orchestrator inside an agent (REVIEW F6)", requestRole({ sidechain: true, agent_id: "a1" }, "agent:a1"), "orchestrator"); + eq( + "the main thread is NOT the orchestrator of a run started in an agent", + requestRole({ sidechain: false, agent_id: null }, "agent:a1"), + "stage-agent" + ); + eq("a stage agent", requestRole({ sidechain: true, agent_id: "b2" }, "agent:a1"), "stage-agent"); + eq("no run context", requestRole({ sidechain: false }, null), "unknown-context"); + eq( + "storedProfile splits by role", + storedProfile({ + membership: { context: "agent:a1" }, + requests: [ + { stage: "pharn-build", sidechain: true, agent_id: "a1", tokens: { output: 1 } }, + { stage: "pharn-build", sidechain: true, agent_id: "b2", tokens: { output: 2 } }, + ], + }).by_stage_role, + { + "pharn-build · orchestrator": { requests: 1, tokens: { output: 1 } }, + "pharn-build · stage-agent": { requests: 1, tokens: { output: 2 } }, + } + ); + eq("bash blocks counted per section", bashBlocksBySection("## A\n```bash\nx\n```\n```text\ny\n```\n## B\n```bash\nz\n```\n"), { + A: 1, + B: 1, + }); + eq("section bytes", sectionBytes("## 5. a\nx\n## 6. b\nyy\n## 7. c\n", 6), 11); + eq("stats lower median", stats([3, 1, 2, 4]), { n: 4, min: 1, median: 2, max: 4 }); + eq("home is redacted", redact(`${HOME}/x`), "~/x"); + const failed = checks.filter((c) => !c.ok); + return { ok: failed.length === 0, checks: checks.length, failed }; +} + +function main(argv) { + const usage = () => { + process.stderr.write( + "usage: audit.mjs --self-test | --ledgers [--fixture ] | --static [--project ] | --prefix [--since ] [--until ]\n" + ); + return 2; + }; + const take = (flag) => { + const i = argv.indexOf(flag); + if (i < 0) return null; + const out = []; + for (let k = i + 1; k < argv.length && !argv[k].startsWith("--"); k++) out.push(argv[k]); + return out; + }; + const one = (flag) => { + const v = take(flag); + return v === null ? undefined : v.length === 1 ? v[0] : null; + }; + let doc; + if (argv.includes("--self-test")) doc = selfTest(); + else if (argv.includes("--ledgers")) + doc = ledgersMode( + (take("--ledgers") ?? []).map((p) => resolve(p)), + (take("--fixture") ?? []).map((p) => resolve(p)) + ); + else if (argv.includes("--static")) { + const project = one("--project"); + if (project === null) return usage(); + doc = staticMode(project === undefined ? undefined : resolve(project)); + } else if (argv.includes("--prefix")) { + const since = one("--since"); + const until = one("--until"); + if (since === null || until === null) return usage(); + const files = (take("--prefix") ?? []).map((p) => resolve(p)); + doc = prefixMode(files, since, until); + } else return usage(); + process.stdout.write(`${JSON.stringify(doc, null, 2)}\n`); + return doc.ok === false ? 1 : 0; +} + +if (import.meta.main) process.exitCode = main(process.argv.slice(2)); diff --git a/.dev/features/pipeline-performance-audit/regression-report.json b/.dev/features/pipeline-performance-audit/regression-report.json new file mode 100644 index 00000000..aa58f00c --- /dev/null +++ b/.dev/features/pipeline-performance-audit/regression-report.json @@ -0,0 +1,33 @@ +{ + "base": "c9737b4486bc47759bd36f43f3430cf86bc068b4", + "inside": [ + ".dev/features/pipeline-performance-audit/GRILL.md", + ".dev/features/pipeline-performance-audit/PLAN.md", + ".dev/features/pipeline-performance-audit/REGRESSION.md", + ".dev/features/pipeline-performance-audit/REVIEW.md", + ".dev/features/pipeline-performance-audit/SHIP.md", + ".dev/features/pipeline-performance-audit/VERIFY.md", + ".dev/features/pipeline-performance-audit/audit.mjs", + ".dev/features/pipeline-performance-audit/regression-report.json", + ".dev/features/pipeline-performance-audit/verify-report.json", + ".dev/measurements/pipeline-performance-audit-2026-09-29.md", + "CHANGELOG.md" + ], + "outside_gates": { + "structural:pharn/pharn-review/trust-fence/evals/expected/expected-injection-comment.json": { + "base": 0, + "head": 0 + }, + "tests": { + "base": 0, + "head": 0 + }, + "validate": { + "base": 0, + "head": 0 + } + }, + "regressions": [], + "pre_existing": [], + "verdict": "no-regressions" +} diff --git a/.dev/features/pipeline-performance-audit/verify-report.json b/.dev/features/pipeline-performance-audit/verify-report.json new file mode 100644 index 00000000..236d30f9 --- /dev/null +++ b/.dev/features/pipeline-performance-audit/verify-report.json @@ -0,0 +1,15 @@ +{ + "feature": "pipeline-performance-audit", + "gates": { + "format:check": 0, + "lint": 0, + "lint:md": 0, + "reconcile": 0, + "structural:pharn/pharn-review/trust-fence/evals/expected/expected-injection-comment.json": 0, + "test": 0, + "validate": 0 + }, + "verdict": "PASS", + "failing_gates": [], + "verifiers": { "registered": 0, "findings": [] } +} diff --git a/.dev/measurements/pipeline-performance-audit-2026-09-29.md b/.dev/measurements/pipeline-performance-audit-2026-09-29.md new file mode 100644 index 00000000..baf6ea15 --- /dev/null +++ b/.dev/measurements/pipeline-performance-audit-2026-09-29.md @@ -0,0 +1,610 @@ +# Post-optimization performance audit of the PHARN delivery pipeline (2026-09-29) + +An analysis-only audit of the product delivery pipeline (`/pharn-ship`, `/pharn-loop`) as it stands at +`SKILLS_VERSION` 6.35.0, commit `c9737b4`. It follows the four optimization increments: 6.32.0 (#294), 6.33.0 (#297), +6.34.0 (#298) and 6.35.0 (#299). + +- **Plan and selection rule:** `.dev/features/pipeline-performance-audit/PLAN.md`, approved at GATE 1 on 2026-09-29 + (route "audit now"). +- **Grill:** `GRILL.md` in the same folder. +- **Review:** `REVIEW.md`. This version applies the review's findings F1–F10 and the re-review's R1–R3. + +**Where the numbers come from.** Figures come from `.dev/features/pipeline-performance-audit/audit.mjs`, and the mode +is named with each one. There are three exceptions: + +- CF1/CF2 are quoted from their `MEASUREMENT.md` files; +- pharn-starter's install date is read from its `pharn.config.json` (`installedAt`); +- the pre-routing request counts (11–426, median about 200) are quoted from the CHANGELOG [6.32.0] entry. + +A derived figure shows its formula. + +## The answer, in one paragraph + +**No real delivery run exists on a post-optimization version.** All 69 real cost ledgers found were written by pharn +6.12.1 (56) or 6.7.0 (13). pharn-starter, the one project with real runs, still records `skillsVersion` 6.12.1, so none +of the four optimizations has run on real work. So this audit cannot say what consumes most measured model usage or +most observed wall-clock time. Section 9 names the smallest step that answers both. + +What the code and the controlled observations do show: + +- **Most PINNED orchestrator calls are deterministic.** A green one-iteration `/pharn-loop` makes at least 70 + orchestrator tool calls that are pinned or mandated (estimate). 57 of them run a pinned script and read its exit code + or one short line. The pinned lines are sequential, because each branches on the exit code of the one before it. So + nearly every call is its own model request carrying the orchestrator's whole context. The share of ALL orchestrator + requests these make in a real run is **not** known: unpinned calls are not counted. +- **Every fresh stage agent pays a large fixed prefix on its first request.** Stage agents are spawned as + `general-purpose`. In 29 such agents in this repo, that first request cache-wrote 68,223–129,646 tokens (median + 91,759), and 15 of the 29 also cache-read about 31k tokens. +- **The project's `test` gate runs at least twice per iteration**, and three or four times when any test file lies + outside the feature: + - inside the build agent, repeated until green, with its output in that agent's context; + - at regress, HEAD and BASE (unless reused), over the outside-feature test files only. It is recorded `no-files`, + and not run, when there are none; + - at verify, where it is never reused. + +Which of these dominates a real run is **unknown**. The best-evidenced structural candidates cut pinned orchestrator +calls (section 6, C1 and C2). Both are contingent on real runs showing that orchestrator requests are a material +share. + +## Labels used throughout + +Each figure carries a **source** and a **precision**, never one mixed label (GRILL G5). + +**Source:** + +- `R` — real, post-optimization user workload. There is none. +- `H` — real, pre-optimization. It is used only to report exclusions, never for a conclusion. +- `CF` — a controlled fixture. It names its fixture and floor. +- `CH` — a controlled harness observation: Claude Code's own transcripts in this repo, not a PHARN delivery run. +- `S` — static: this checkout at `c9737b4`. + +**Precision:** `m` measured; `e` an estimate derived from measured numbers, with its formula. + +So `[S·m]` is a byte count read from a file, and `[S·e]` is a figure derived from such counts. + +- **Never compared across sources.** Figures of different sources are never expressed as percentages or ratios of + each other. +- **Classes kept separate.** Token classes (uncached input, cache write, cache read, output, thinking output) are + reported separately. The one all-classes total the helper prints is reference only, and no conclusion uses it. +- **No prices.** No figure is priced. + +## 1. Evidence set + +### Real runs: found, included, excluded + +`audit.mjs --ledgers ~/Projects/pharn-starter/pharn/features/*/cost.json`. All values `[H·m]`: + +| measure | value | +| ----------------------------------------------------- | ------------------------------------------------------------------------------------- | +| ledgers found (the rule's universe) | **69** | +| included, full profile (≥ 6.35.0) | **0** | +| included, model usage only (6.32.0–6.34.x) | **0** | +| denominators: tokens / elapsed / work | **0 / 0 / 0** | +| `skills_version` | 6.12.1 ×56, 6.7.0 ×13 | +| schema · membership | `pharn-cost-ledger/2` · `run-window/1` ×56; `pharn-cost-ledger/1` · no membership ×13 | +| `coverage` | `partial` ×69 | +| with `executions` or `work` (the 6.35.0 keys) | 0 | +| `check-cost-ledger.mjs` today | GREEN ×69 (internal consistency only, L43) | +| checker WARN: the rows come from two or more contexts | 49 (all 6.12.1) | +| ledgers sharing one `window_start` | 10 at 2026-09-22T17:06:51Z, 2 at 2026-09-23T08:56:20Z, 2 at 2026-09-25T20:39:38Z | + +**Exclusions, by the frozen rule's closed reasons:** + +- `pre-optimization-version` ×69; +- `membership-run-window-1` ×56. + +No other reason fired, and no ledger was eligible, so no per-metric exclusion applies. The 13 schema-`/1` ledgers have +no `membership` at all, so they fail the rule's `run-window/2` condition. The frozen reason set has no member for that +(GRILL G2). It is recorded here, not patched, because every one of them is already out on version. + +**Two independent signals agree that none is post-optimization:** + +- **The version field.** It is advisory under the cost-ledger contract, because it records the configured version + (GRILL G1). +- **The structure.** No ledger has the 6.35.0 `executions`/`work` keys, and none has `run-window/2` membership + (6.29.0). + +**The ledgers are also weak on their own terms:** + +- **Multi-context rows.** The checker's WARN on 49 `run-window/1` ledgers says their rows come from two or more + contexts. A run's own spawned agents are contexts too, so this shows membership that `run-window/1` could not scope + to the run. It does not prove that foreign rows are present. +- **Output under-counted.** Before 6.24.1, output was read from a request's first transcript line (L63). + +**Other real run records, not ledgers (`H`, excluded):** + +- this repo's `pharn/features/loop-decision-integrity/`, a 6.3.0-era `LOOP.md` and its reports; +- one `/pharn-loop` transcript in this repo's `pharn-loop-run` worktree project, dated 2026-09-21. The PLAN's inventory + named it. It is excluded as pre-optimization by date: 6.32.0 was released on 2026-09-28; +- the pharn-starter orchestrator transcripts. pharn-starter's install record reads 6.12.1, installed 2026-09-23, so + every one of them predates the four optimizations; +- the "37 orchestrator transcripts on the maintainer's machine" that the CHANGELOG [6.32.0] entry measured. That + increment's plan (`.dev/features/orchestrator-context/PLAN.md`) says those transcripts "predate 6.27.0 routing". + +The rule's universe named ledgers only. Leaving transcripts out is a scope choice recorded here (GRILL G7). None of +them could enter anyway. + +**Fixtures.** No fixture delivery run was made: GATE 1 chose "audit now", so `fixture-not-workload` never fired. + +**Limits of the rule, disclosed (GRILL G1–G3).** + +- **The freeze and the "fixture" call.** The freeze is advisory: no floor pin held `PLAN.md` between GATE 1 and the + build. `fixture-not-workload` is the caller's declaration (`--fixture`), because nothing in a `cost.json` records + it. +- **The sample above 20** is not sized, and "mode" is not a top-level ledger field. The human should settle both + before the first real-run pass. +- **What the planner saw before freezing the rule.** The planner read one pre-optimization ledger's totals and the + version and schema fields of all 69. No post-optimization cost existed to bias the rule. + +### Controlled evidence (used as controlled, never as workload, L4) + +| id | source | what it measures | reps | +| --- | ----------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------- | +| CF1 | `.dev/features/regress-base-reuse/MEASUREMENT.md` | BASE reuse on a 2nd regress. Floors `a2b5f6b` (before) vs its branch (after). Fixture: 3 gates (0.8–1 s) + a 1.5 s `postinstall` | 3 per variant | +| CF2 | `.dev/features/verify-head-gate-reuse/MEASUREMENT.md` | VERIFY reuse of REGRESS/HEAD executions. Floors `2cf0e85` vs its branch. Fixture: 4 gates (0.3–1 s), `--no-install` | 3 per variant | +| CH1 | `audit.mjs --prefix`, input set below | the FIRST request of each subagent in this repo from 2026-09-23 to 2026-09-29T07:00. It is grouped by the agent type the transcript's `.meta.json` records, per token class | 37 agents | +| — | `.dev/features/run-performance-breakdown/DEMO.md` | **synthetic** token counts. Used for nothing in this audit | — | + +- **CH1's input set.** Every `subagents/agent-*.jsonl` under this checkout's Claude Code project directory, run as: + + ```text + find ~/.claude/projects/ -path '*/subagents/agent-*.jsonl' \ + -exec node .dev/features/pipeline-performance-audit/audit.mjs --since 2026-09-23 --until 2026-09-29T07:00 --prefix {} + + ``` + + The window is fixed, so the set does not move as later agents are spawned. It holds 29 `general-purpose`, 6 + `Explore` and 2 `claude-code-guide` agents. PHARN spawns stage agents as `general-purpose`, so only that group + describes them. + +- **CF1 and CF2 are not pooled.** Both ran on floors older than 6.35.0, each with its own fixture. + +### Static evidence + +`audit.mjs --static --project ~/Projects/pharn-starter`, at commit `c9737b4`. It covers: + +- the bytes of every orchestrator command and part, every stage command and every stage-agent brief (the build brief + at iterations 1 and 2); +- the pinned shell blocks per command section, and the size of `ARCHITECTURE.md §6`; +- pharn-starter's `CLAUDE.md`, and the route its current `models` block would get. + +## 2. Current execution profile (as the code stands, not as measured) + +**Routing depends on the install's config (GRILL G8).** + +- **This repo's `models.stages`:** + - spec, plan, grill and `ac-test` (the test stage) request opus; + - build, regress, verify, ship and loop request sonnet; + - effort is `high` everywhere. +- **pharn-starter's current block** (the pre-0.7.0 shape, `opus-4-8`/`sonnet-5`, top-level `default`) makes every + routed cell print `inline:config-red` [S·m]. Updated to 6.35.0 **without** a CLI ≥ 0.7.0 migrating that block, it + would run every stage inline on the orchestrator's model. Stage-agent routing, and the context separation below, + would then not happen. +- **Effort is never routed.** A stage agent inherits the parent session's effort (`stage-agent-core.mjs`, header). + +Byte figures in the table are `[S·m]`. + +| stage | model or deterministic | where it runs (full / quick) | requested model (this repo) | major inputs | deterministic work | repeats when | +| ------------------------------------------- | ------------------------------------------------------------------------------------------- | ---------------------------------------------------------- | --------------------------------- | ---------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------- | +| orchestrator (`/pharn-loop`, `/pharn-ship`) | model (sequencing, the stuck-point mapping, ship's GATE-2 text); mostly runs pinned scripts | the invoking session, inline | frontmatter `sonnet` for the turn | its command (43,164 / 35,982 B), `CONSTITUTION.md`, parts at their points, every tool result | markers, route/read lines, freshness, stop decision, ledger/report/commit (loop) | the whole run | +| SPEC | model | loop: agent / agent; ship: inline (it IS GATE 1) | opus | description, template, the project | `check-spec`, pin, `check-spec-approved` | once per run | +| PLAN | model | agent / agent | opus | SPEC, lessons (index then canon), the §6 plan row, the project | `check-plan-lessons`, AC-TESTS mapping | once per run | +| GRILL | model (full), checkers only (quick) | agent / inline `floor-only` | opus (full) | PLAN, SPEC, lessons, contracts, the §6 grill row | `check-plan-spec-agree`, `check-plan-lessons` | once per run | +| TEST | model + deterministic red run | agent / agent | opus (`ac-test` key) | SPEC, PLAN, AC-TESTS.md | `run-gates --stage ac-test`, `check-red-run`, the lock | once per run (never re-run after build) | +| BUILD | model + the project's own gate | agent / agent | sonnet | PLAN.md only (not SPEC, not GRILL), the §6 build row, the project; iter ≥ 2 (loop): report fix-list fields | `check-test-stage`, scope set, anchor, **the user's `test`/`lint`, re-run until green** (un-stamped, un-budgeted) | every iteration (loop cap default 3); ship: one bounded rebuild on `INCOMPLETE` (Step 2b) | +| REGRESS | deterministic (`stage-regress.mjs`) | inline `floor-only` / skipped | (orchestrator's) | PLAN `## Files`, git | HEAD gates (outside-feature `test` subset + other discovered gates); BASE worktree + install + base gates unless the BASE requirement is unchanged within the run (6.33.0) | every iteration; + ≤ 1 freshness re-run per iteration; + a `continue` resume per 570 s budget | +| VERIFY | deterministic (`stage-verify.mjs`) | inline `floor-only` / inline | (orchestrator's) | stamps, AC evidence | every discovered gate at HEAD (typecheck/build reused from REGRESS/HEAD when eligible, 6.34.0; `test`, style gates, `reconcile` never) + AC gate + completeness | as regress | +| retries | the stop is deterministic (`check-loop.mjs`); the fix is model work | — | — | loop: structured fix-list fields only | `check-loop-fresh.mjs` before every stop read | `CONTINUE` → a new build agent + regress + verify; `RERUN` → the named stage again, same iteration | +| closeout | mostly deterministic; `LOOP.md` Handoff is model text | inline, from the close part (31,211 / 28,179 B, read once) | (orchestrator's) | reports, markers | loop green path: record checks, run-stop marker, ledger render + check, run report, commit-gate freshness, stage list, branch, add, commit, release | once | + +**What one routed stage costs the orchestrator** [S·m]: + +- four pinned lines: `route`, the stage-start marker, `read`, and the return marker; +- one Agent call. + +**What a routed stage agent receives:** + +- first, its harness prefix, on its first request (CH1); +- then, through its own tool calls: + - its brief, 2,498–4,063 B [S·m]. The upper end is a loop rebuild at iteration ≥ 2; a first loop build's brief is + 3,598 B; + - `CONSTITUTION.md`, 7,958 B [S·m]; + - its stage command, 19,125–26,261 B [S·m]; + - for plan, grill and build, the `ARCHITECTURE.md §6` row its command names (all of §6 is 4,754 B [S·m]); + - whatever else the stage reads and runs. + +## 3. Measured cost profile + +### Model usage + +- **Real:** no measurement (0 runs, all three denominators 0). +- **CH1, a fresh `general-purpose` agent's first request** [CH·m], 29 agents, by token class: + + | token class | min | median | max | note | + | -------------- | ------ | ------ | ------- | ------------------------ | + | uncached input | 2 | 2 | 2 | | + | cache write | 68,223 | 91,759 | 129,646 | | + | cache read | 0 | 30,927 | 31,428 | non-zero in 15 of the 29 | + + The other agent types differ: `Explore` cache-wrote 34,934–46,732, and `claude-code-guide` (haiku) 72,869–73,357. + They are not stage agents. The prefix's composition (system prompt, tool schemas, `CLAUDE.md`, memory) is **not** + measured. + + This repo's `CLAUDE.md` is 172,421 B [S·m], and pharn-starter's is 387,549 B [S·m], 2.25× larger. The prefix + there is **not measured**. Whether `CLAUDE.md` drives the prefix cannot be told apart from the rival causes with this + data (L40). + +- **PHARN-owned instruction bytes per routed stage agent** [S·m]: + - brief + `CONSTITUTION.md` + stage command is 29,581 B (test, ship) to 37,485 B (spec, quick loop); + - plan, grill and build add up to 4,754 B of §6. + + These enter through the agent's tool results after its first request, so they are not part of CH1's figure. No + ratio of the two is stated (different sources). + +- **Orchestrator instruction bytes** [S·m]: + - `pharn-loop.md` 43,164 and `pharn-ship.md` 35,982 at invocation, carried by every later orchestrator request; + - quick parts 11,091 / 11,223, only on a `--quick` run; + - close parts 31,211 / 28,179, read once at the stop; + - the inline thin callers `pharn-regress.md` 19,848 and `pharn-verify.md` 17,929. How the orchestrator loads these + (a slash-command invocation per call, or one Read) is not pinned and not measured (section 9). + + The orchestrator's harness prefix is not measured: CH1 covers subagents only. + +### Observed elapsed time + +- **Real:** no measurement. The 6.35.0 `executions` view exists in no real ledger. +- **CF1 / CF2 whole-stage wall-clock**, for their fixtures only [CF·m]: + - regress 10.1–10.6 s → 4.4 s on a BASE-reuse hit; + - verify 4.8–4.9 s → 2.5 s with 2 of 4 gates reused. + + They illustrate what the counts save on those shapes. They say nothing about a real project. + +### Deterministic work + +- **Real:** no `work[]` record exists. +- **CF1** [CF·m]. On the 2nd regress of one run, per invocation: + - worktree checkouts 1 → 0, installs 1 → 0, base gate processes 3 → 0; + - all gate processes 6 → 3. + + The 1st invocation is unchanged. + +- **CF2** [CF·m]. Per regress + verify pair: + - verify gate processes 4 → 2 (`typecheck`, `build` reused); + - total gate processes 12 → 10. + + `test` runs at both stages by design, and `lint` runs at both because style gates are never reused. + +### Retries + +- **Real:** no measurement. Frequency of `CONTINUE`, freshness `RERUN`, `continue` resumes and ship's Step 2b rebuild: + unknown. + +## 4. Dominant remaining costs + +**Measured model usage: cannot be ranked (0 real runs).** **Observed wall-clock: cannot be ranked (0 real runs).** + +What can be ranked is **static structure**, and it is ranked as structure only: + +1. **Pinned orchestrator calls** [S·e]. The inputs are pinned shell blocks in the loop family [S·m]. In + `pharn-loop.md`: + - Step 1a: 5; + - Step 3: 5; + - Step 4: 14, of which a green run executes 13, because `check-red-run --preflight` runs only when + `check-test-stage` fails; + - Step 5: 10 per iteration. + + The green-path close part adds 16 (6b 8, 6c 5, Step 7 1, Final 2). The two inline stages, `/pharn-regress` and + `/pharn-verify`, each cost 6 calls per iteration, because the loop pins them without the feature name: + - the invocation; + - the slug Write and the `feature-name.mjs` line; + - the scope set, the run and the release. + + The Agent calls are 5 (spec, plan, grill, test, build). The model file operations are 4: the slug Write, the + `CONSTITUTION.md` Read, the close-part Read and the `LOOP.md` Write. + + **A green one-iteration full loop:** 5 + 5 + 13 + 10 + 16 + (2 × 6) + 5 + 4 = **70 calls** [S·e]. + - **57 of them** run a pinned deterministic line and read its exit code or one short line: the 49 pinned blocks + plus 4 per inline stage. + - **The other 13** are the Agent calls, the invocations, the slug Writes and the file Reads and Writes. + + **Each further iteration:** 10 + 12 + 1 = **23 calls** [S·e], 18 of them deterministic. + + **What is NOT counted:** + - grill's verdict reads, which the loop cites from ship and does not pin; + - further mandated green-path reads and lines: Step 1b's lookup of a prior `LOOP.md`, Step 6b's Read of the + loop-record contract, Step 6c's post-commit `git rev-parse HEAD`, and Step 7's reads of `RUN-REPORT.md` and the + pre-run snapshot. With these the count is about 75 [S·e]; + - `continue` resumes, freshness re-runs, and any extra reads the model makes. + + So 70 is a **lower bound** on the pinned or mandated calls. It is not the run's request count. For scale, the + 6.32.0 entry measured pre-routing loop runs at 11–426 requests, median about 200 `[H·m]`. It describes a different + pipeline and is not a denominator here. + + The pinned lines run one per turn, since each branches on the previous exit code. Each is therefore one model + request that re-sends the orchestrator's whole context: its prefix, the 43 KB command, and every tool result and + stage-agent reply so far. Independent calls issued in one turn would share a request (the loop command says so of + its two entry Reads). + +2. **Fresh-context prefixes** [CH·m, S·m]. The number of stage agents spawned per run [S·m]: + + | run | stage agents | + | ------------------------------- | ------------------------------- | + | `/pharn-loop` full, 1 iteration | 5, plus 1 per further iteration | + | `/pharn-ship` full | 4 | + | `/pharn-loop --quick` | 4 | + | `/pharn-ship --quick` | 3 | + + Each `general-purpose` first request cache-wrote a median of 91,759 tokens in CH1. For 5 agents that is about + 459k cache-write tokens [CH·e] (5 × 91,759). Where the shared part hits, add about 155k cache-read tokens [CH·e] + (5 × the median hit 30,927; at most 157k, 5 × 31,428). This is before any stage work, and it is this repo's + harness, not pharn-starter's. + +3. **Repeated `test` executions** [S·m]. Per iteration, `test` runs: + - inside the build agent, repeated until green, with its output in the agent's context; + - at regress HEAD, as the outside-feature subset; + - at regress BASE, unless reused; + - at verify, the whole suite; it is never reused, because it is an AC level gate. + + Regress records `test` as `no-files`, and does not run it, when no test file lies outside the feature. So the floor + is two runs per iteration (build and verify), and three or four when such a file exists. + + `typecheck`/`build` run at regress HEAD and BASE, and at verify when not reused. Real durations: unmeasured. No + per-gate duration is recorded anywhere (`gate-process-duration`). + +## 5. Root-cause analysis + +- **Pinned orchestrator calls (mechanism confirmed, magnitude unknown).** + - The mechanism: each pinned line is its own Bash call by design (`pharn-loop.md`: "Each fenced block runs as its + own shell and carries no state into the next"). Each call returns to the model, which issues the next request with + the full context. + - The design reason: it pins every deterministic step as a separate, testable, exit-code-branched line. That serves + the P5 membership branching and L44, since a multi-block procedure must not carry shell state. + - What is not known: whether this is a material share of a real run. The rival explanation is that stage agents' + own requests dwarf it. + + Real ledgers settle it. Each stored row names its context (`sidechain` + `agent_id`), and the run's own context is + `membership.context`. That may be `main`, or `agent:` when the run itself started inside an agent. So + `audit.mjs` splits requests per stage into `orchestrator` (the run's own context) and `stage-agent` (any other + admitted context) (REVIEW F6). + +- **Fresh-context prefix (mechanism confirmed in this repo, composition unknown).** + - The mechanism: stage-agent routing (6.27.0) gives each routed stage a new context, which is the only way to + request a stage's configured model. In CH1 the first request is mostly cache write, and cache sharing across agents + is partial (about 31k in 15 of 29). + - The cause of the prefix size is **not** separated. Candidates are `CLAUDE.md` (larger in pharn-starter), the tool + schemas and MCP instructions of the session, and the skills listing. CH1's 68k–130k spread across one week in one + repo points at session configuration as much as at `CLAUDE.md` (L40). +- **Repeated `test` (mechanism confirmed, magnitude unknown).** + - The build agent's own gate is `/pharn-build` Step 4, "the user's `test` / `lint`", fix within scope until green. + - Regress HEAD runs an outside-feature subset. It is a different execution from verify's whole suite, so 6.34.0 + cannot reuse it. + - Verify never reuses an AC level gate (`NON_REUSABLE_IDS`). + + Real suite duration, and how many fix-and-rerun cycles a build makes, are unknown. + +- **Artifacts are not a demonstrated cause.** + - The loop's rebuild reads only the reports' structured fix-list fields (brief rule 7). + - `/pharn-build` reads `PLAN.md` only. + - The closeout cites `GRILL.md`/`REGRESSION.md`/`VERIFY.md` by pointer. + - Verdicts are read from JSON `.verdict` by checkers. + + No case was found of a large human-readable artifact re-read by a model where a structured subset would do. Artifact + sizes in post-optimization runs are unmeasured (0 runs). + +### Stage boundaries (what crosses, what is re-acquired, what the separation is for) + +| boundary | crosses (on disk) | receiver re-acquires | semantic role today | risk if merged (quality effect: **unknown**, no comparative evidence) | +| ----------------------------- | ------------------------------------- | ------------------------------------------------------------ | ------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------- | +| spec → plan | `SPEC.md` (pinned hash) | prefix, constitution, command, SPEC, lessons, repo discovery | model routing; a plan written against a pinned, approved intent | the planner inherits the spec writer's framing (anchoring); larger persistent context | +| plan → grill | `PLAN.md`, `AC-TESTS.md` | the same + contracts | **independent review**: the grill treats the PLAN as `trust: untrusted` (P2) and did not write it | loss of independence; confirmation bias; a trust boundary becomes self-review | +| grill → test | none the test reads from grill | SPEC, PLAN, AC-TESTS | tests written before, and apart from, the code (the red run proves each can fail) | tests shaped by the implementer's intent; weaker red-run meaning | +| test → build | the lock, the pinned tests | PLAN only + repo | the builder cannot rewrite the tests (the pinned tests sit outside PLAN `## Files`) | the builder sees the tests' authoring context; stale assumptions carried across | +| build iter N → build iter N+1 | reports' fix-list fields (structured) | everything, in a new agent | a fresh context per fix attempt; failure data arrives as quoted DATA | resuming one agent keeps its earlier reasoning (anchoring on a failed approach) and grows its context | + +Every boundary re-pays the fixed prefix (CH1) and re-discovers the repo. How many requests that re-acquisition costs +is not measured. + +## 6. Candidate optimizations + +Each candidate cites the section above that motivates it. **None rests on real evidence**, so every one is +_contingent on real-run confirmation_. The confirming measurement is named with each. + +### C1 — Fold each routed stage's four pinned lines into two (structural, low semantic risk) + +- **Evidence.** §4.1: each routed stage costs the orchestrator 4 pinned lines + 1 Agent call [S·m]. A one-iteration + loop has 5 routed stages, so 20 of its 70 counted calls are this ceremony [S·e]. +- **Mechanism.** Pair `route` with its stage-start marker, and `read` with the return marker, as one tested CLI call + each. Stage boundaries, markers, routes and exit mapping all stay as they are. +- **Expected benefit.** + - Requests: −10 per one-iteration loop (5 × 2), −2 per further iteration [S·e]. + - Tokens: each removed request's full read of the orchestrator's context, mostly as cache read. That context's size + in a post-optimization run is not measured. Its PHARN-owned floor is the 43,164 B command (about 10.8k tokens at + 4 B per token [S·e], no class). Output savings are small. + - Wall-clock: one model round trip per removed call (unmeasured). +- **Quality / correctness risk.** + - The marker timing must stay exactly where attribution and `executions` need it: start before the Agent call, + return after `read`. + - Two exit codes merge into one call, so route exit 3 (inline) versus a marker refusal must stay distinguishable. + - The hygiene pins (`STAGE_AGENT_WIRING`, `PHASE_MARKER_WIRING`) move. +- **Scope.** Medium: `stage-agent.mjs`/`mark-phase.mjs` (or a small wrapper), `pharn-loop.md`, `pharn-ship.md`, the + hygiene tests. +- **Reversibility.** High; it is one revert, since the ledger format is unchanged. +- **Validation.** On one fixed task and base, run current vs candidate sequentially. Compare: + - `orchestrator`-role requests per stage (`by_stage_role`); + - identical `executions` rows (same stages, same run numbering); + - identical verdicts. +- **Confirm first.** In real ledgers, `orchestrator`-role requests (the run's own context, not "the main thread") are + at least 20% of a run's requests. That threshold is pre-registered here. + +### C2 — A closeout script for the green path (structural, low–medium semantic risk) + +- **Evidence.** §4.1: the loop's green close runs 16 pinned blocks [S·m]. Two of them must precede the model's + `LOOP.md` Write: the one block holding the `LOOP.md` scope set and its amend, and Step 6b's `git rev-parse HEAD` + commit capture. The other 14 are a deterministic tail. The named follow-up `ship-closeout-script` already exists. +- **Mechanism.** After the model writes `LOOP.md`, one tested script runs the ordered tail: + - the record checks; + - the run-stop marker; + - the ledger render and check, and the run report; + - the commit-gate freshness check; + - the scope and amend steps; + - the stage list, branch, add and commit; + - the releases. + + It returns one closed outcome (`committed` or `not committed: `). + +- **Expected benefit.** Requests: 14 − k, where k (1–4) is the number of calls a design keeps separate. That is + **−10 to −13** per run [S·e], at the orchestrator's largest context, the end of the run. Tokens: the same + per-request logic as C1. +- **Risk.** + - The ordering is load-bearing: the ledger is emitted before the attestation (ship), and the decision is + re-derived before the commit. + - Every `not committed:` path must keep its own outcome. + - A script that commits is a larger blast radius than a line that commits. +- **Scope.** Medium: a new floor script + tests, the two close parts, the hygiene pins. +- **Reversibility.** High. +- **Validation.** On green, cap and blocked fixtures, compare the same commit, record, ledger and exit outcomes with + fewer `orchestrator`-role requests. +- **Confirm first.** As for C1. + +### C3 — Run the regress/verify stage scripts from the orchestrator directly (structural if the mechanism holds) + +- **Evidence.** §4.1: each inline stage costs 6 orchestrator calls per iteration, 2 of which (the slug Write and the + `feature-name.mjs` line) only re-derive a `` the orchestrator already holds [S·m]. The loop already owns the + stage-exit mapping (`pharn-loop.md` Step 2). **If** each invocation loads the command body, 37,777 B of thin-caller + text [S·m] also enters the orchestrator's context per iteration. +- **Mechanism.** Pin `stage-regress.mjs` / `stage-verify.mjs` and their scope lines in the orchestrator, with + `--feature ''`, instead of invoking the commands. +- **Expected benefit.** −6 calls per iteration (2 invocations, 2 slug Writes, 2 `feature-name.mjs` lines) [S·e]. + Also up to about 9.4k tokens [S·e] (37,777 B at 4 B per token, no class) of carried text per iteration, which is **unknown until the invocation form is seen + in a transcript**. +- **Risk.** + - A second copy of each pinned line (L35), which needs a parity pin. + - The thin callers' per-exit guidance (question relay, crash and timeout handling) must be carried or cited. +- **Scope.** Small–medium. +- **Reversibility.** High. +- **Validation.** Count `orchestrator`-role requests and cache-read per iteration, before and after. +- **Confirm first.** The tool-use names in a real orchestrator transcript. + +### C4 — Bound the build agent's own gate output (EXPERIMENTAL: changes what BUILD sees) + +- **Evidence.** §4.3/§5: the build agent runs the project's `test`/`lint` itself, repeatedly, with its output entering + its context [S·m]. The volume is unknown. +- **Mechanism.** Run the build's gate through the runner (logs to files). The agent then sees an exit code and the + failing test ids and titles (the per-test results record), and opens a log only on demand. +- **Expected benefit.** Tokens: unknown (it scales with the project's reporter verbosity and the fix-and-rerun count). + Wall-clock: none by itself. +- **Risk.** The builder loses diagnostic detail it may need, so iterations may rise. This is exactly the risk class + the prompt names. +- **Scope.** Medium. +- **Reversibility.** High. +- **Validation.** The counterfactual E1 below. + +### C5 — Resume the build agent across loop iterations (EXPERIMENTAL: independence and context) + +- **Evidence.** §4.2: each iteration spawns a fresh build agent, which pays the prefix (CH1) and rediscovers the repo. +- **Mechanism.** Iteration N+1 continues iteration N's agent, with the fix-list delivered as DATA. +- **Expected benefit.** One fewer first-request cache write per further iteration (CH1 median 91,759 tokens), plus the + rediscovery requests. Unknown. +- **Risk.** Anchoring on a failed approach; a larger persistent context; fix-list DATA arriving in a context that + "believes" it finished; a trust channel via the agent's own earlier output. +- **Scope.** Medium–large. The Agent tool's resume semantics are harness behaviour. +- **Reversibility.** High. +- **Validation.** E1 below. + +### C6 — Model or effort per stage (EXPERIMENTAL; no candidate yet) + +No stage's share is measured, so no stage is nominated (the prompt: never from a stage's name). Effort is not routed +at all today: stage agents inherit the parent's effort. So an effort experiment first needs effort routing +(`stage-agent-effort`, a named residual). When real ledgers exist, a stage becomes a candidate only if it holds a +material share of requests or output. Its experiment is E1 with the model or effort as the one varied factor. The +quality dimensions at risk are the stage's own: + +- plan/grill: findings a later external review surfaces; +- build: iterations and verify `FAIL`s; +- test: red-run validity. + +### Counterfactual experiment E1 (required before adopting C4, C5 or any C6) + +- **Design.** Choose 6 comparable pharn-starter tasks before running anything: 3 small fixes and 3 medium features, + sized by the PLAN `## Files` count of a dry plan. Run each task under A (current) and B (candidate), from the same + base commit, in separate worktrees, **sequentially** (one run at a time, so no two share a measurement window). Keep + the models and the config identical except for the one varied factor, and randomize A/B order per task. +- **Outcomes per pair:** + - completion (`STOP_GREEN`) and iterations; + - verify/regress verdicts and AC gate results; + - blocking findings from one fixed external review (`/pharn-review` with a fixed lens set, or the human) on the + final diff; + - tokens by class (`cost.json`) and elapsed (`executions`). +- **Reading.** Report the per-task pairs. Six pairs support "no visible regression" at most, never "non-inferior". + Adoption needs the human's call on those pairs. + +## 7. Recommended next increments + +The evidence supports **one prerequisite and two implementation increments**, in this order. It does not support a +fourth. + +1. **Prerequisite (no code): produce post-optimization evidence.** + - Run `pharn update` in pharn-starter with `@pharn-dev/pharn` ≥ 0.7.0. The CLI version is not optional: an older + CLI leaves the `models` block unmigrated, so every stage runs `inline:config-red` and routing is off (§2). + - Then run the normal workload. + - After at least 5 full runs, and ideally some quick ones, run `audit.mjs --ledgers` over them. That fills sections + 3–4 with real, denominated numbers, including `by_stage_role` (orchestrator versus stage agents per stage), + `executions` and `work`. + - This is the only step that answers the success criterion's first two questions. +2. **C1 — the routed-stage pinned-line fold.** It is the smallest, most mechanical and most reversible change, and it + preserves every boundary and verdict. Adopt it only if item 1 confirms that `orchestrator`-role requests are at + least 20% of run requests. Otherwise stop here. +3. **C2 — the green-path closeout script.** It has the same confirmation gate. It is the named follow-up + `ship-closeout-script`, extended to the loop. + +C3 waits on the one transcript observation it needs (section 9, gap 3). C4–C6 are experiments, gated on E1. + +## 8. Do not optimize yet + +| idea | why it waits | +| --------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| trim the stage command texts further | there is no measured share: no post-optimization run has measured a stage agent's requests. 6.28.2 and 6.32.0 already cut these texts. Insufficient evidence | +| cross-run BASE cache (roadmap 4.2) | there are no real install or base-gate durations. `work[].base.install.ms` now records them, so wait for it. How often 6.33.0 reuse hits in real runs is also unknown | +| reuse `test` between regress and verify | these are different executions (a subset versus the whole suite), and `test` is an AC level gate the AC gate must observe. There is no real duration | +| run regress and verify in parallel, or the BASE side during the build | verify's reuse needs regress's final offer, and `reconcile` must run last. The BASE requirement includes the build's partition. Real stage durations: unmeasured | +| merge stages (plan+grill, spec+plan, test+build) | the quality effect is unknown, and it removes independence and trust boundaries (§5). It needs E1 | +| cheaper model or lower effort for any stage | no stage's share is measured, and effort is not even routed yet | +| a state capsule or handoff summary between stages | re-acquisition cost is unmeasured, and it changes what each agent sees (experimental) | +| slim `SPEC.md`/`PLAN.md`/reports | no evidence that a model re-reads a large artifact where a structured subset would do (§5). Sizes are unmeasured | +| act on the 6.35.0 DEMO's numbers | its tokens are synthetic | +| act on the 69 pre-optimization ledgers | they describe a pipeline that no longer exists (no test stage, no routing, prose regress/verify), and their `run-window/1` membership is not run-scoped (L24) | + +## 9. Measurement gaps + +| gap | why it matters | smallest improvement | +| ------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| 1. no post-optimization real run | blocks every model-usage and elapsed conclusion | the section-7 prerequisite. No code | +| 2. the prefix composition of a fresh agent in the target project | decides whether C5-style consolidation could matter, and how much `CLAUDE.md` weighs | a controlled probe: spawn a one-line `general-purpose` subagent in pharn-starter, then again with a trimmed `CLAUDE.md` in a scratch copy, and read both first requests with `audit.mjs --prefix`. Analysis-only | +| 3. how the orchestrator invokes an inline stage (slash-command body per call, or one Read) | decides part of C3's value | read the tool-use names in one real 6.35.0 orchestrator transcript. Analysis-only | +| 4. per-gate durations | ranks the deterministic work; decides BASE caching and `test` reuse | the runner times each gate process into its stamp (`gate-process-duration`, a named residual). A measurement-only runtime change, when the first real `work[]` rows show regress or verify elapsed is material | +| 5. tool-result volume per request | the prompt's §12 (full test output in BUILD's context) | cost.json carries token classes per request, not what produced them. A transcript is perishable, so an analysis-only transcript pass soon after a run, attributing each cache-write jump to the preceding tool result, is the smallest step | +| 6. the orchestrator's own prefix and per-request context | sizes C1–C3's per-request saving | the same `--prefix` read over the orchestrator's transcript (its first request) in a real run, plus the stored `requests[]` rows of its context | +| 7. served effort | effort experiments cannot be read | none available: the platform does not report it | +| 8. the rule's own gaps (G1–G3) | the first real pass must be decidable | the human sizes the sample above 20, names a membership-absent reason, and states who declares a fixture. Then the rule is frozen again before that pass | + +## What this audit may claim (P0) + +- **Floor.** Nothing this audit concludes is a floor guarantee. The one floor fact near it is `check-cost-ledger.mjs`'s + GREEN on each ledger, which is internal consistency, never accuracy (L43). The build's own guard facts are the + dev chain's (writes-scope hook, `reconcile`), not this report's. +- **Measured.** + - byte counts, pinned-block counts, route tokens, and the ledgers' enum and version fields; + - CH1's first-request usage per class as the transcripts record it; + - CF1/CF2 as their files record them. +- **Advisory.** + - Every estimate: 70 calls (a lower bound; about 75), 57, 23 per iteration, about 459k, about 155k, −10, −10 to + −13, −6, about 9.4k, about 10.8k. + - Every candidate's expected benefit and risk. + - Every root-cause statement. + - The mapping from static structure to cost. +- **Struck.** + - "X is the most expensive stage"; + - "most orchestrator requests are ceremony" (only the PINNED calls are counted); + - "C1 saves N% of a run"; + - "merging stages does not hurt quality". + + None of these has evidence. diff --git a/CHANGELOG.md b/CHANGELOG.md index 9ba48fa2..baed446e 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -23,6 +23,29 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), `npm run check:changelog` holds this file's shape; the CI step "CHANGELOG per-PR entry check" holds each PR's diff. Details and known costs: CONTRIBUTING.md, "CHANGELOG entries". --> +### Added + +- 2026-09-29: **A post-optimization performance audit of the delivery pipeline, analysis only + (`.dev/measurements/pipeline-performance-audit-2026-09-29.md`).** Its headline is that no real run exists on a + post-optimization version: all 69 real cost ledgers were written by pharn 6.12.1 or 6.7.0, so the question of what + dominates measured model usage and wall-clock time stays open. From static and controlled evidence it finds three + structural costs: + - at least 70 pinned or mandated orchestrator tool calls in a green one-iteration `/pharn-loop`, 57 of them + deterministic. That is an estimate from the command text; how many requests a real run makes is not measured; + - a fresh `general-purpose` agent's first request cache-writing 68,223–129,646 tokens (median 91,759), over 29 + such subagents in this repo; + - the project's `test` gate running at least twice per iteration, and three or four times when any test file lies + outside the feature. + + It recommends one prerequisite and two contingent increments. The prerequisite is to update pharn-starter with a CLI + at 0.7.0 or later and collect real runs. The two increments fold the routed-stage pinned lines, and move the + green-path closeout into a script. Both wait for real ledgers to show that the run's own context (the orchestrator) + makes at least 20% of a run's requests. The build-output, agent-resume and model/effort candidates are experiments, + gated on a defined counterfactual. The audit also lists what not to optimize yet. + `.dev/features/pipeline-performance-audit/audit.mjs` is the read-only helper behind the report's figures; its + `--ledgers` mode applies the selection rule frozen at GATE 1. The review's ten advisory findings and the re-review's + three were fixed before merge. No product-surface byte changed, so there is no `SKILLS_VERSION` bump. + ## [6.35.0] - 2026-09-28 ### Added