Skip to content

docs(measurements): post-optimization performance audit of the delivery pipeline (apparatus) - #300

Merged
PrzemekGalarowicz merged 1 commit into
mainfrom
pipeline-performance-audit
Sep 29, 2026
Merged

PrzemekGalarowicz merged 1 commit into
mainfrom
pipeline-performance-audit

Conversation

@PrzemekGalarowicz

Copy link
Copy Markdown
Contributor

Summary

This is an analysis-only audit of the product delivery pipeline (/pharn-ship, /pharn-loop) at 6.35.0, after the four optimization increments (6.32.0 #294, 6.33.0 #297, 6.34.0 #298, 6.35.0 #299). No optimization is implemented. No product-surface byte changes, so there is no SKILLS_VERSION bump; the one [Unreleased] CHANGELOG entry is apparatus-only.

  • Report: .dev/measurements/pipeline-performance-audit-2026-09-29.md, in the nine sections the prompt asked for. Every figure carries a source label (real, pre-optimization, controlled fixture, controlled harness, static) and a precision label (measured, estimate). Token classes are never added together, and figures from different sources are never compared as ratios.
  • Helper: .dev/features/pipeline-performance-audit/audit.mjs, read-only. Its modes are --ledgers (the run-selection rule frozen at GATE 1), --static, --prefix and --self-test (25 checks).

Findings

  • No real post-optimization evidence. 0 of 69 real cost ledgers are post-optimization: 56 were written by 6.12.1 and 13 by 6.7.0. pharn-starter is still installed at 6.12.1. So what dominates measured model usage and wall-clock time cannot be ranked yet.
  • What the code and controlled observations show:
    • At least 70 pinned or mandated orchestrator calls in a green one-iteration /pharn-loop, 57 of them deterministic.
    • A fresh general-purpose stage agent's first request cache-writes 68k–130k tokens (median 91,759; 29 agents in this repo).
    • The project's test gate runs at least twice per iteration.
  • Recommended:
    • a prerequisite: pharn update pharn-starter with a CLI at 0.7.0 or later, then collect at least 5 real runs;
    • two contingent increments: C1 folds each routed stage's four pinned lines into two, and C2 moves the green-path closeout into one script. Both are gated on real ledgers showing the orchestrator's own context makes at least 20% of a run's requests.
  • Experiments only, each gated on a defined counterfactual (E1): bounding the build agent's gate output, resuming the build agent across iterations, and changing a stage's model or effort. A "do not optimize yet" list is included.

Chain (/pharn-dev-ship)

stage result
GATE 1 maintainer approved the plan, route "audit now"
grill check-plan-lessons exit 0; 15 advisory findings
build validate exit 0
regress no-regressions
verify PASS (7 gates, reconcile CLEAN)
review 0 floor-gate findings, 10 advisory
GATE 2 maintainer chose "Fix, then PR"
fix pass F1–F10 fixed; regress no-regressions; verify PASS
re-review F1–F10 confirmed fixed; 3 minor findings (R1–R3), also fixed

The R1–R3 fix changed prose only. After it, the style, CHANGELOG, validate, self-test and reconcile gates were re-run, but npm test was not; CI runs the full suite here. Details: .dev/features/pipeline-performance-audit/SHIP.md.

Test plan

  • node pharn/floor/validate.mjs . GREEN
  • npm test (full suite) passed in the fix-pass verify
  • npm run lint, format:check, lint:md, check:changelog, check:changelog-entry clean
  • node .dev/features/pipeline-performance-audit/audit.mjs --self-test passes 25/25
  • CI green on this PR

🤖 Generated with Claude Code

…ry pipeline (apparatus)

An analysis-only audit of /pharn-ship and /pharn-loop at 6.35.0, after the four
optimization increments (6.32.0-6.35.0). No product-surface byte changes, so no
SKILLS_VERSION bump.

- .dev/measurements/pipeline-performance-audit-2026-09-29.md: the report.
  0 of 69 real cost ledgers are post-optimization (all 6.12.1 / 6.7.0), so
  measured model usage and wall-clock cannot be ranked yet. From static and
  controlled evidence: at least 70 pinned orchestrator calls in a green
  one-iteration loop (57 deterministic), a 68k-130k-token first-request cache
  write per fresh stage agent, and the project's test gate at least twice per
  iteration. Recommends a prerequisite (update pharn-starter with CLI >= 0.7.0
  and collect real runs) and two contingent increments (C1, C2).
- .dev/features/pipeline-performance-audit/audit.mjs: the read-only helper
  behind the figures (--ledgers, --static, --prefix, --self-test).
- The /pharn-dev-ship stage artifacts (PLAN, GRILL, REGRESSION, VERIFY,
  REVIEW, SHIP) and one [Unreleased] CHANGELOG entry.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 29, 2026

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: 1091f4bd-d416-423e-b33b-7d0b4c642ff2


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@PrzemekGalarowicz
PrzemekGalarowicz merged commit d10fc1c into main Sep 29, 2026
8 checks passed
@PrzemekGalarowicz
PrzemekGalarowicz deleted the pipeline-performance-audit branch September 29, 2026 08:26
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant