Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
46 changes: 46 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,51 @@
# Changelog

## v2.5.2

Makes v2.5 work on a real run. Every fix has a test that fails when the fix is reverted.

Correction to v2.5.1: its note "the provenance test now checks one ported test per row" was wrong. The test
checks each ported file's header and notice, not the rows of `research/upstream.lock`, and that lock file
sits outside this repo, so no test here can read it. `THIRD_PARTY_NOTICES.md` plus the provenance test are
the in-repo record; the lock's `shipped:` line was corrected by hand.

- Verify no longer runs on a red gate: a failing fresh gate goes back to build, or to failed at the round cap.
- The gate tree now hashes the working tree, so uncommitted build edits no longer look like "same tree".
- Read-only stages may run `git rev-parse`; a test checks every command a skill tells an agent to run
against its stage allow-list.
- Step artifacts are validated for required fields and enums from `src/schemas.ts`. `outcome: blocked`
without a verdict routes to needs-human. Review rounds live in the runner, not the artifact.
- Rehydrate writes the build object shape. The per-run event budget is evicted when the run ends.
- `factory doctor` warns when installed skills differ from this runner, and when a preset is not verified.
- The dashboard cookie gets `Secure` over https. Durations under a minute show whole seconds (`42s`).
Board titles and stage rows pass through `plain()`, with a test that walks every read route.
- Finding re-check (P40): after verify, Claude re-checks each must/should finding against the diff in a
tool-free call and drops the unsupported ones. Fails open. Runs only when Claude is the verifier.
- Ported assembler runtime cases: literal prompt args, stdin/stderr/exit preserved, hung-process timeout.
- `bin/factory` is now type-checked (it has no `.ts` extension, so `tsc` skipped it). That found an
out-of-scope `config` in `buildWatchDeps` and a type error in `park`.
- Participant verification path: presets carry `verified`; `factory verify-agent <name>` runs one issue on
that agent, records a scrubbed fixture and prints pass/fail, cost and tokens; `make agent-matrix` runs it
for every installed agent; `docs/verify-an-agent.md` is the runbook. Only Claude is verified by us.

Live proof (Claude, splitbill issue #61, cent split): triage, plan, build, verify and pr all passed with
no manual step except approving the plan. 5 stages, about 4.5 minutes, $1.68. Verify passed 5 of 5
acceptance criteria and the PR changed only `src/money/cents.ts` and its test. `--json-schema` with no tools
returns `structured_output`, which the P40 re-check reads. Recorded fixtures: `tests/fixtures/agents/claude/`
(triage and plan) with a replay test; a preset can only say `verified: true` with a real fixture.

Not in this release:

- Claude fixtures for build, verify and pr: that run was not recorded. The next live Claude run adds them.
- The P40 re-check has not fired in a live run (verify raised no findings); only its call was probed live.
- R10: `cost_usd` stays `NOT NULL DEFAULT 0` in old databases. Incomplete usage is flagged by `usage_complete`.
- Assembler `readCommandDecision` and `runSDK` (dead code for a CLI runner) and `test/delivery.test.ts`
(task-to-pr workflow, v2.8).
- Machinist `runs-view.test.js` (needs React, jsdom and vite; our dashboard has no build step),
`artifacts_test.go` and the control-plane auth tests (a leased SQLite store and CSRF-protected HTTP
artifacts; we keep artifacts as files under `.factory/runs/`). Nothing to port to.
- Live runs of any agent except Claude; participants verify those.

## v2.5.1

Finishes v2.5: the pieces it promised and did not ship.
Expand Down
6 changes: 5 additions & 1 deletion Makefile
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
.PHONY: install check typecheck test skills-check up watch dashboard reset doctor scan
.PHONY: agent-matrix install check typecheck test skills-check up watch dashboard reset doctor scan

install:
bun install
Expand Down Expand Up @@ -53,3 +53,7 @@ doctor:

scan:
bun bin/factory scan --repo-dir "$${REPO_DIR:-.}"

# Spends tokens, so it is not part of `make check`. See docs/verify-an-agent.md.
agent-matrix:
bun scripts/agent-matrix.ts
1 change: 1 addition & 0 deletions THIRD_PARTY_NOTICES.md
Original file line number Diff line number Diff line change
Expand Up @@ -33,6 +33,7 @@ Source: https://github.com/owainlewis/assembler (MIT, Copyright (c) 2026 Owain L
- `src/logs.ts` from `src/runs.ts` [tested by `tests/ported/assembler/logs.test.ts`]
- `src/config.ts` from `src/index.ts` [tested by `tests/ported/assembler/config.test.ts`]
- `src/agents/reply.ts` from `src/index.ts` [tested by `tests/ported/assembler/outputs.test.ts`]
- `src/recheck.ts` from `examples/review-pr.ts` [no upstream test]

## License text (both projects)

Expand Down
62 changes: 55 additions & 7 deletions bin/factory
Original file line number Diff line number Diff line change
Expand Up @@ -3,8 +3,8 @@
// Thin dispatcher; all logic lives in src/ so it stays testable without a
// child process.

import { accessSync, constants } from "node:fs";
import { resolve } from "node:path";
import { accessSync, constants, readdirSync, readFileSync } from "node:fs";
import { join, relative, resolve } from "node:path";
import { GitHub } from "../src/github";
import { GitCommandRunner, Git } from "../src/git";
import { FactoryState, DEFAULT_DB_PATH } from "../src/state";
Expand All @@ -13,6 +13,7 @@ import { loadConfig, type FactoryConfig } from "../src/config";
import { EXIT, UsageError, failureJson, successJson } from "../src/cli-output";
import { advanceIssue, pollOnce, recoverInFlight, startWatch, type WatchDeps } from "../src/watch";
import { ShellGateRunner } from "../src/gates";
import { BinaryRunner, ClaudeRechecker } from "../src/recheck";
import { scan } from "../src/scan";
import { streamLogs } from "../src/logs";
import { plain } from "../src/display";
Expand All @@ -24,6 +25,8 @@ import { workspacesDir as defaultWorkspacesDir } from "../src/paths";
import { LABEL } from "../src/labels";
import { ensureRepoClone } from "../src/repo";
import { helpText } from "../src/help";
import { FixtureRecorder } from "../src/agents/record";
import { configFor, formatReport, reportFor } from "../src/verify-agent";

const args = process.argv.slice(2);
const command = args[0];
Expand Down Expand Up @@ -66,13 +69,20 @@ async function resolveCloneDir(): Promise<string> {
// Shared by watch/run/tick so all three modes (long-lived poll, one-shot CI
// step, cron tick) resolve the exact same paths and construct the exact same
// gate runner — one place, not three copies to drift (audit finding #1).
function buildWatchDeps(cloneDir: string): WatchDeps {
// The tool-free re-check calls `claude` directly, so it runs only when Claude does the verify stage.
function verifierIsClaude(config: FactoryConfig): boolean {
const name = config.stages.verify ?? config.stages.default ?? "claude";
return config.agents[name]?.preset === "claude";
}

function buildWatchDeps(cloneDir: string, config: FactoryConfig): WatchDeps {
return {
github: new GitHub(),
git: new Git(new GitCommandRunner()),
state: new FactoryState(flag("db") ?? DEFAULT_DB_PATH),
executor: new CommandExecutor(config.agents, config.stages),
gateRunner: new ShellGateRunner(),
rechecker: verifierIsClaude(config) ? new ClaudeRechecker(new BinaryRunner("claude"), cloneDir) : undefined,
cloneDir,
workspacesDir: flag("workspaces") ?? defaultWorkspacesDir(),
};
Expand Down Expand Up @@ -143,7 +153,7 @@ async function cmdPark(): Promise<void> {
if (!issueNumber) throw new UsageError("park: --issue <N> is required");
const github = new GitHub();
const issue = await github.getIssue(config.repo, issueNumber);
const runningLabels = [LABEL.triaging, LABEL.planning, LABEL.building, LABEL.verifying, LABEL.inReview];
const runningLabels: string[] = [LABEL.triaging, LABEL.planning, LABEL.building, LABEL.verifying, LABEL.inReview];
const current = issue.labels.map((l) => l.name).find((n) => runningLabels.includes(n));
if (current) await github.setStateLabel(config.repo, issueNumber, [current], LABEL.needsHuman);
await github.commentIssue(config.repo, issueNumber, `Parked by \`factory park\`: ${reason}`);
Expand Down Expand Up @@ -236,6 +246,21 @@ async function cmdRebaseline(): Promise<void> {
for (const c of moved) console.log(` ${c}`);
}

// The skills this runner ships, keyed by their path in a target repo.
function templateSkills(): Record<string, string> {
const root = resolve(import.meta.dir, "..", "template");
const out: Record<string, string> = {};
const walk = (dir: string): void => {
for (const e of readdirSync(dir, { withFileTypes: true })) {
const full = join(dir, e.name);
if (e.isDirectory()) walk(full);
else out[relative(root, full)] = readFileSync(full, "utf8");
}
};
walk(join(root, ".claude", "skills"));
return out;
}

async function cmdDoctor(): Promise<void> {
const cloneDir = await resolveCloneDir();
let config;
Expand Down Expand Up @@ -263,11 +288,11 @@ async function cmdDoctor(): Promise<void> {
}
},
},
{ repo: config.repo, cloneDir, baselineTag: config.baselineTag, factoryMode: process.env.FACTORY_MODE, agents: config.agents, stages: config.stages },
{ repo: config.repo, cloneDir, baselineTag: config.baselineTag, factoryMode: process.env.FACTORY_MODE, agents: config.agents, stages: config.stages, templateSkills: templateSkills() },
);
const allOk = checks.every((c) => c.ok);
const allOk = checks.every((c) => c.ok || c.warn);
if (has("json")) console.log(successJson({ checks }, allOk));
else for (const c of checks) console.log(` [${c.ok ? "ok" : "FAIL"}] ${c.name} — ${c.detail}`);
else for (const c of checks) console.log(` [${c.ok ? "ok" : c.warn ? "warn" : "FAIL"}] ${c.name} — ${c.detail}`);
if (!allOk && has("fix")) {
console.error("factory doctor --fix: creating missing labels");
await fixDoctor(github, config.repo);
Expand Down Expand Up @@ -325,6 +350,27 @@ async function cmdUp(): Promise<void> {
await new Promise(() => {}); // keep the process alive
}

// Runs the whole loop for one issue on one preset, with the raw output recorded as a fixture.
async function cmdVerifyAgent(): Promise<void> {
const name = args[1];
if (!name || name.startsWith("--")) throw new UsageError("verify-agent: <name> is required");
const cloneDir = await resolveCloneDir();
const issueNumber = Number(flag("issue"));
if (!issueNumber) throw new UsageError("verify-agent: --issue <N> is required");
const loaded = await loadConfig(cloneDir);
const config: FactoryConfig = { ...loaded, ...configFor(name) };
const out = resolve(flag("out") ?? `tests/fixtures/agents/${name}`);
const deps = { ...buildWatchDeps(cloneDir, config), executor: new CommandExecutor(config.agents, config.stages, new FixtureRecorder(out, name)) };
const outcome = await advanceIssue(deps, config, await deps.github.getIssue(config.repo, issueNumber));
if (outcome === "awaiting-approval") {
console.log(`Plan gate: read the plan on #${issueNumber}, approve it (/factory approve), then run this command again to finish and record the rest.`);
}
const report = reportFor(name, deps.state.listStageRuns(config.repo, { issue: issueNumber }), outcome);
console.log(formatReport(report));
console.log(`fixture: ${out}`);
if (!report.pass) process.exit(EXIT.runFailed);
}

async function cmdInstall(): Promise<void> {
const target = args[1];
if (!target) throw new UsageError("install: usage: factory install <target-dir> [--dry-run] [--update] [--ci]");
Expand Down Expand Up @@ -367,6 +413,8 @@ async function main(): Promise<void> {
return cmdDoctor();
case "install":
return cmdInstall();
case "verify-agent":
return cmdVerifyAgent();
default:
console.log(helpText());
if (command && !["help", "--help", "-h"].includes(command)) process.exit(EXIT.error);
Expand Down
3 changes: 2 additions & 1 deletion dashboard/board.ts
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@
// /api/runs feed is what still shows recently-shipped rows locally.

import type { GhIssue } from "../src/github";
import { plain } from "../src/display";
import { LABEL, isParkedLabel, isStateLabel } from "../src/labels";

export type BoardColumn = "intake" | "triage" | "plan" | "build" | "verify" | "pr" | "attention";
Expand Down Expand Up @@ -52,7 +53,7 @@ export function buildBoard(issues: readonly GhIssue[]): BoardCard[] {
const parkedLabel = names.find(isParkedLabel) ?? null;
return {
issue: issue.number,
title: issue.title,
title: plain(issue.title),
column: columnFor(names),
stateLabel,
parkedLabel,
Expand Down
8 changes: 4 additions & 4 deletions dashboard/public/app.js
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@

import { routeFromHash } from "/lib/routes.js";
import { createStatusLoader } from "/lib/status-loader.js";
import { formatDurationMillis } from "/lib/run-metrics.js";
import { formatElapsed } from "/lib/run-metrics.js";

const STAGES = ["triage", "plan", "build", "verify", "pr"];
const WAITING = new Set(["needs-info", "awaiting-approval", "needs-human", "failed"]);
Expand Down Expand Up @@ -86,7 +86,7 @@ function stationsFor(row) {
const failed = attempts.length > 0 && !attempts[attempts.length - 1].ok && (i < at || run.status === "failed");
const done = i < at || run.status === "shipped" || (i === at && attempts.length > 0 && !live && !failed && !WAITING.has(run.status));
const grow = Math.max(1, Math.min(8, Math.round(ms / 30000)));
return { name, node: h("div", { class: "station", style: `flex-grow:${grow}`, "data-done": done, "data-live": live, "data-failed": failed, title: `${name}${ms ? ` · ${formatDurationMillis(ms)}` : ""}` }) };
return { name, node: h("div", { class: "station", style: `flex-grow:${grow}`, "data-done": done, "data-live": live, "data-failed": failed, title: `${name}${ms ? ` · ${formatElapsed(ms)}` : ""}` }) };
});
}

Expand Down Expand Up @@ -214,7 +214,7 @@ function renderSheet() {
run.reason && [h("dt", null, "Reason"), h("dd", null, run.reason)]),
h("h2", null, "Stages"),
!sheet.stages ? h("p", { class: "muted" }, "Loading") : sheet.stages.length ? h("div", { class: "table-wrap" }, h("table", null, h("thead", null, h("tr", null, ["Stage", "Agent", "Took", "Tokens", "Cost"].map((c, i) => h("th", { class: i >= 2 ? "num" : "" }, c)))),
h("tbody", null, sheet.stages.map((s) => h("tr", null, h("td", null, h("span", { class: "status", "data-tone": s.exit_code === 0 && !s.killed_reason ? "ok" : "bad" }, s.stage)), h("td", null, s.agent), h("td", { class: "num" }, formatDurationMillis(s.duration_ms)),
h("tbody", null, sheet.stages.map((s) => h("tr", null, h("td", null, h("span", { class: "status", "data-tone": s.exit_code === 0 && !s.killed_reason ? "ok" : "bad" }, s.stage)), h("td", null, s.agent), h("td", { class: "num" }, formatElapsed(s.duration_ms)),
h("td", { class: "num" }, s.usage_complete !== 0 && s.tokens_in + s.tokens_out ? compact(s.tokens_in + s.tokens_out) : "Not reported"), h("td", { class: "num" }, s.usage_complete === 0 ? "Not reported" : money(s.cost_usd))))))) : h("p", { class: "muted" }, "No stage has finished yet."),
h("h2", { style: "margin-top:1.5rem" }, "Files from the run"),
!sheet.artifacts ? h("p", { class: "muted" }, "Loading") : sheet.artifacts.length ? h("div", { class: "actions" }, sheet.artifacts.map((f) => h("button", { class: "btn", type: "button", onclick: () => preview(f.name) }, `${f.name} (${compact(f.size)}B)`))) : h("p", { class: "muted" }, "This run left no files on this machine."),
Expand All @@ -234,7 +234,7 @@ function bars(title, buckets, unit) {
const max = Math.max(...buckets.map((b) => b.costUsd), 0.0001);
return h("div", null, h("h2", null, title), h("div", { class: "bars" }, buckets.map((b) =>
h("div", { class: "bar" }, h("span", null, b.key), h("div", { class: "bar-track" }, h("div", { class: "bar-fill", style: `width:${Math.max(2, (b.costUsd / max) * 100)}%` })),
h("span", { class: "bar-note" }, `${money(b.costUsd)} · ${b.attempts} ${b.attempts === 1 ? "attempt" : "attempts"} · ${formatDurationMillis(b.avgDurationMs)} avg`)))));
h("span", { class: "bar-note" }, `${money(b.costUsd)} · ${b.attempts} ${b.attempts === 1 ? "attempt" : "attempts"} · ${formatElapsed(b.avgDurationMs)} avg`)))));
}

function analyticsView() {
Expand Down
6 changes: 6 additions & 0 deletions dashboard/public/lib/run-metrics.js
Original file line number Diff line number Diff line change
Expand Up @@ -58,6 +58,12 @@ export function formatDurationMillis(milliseconds) {
return `${Math.floor(minutes / 60)}h ${minutes % 60}m ${remainingSeconds}s`;
}

// The cockpit shows whole seconds ("42s"); the ported formatter above keeps machinist's milliseconds.
export function formatElapsed(milliseconds) {
if (!Number.isSafeInteger(milliseconds) || milliseconds < 1000) return formatDurationMillis(milliseconds);
return formatDurationMillis(Math.round(milliseconds / 1000) * 1000);
}

export function formatTokenUsage(value) {
return typeof value === "string" && /^(0|[1-9]\d*)$/.test(value) ? value.replace(/\B(?=(\d{3})+(?!\d))/g, ",") : "Unavailable";
}
Expand Down
14 changes: 11 additions & 3 deletions dashboard/server.ts
Original file line number Diff line number Diff line change
Expand Up @@ -67,6 +67,11 @@ function plainRun<T extends { title: string }>(run: T): T {
return { ...run, title: plain(run.title) };
}

// The agent name and kill reason come from config and from what an agent printed.
function plainStage<T extends { agent: string; killed_reason: string | null }>(s: T): T {
return { ...s, agent: plain(s.agent), killed_reason: s.killed_reason === null ? null : plain(s.killed_reason) };
}

function plainIssue<T extends { title: string; body: string; comments: { body: string }[] }>(issue: T): T {
return { ...issue, title: plain(issue.title), body: plain(issue.body), comments: issue.comments.map((c) => ({ ...c, body: plain(c.body) })) };
}
Expand Down Expand Up @@ -111,7 +116,7 @@ export function createDashboard(state: FactoryState, github: GitHub, repo: strin
const out: ReturnType<FactoryState["listStageRuns"]> = [];
for (let after = 0; ; ) {
const page = state.listStageRuns(repo, { after, limit: 500 });
out.push(...page);
out.push(...page.map(plainStage));
if (page.length < 500) return out;
after = page[page.length - 1]!.id;
}
Expand Down Expand Up @@ -209,7 +214,10 @@ export function createDashboard(state: FactoryState, github: GitHub, repo: strin
for (const [k, expires] of sessions) if (expires <= now) sessions.delete(k);
if (sessions.size >= MAX_SESSIONS) sessions.delete(sessions.keys().next().value!);
sessions.set(id, now + SESSION_TTL_MS);
return json({ ok: true }, { headers: { "set-cookie": `${SESSION_COOKIE}=${id}; HttpOnly; SameSite=Strict; Path=/; Max-Age=${SESSION_TTL_MS / 1000}` } });
// Behind a TLS-terminating proxy the request arrives as http, so X-Forwarded-Proto counts too.
const https = new URL(req.url).protocol === "https:" || req.headers.get("x-forwarded-proto") === "https";
const secure = https ? "; Secure" : "";
return json({ ok: true }, { headers: { "set-cookie": `${SESSION_COOKIE}=${id}; HttpOnly; SameSite=Strict; Path=/; Max-Age=${SESSION_TTL_MS / 1000}${secure}` } });
},
},
{ method: "GET", pattern: /^\/api\/runs$/, label: "GET /api/runs", handler: () => json({ repo, runs: state.listRuns(repo || undefined).map(plainRun) }) },
Expand Down Expand Up @@ -261,7 +269,7 @@ export function createDashboard(state: FactoryState, github: GitHub, repo: strin
handler: (_req, _url, m) => {
const run = state.listRuns(repo || undefined).find((r) => r.id === Number(m[1]));
if (!run) return json({ error: "no such run" }, { status: 404 });
return json({ run: plainRun(run), stages: state.listStageRuns(run.repo, { issue: run.issue }) });
return json({ run: plainRun(run), stages: state.listStageRuns(run.repo, { issue: run.issue }).map(plainStage) });
},
},
{
Expand Down
Loading
Loading