Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
19 commits
Select commit Hold shift + click to select a range
b2b2257
Fix preset prompt delivery evidence, OpenCode read-only, Pi skills di…
Param-Harrison Sep 24, 2026
61abd2e
Doctor: fail custom agents without a bypass flag, time-box --version
Param-Harrison Sep 24, 2026
fac9d6c
Port Mastra Code flags (Apache-2.0) and build its argv through them
Param-Harrison Sep 24, 2026
07e6c42
CI template: install and pass keys for every preset, test it
Param-Harrison Sep 24, 2026
074db33
Doctor: do not probe --version on a CLI that has none
Param-Harrison Sep 24, 2026
00d903d
R10: stage_runs.cost_usd is nullable; an unknown cost is NULL, not $0
Param-Harrison Sep 24, 2026
52a1962
Scenario 4c expects NULL for an unknown cost
Param-Harrison Sep 24, 2026
0f66a33
Pin Codex read-only sandbox per stage (R11 is already covered)
Param-Harrison Sep 24, 2026
2d5e6ef
Pin install.sh's agent map to the presets and src/agent-dirs.ts; fix …
Param-Harrison Sep 24, 2026
4cc3d2f
Agents page: installed version and doctor rows; stages fall back to c…
Param-Harrison Sep 24, 2026
baf20c6
verify-an-agent: one section per preset, pinned by a test
Param-Harrison Sep 24, 2026
73d4d1f
Teaching map: teach/sessions.json and its test
Param-Harrison Sep 24, 2026
361e091
Origin guard: repo comes from the clone's origin; reset stays inside …
Param-Harrison Sep 24, 2026
bb19587
CHANGELOG v2.6.1, correct v2.6.0 overclaims, bump version and runner ref
Param-Harrison Sep 24, 2026
7cefe9d
reset: create labels before the issues that carry them
Param-Harrison Sep 24, 2026
ed982da
CHANGELOG: reset label order
Param-Harrison Sep 24, 2026
6074cb1
Record live Claude fixtures for every stage; scrub dash-encoded home …
Param-Harrison Sep 24, 2026
f9c6858
Teach map uses checkpoint tags; CHANGELOG records Docker and live-run…
Param-Harrison Sep 24, 2026
d40df1f
CHANGELOG: OpenCode and Pi keep stdin; list the prompt evidence and s…
Param-Harrison Sep 24, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
43 changes: 42 additions & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,46 @@
# Changelog

## v2.6.1

Residue fixes for v2.6.0. No new features.

Corrections to v2.6.0:
- The doctor did not fail a custom agent that had no bypass flag; it does now, and `--version` has a timeout.
- The CI template did not install every preset or pass its keys; it does, and a test checks it.
- The Docker note "a build with two agents was run" had no recorded evidence. This release built the default image
and ran `--version` in it: claude 2.1.281, codex 0.156.1, gemini 0.61.0, opencode 1.18.32, all equal to the pins.
- OpenCode and Pi were read as positional-only from their `--help`, but their docs show both take the prompt on
stdin, so it stays on stdin. Each preset now names the doc line that shows how it takes the prompt
(`promptVia`, `evidence`), and a test reads that line. Pi's skills dir is `.pi/skills`. OpenCode's read-only
stages run `--agent plan`; Cursor's run `--mode plan`.
- A prompt passed as an argument (Cursor) over 120 KiB is refused with a clear error, not an E2BIG failure.
- Mastra Code has no `--version` (it reads it as a prompt); doctor skips the drift check for it. Its argv is
built through flags ported from mastra under Apache-2.0 (`ee/` was not read).
- R11 was already covered: Codex gets `-s read-only` per stage, and a test pins it.

Fixes:
- A live Claude run on a fresh sandbox (`factory verify-agent claude`) passed all six stages in 281 s for $1.77;
its build, verify and pr output is recorded in `tests/fixtures/agents/claude/`. The recorder now also scrubs
Claude's dash-encoded home paths (`-Users-name-`), which it used to leave in.
- `teach/sessions.json` names immutable splitbill tags `checkpoint/<name>` instead of force-pushed branches.
- R10: `stage_runs.cost_usd` is nullable in old databases too; an unknown cost stays NULL, not $0.
- `factory reset` and `watch` act on the clone's origin. A config whose `repo` disagrees with origin is
refused, and a missing `repo` is read from origin. Before this, a copied config could close and reseed
issues on the wrong repo.
- `factory reset` creates labels before seeding issues, so it works on a fresh repo (found on the first sandbox reset).
- `factory reset` checks the baseline tag before closing anything, and closes only issues labelled
`factory:*` or seeded from `.factory/issues`. `--all-issues` keeps the old behaviour for a sandbox.
- `tests/app-agnostic.test.ts`: runtime code never names the demo app outside comments, and a Python repo on
a `trunk` branch loads, gates and refuses a wrong origin.
- The Agents page shows the installed version and doctor rows, and stages fall back to Claude like the executor.
- `install.sh`'s agent map is pinned to `src/agent-dirs.ts` and every preset's `skillsDir`.
- `docs/verify-an-agent.md` has a section per preset. `teach/sessions.json` maps each session to its release.

Not in v2.6.1:
- Live runs of any agent except Claude.
- The P40 re-check has still never fired live: the recorded run's verify stage raised no findings.
- The `resettable` opt-in, per-repo state databases and a base default from origin/HEAD (v2.6.2).

## v2.6.0

One factory, any agent. Every preset except Claude ships `verified: false`; participants verify them with
Expand Down Expand Up @@ -110,7 +151,7 @@ Not in this release:
- `gpt-5.6-terra` is unpriced, so its cost is not reported. Incomplete usage stores `cost_usd = 0` with
`usage_complete = 0`, not NULL.
- A read-only reply's artifact is capped at 16 KiB.
- The tool-free finding re-check call (P40) is still prompt text only.
- The tool-free finding re-check call (P40) was still prompt text only here; v2.5.2 made it a second tool-free call.
- Sub-minute durations keep the upstream `42.5s` format.
- Upstream `runs-view.test.js`, `artifacts_test.go` and the auth tests are not ported.
- Real Claude fixtures per stage are not recorded; they spend tokens.
Expand Down
7 changes: 4 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -295,11 +295,12 @@ Also `--db`, `--workspaces`, `--port` and the `FACTORY_*` env vars; `factory --h

## Reset any time

> **Danger.** `reset` force-pushes the base branch back to the baseline tag and closes every open
> issue and PR. Run `--dry-run` first, and never against a repo with real work on it.
> **Danger.** `reset` force-pushes the base branch back to the baseline tag and closes factory PRs
> and the factory's own issues. Run `--dry-run` first, and never against a repo with real work on it.

`factory reset` (or `--dry-run` first) closes factory PRs, deletes `factory/*` branches, forces
the base branch (`config.base`) back to the baseline tag, closes open issues, recreates the seeded ones from
the base branch (`config.base`) back to the baseline tag, closes issues that carry a `factory:*` label or match a seed
(`--all-issues` closes every open issue), recreates the seeded ones from
`.factory/issues/*.md`, and wipes worktrees and local state. Idempotent, run it before every
workshop or demo run. The dry-run lists each commit the force push would drop as `drop-commit`.

Expand Down
8 changes: 7 additions & 1 deletion THIRD_PARTY_NOTICES.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

Code ported from other projects. Each ported file starts with a `Ported from` header naming the
upstream repo, commit, path and lines; `tests/provenance.test.ts` keeps this file and those headers
in step. Only MIT-licensed code is copied.
in step. Only MIT and Apache-2.0 code is copied; mastra's `ee/` directories are never read or copied.

## owainlewis/machinist@3943516

Expand Down Expand Up @@ -78,3 +78,9 @@ SOFTWARE.
of owainlewis/machinist@3943516. Copyright 2018 The Manrope Project Authors
(https://github.com/sharanda/manrope), SIL Open Font License 1.1. The licence text is in
`dashboard/public/fonts/OFL.txt`, shipped beside the font.

## mastra-ai/mastra@68fece5

Source: https://github.com/mastra-ai/mastra (Apache-2.0, Copyright (c) 2025 Kepler Software, Inc.). Only `mastracode/` was used; the licence excludes `ee/` directories and none was read.

- `src/agents/presets/mastracode-flags.ts` from `mastracode/sdk/src/headless/flags.ts` [tested by `tests/ported/mastra/flags.test.ts`]
17 changes: 2 additions & 15 deletions bin/factory
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,7 @@ import { act, buildInbox, inboxPositionals, type InboxAction } from "../src/inbo
import { reset, rebaseline } from "../src/reset";
import { runDoctor, fixDoctor } from "../src/doctor";
import { createDashboard } from "../dashboard/server";
import { versionOf, which } from "../src/probes";
import { workspacesDir as defaultWorkspacesDir } from "../src/paths";
import { LABEL } from "../src/labels";
import { ensureRepoClone } from "../src/repo";
Expand All @@ -41,21 +42,6 @@ function has(name: string): boolean {
return args.includes(`--${name}`);
}

async function versionOf(bin: string): Promise<string | undefined> {
try {
const proc = Bun.spawn([bin, "--version"], { stdout: "pipe", stderr: "ignore" });
const out = await new Response(proc.stdout).text();
return (await proc.exited) === 0 ? out.trim() : undefined;
} catch {
return undefined;
}
}

async function which(bin: string): Promise<boolean> {
const proc = Bun.spawn(["which", bin], { stdout: "pipe", stderr: "pipe" });
return (await proc.exited) === 0;
}

async function fileExists(path: string): Promise<boolean> {
return await Bun.file(path).exists();
}
Expand Down Expand Up @@ -231,6 +217,7 @@ async function cmdReset(): Promise<void> {
issuesDir: `${cloneDir}/.factory/issues`,
workspacesDir: flag("workspaces") ?? defaultWorkspacesDir(),
statePath: flag("db") ?? DEFAULT_DB_PATH,
allIssues: has("all-issues"),
};
const deps = { github: new GitHub(), git: new GitCommandRunner() };
const summary = await reset(deps, ctx, dryRun);
Expand Down
4 changes: 2 additions & 2 deletions dashboard/analytics.ts
Original file line number Diff line number Diff line change
Expand Up @@ -35,7 +35,7 @@ function bucketBy(rows: readonly StageRun[], key: (r: StageRun) => string): Buck
key: k,
attempts: g.length,
failures: g.filter((r) => r.exit_code !== 0 || r.killed_reason).length,
costUsd: g.reduce((n, r) => n + r.cost_usd, 0),
costUsd: g.reduce((n, r) => n + (r.cost_usd ?? 0), 0),
tokensIn: g.reduce((n, r) => n + r.tokens_in, 0),
tokensOut: g.reduce((n, r) => n + r.tokens_out, 0),
avgDurationMs: Math.round(g.reduce((n, r) => n + r.duration_ms, 0) / g.length),
Expand All @@ -47,7 +47,7 @@ export function analytics(runs: readonly Run[], stageRuns: readonly StageRun[],
const finished = runs.filter((r) => FINISHED.has(r.status));
const shipped = runs.filter((r) => r.status === "shipped").length;
const spendSince = (days: number) =>
stageRuns.filter((r) => now.getTime() - Date.parse(r.finished_at) <= days * DAY_MS).reduce((n, r) => n + r.cost_usd, 0);
stageRuns.filter((r) => now.getTime() - Date.parse(r.finished_at) <= days * DAY_MS).reduce((n, r) => n + (r.cost_usd ?? 0), 0);
return {
runs: runs.length,
shipped,
Expand Down
5 changes: 3 additions & 2 deletions dashboard/public/app.js
Original file line number Diff line number Diff line change
Expand Up @@ -250,11 +250,12 @@ function analyticsView() {
function agentsView() {
return h("section", null, heading("Agents", "The coding agents the factory can run, and which stages each one serves."),
stateOr("agents", (a) => !a.agents.length ? quiet("No agents configured", "Add one under agents in .factory/config.json.") : h("div", { class: "table-wrap" }, h("table", null,
h("thead", null, h("tr", null, ["Agent", "Command", "Pinned version", "Verified live", "Stages"].map((c) => h("th", null, c)))),
h("thead", null, h("tr", null, ["Agent", "Command", "Installed", "Pinned version", "Verified live", "Doctor", "Stages"].map((c) => h("th", null, c)))),
h("tbody", null, a.agents.map((r) => h("tr", null,
h("td", null, h("span", { class: "status", "data-tone": r.configured ? "ok" : "" }, r.name), r.configured ? null : h("span", { class: "muted" }, " not configured")),
h("td", null, r.binary), h("td", null, r.pin || "Not pinned"),
h("td", null, r.binary), h("td", null, !r.configured ? "Not configured" : r.installed ? (r.version || "Installed") : "Not installed"), h("td", null, r.pin || "Not pinned"),
h("td", null, h("span", { class: "status", "data-tone": r.verified ? "ok" : "warn" }, r.verified ? "Verified" : "Verified by participants: not yet")),
h("td", null, r.checks.length ? h("ul", { class: "plain-list" }, r.checks.filter((c) => !c.ok).map((c) => h("li", { title: c.detail }, c.name))) : "", r.checks.length && r.checks.every((c) => c.ok) ? "All checks pass" : null),
h("td", null, r.stages.length ? r.stages.join(", ") : "None"))))))));
}

Expand Down
2 changes: 2 additions & 0 deletions dashboard/public/styles.css
Original file line number Diff line number Diff line change
Expand Up @@ -114,6 +114,8 @@ h1 { margin: 0; font-size: 1.5rem; line-height: 1.35; font-weight: 600; letter-s
h2 { margin: 0 0 .75rem; font-size: 1rem; font-weight: 600; letter-spacing: -.012em; }
.lede { margin: .25rem 0 0; color: var(--muted-foreground); }
.muted { color: var(--muted-foreground); }
.plain-list { margin: 0; padding: 0; list-style: none; }
.plain-list li { color: var(--muted-foreground); }

/* Controls */
.btn { min-height: 2.25rem; padding: 0 .9rem; border: 1px solid var(--border); border-radius: 8px; background: var(--surface); transition: border-color .15s ease, background .15s ease; }
Expand Down
32 changes: 30 additions & 2 deletions dashboard/server.ts
Original file line number Diff line number Diff line change
Expand Up @@ -18,6 +18,8 @@ import { LABEL } from "../src/labels";
import { InboxError, act, buildInbox, type InboxAction } from "../src/inbox";
import { plain } from "../src/display";
import { agentCatalog } from "../src/agents/docs";
import { agentChecks } from "../src/doctor";
import { versionOf, which } from "../src/probes";
import { DEFAULT_CONFIG, type FactoryConfig } from "../src/config";
import { workspacesDir } from "../src/paths";
import { runDir } from "../src/artifacts";
Expand Down Expand Up @@ -78,13 +80,39 @@ function plainIssue<T extends { title: string; body: string; comments: { body: s
return { ...issue, title: plain(issue.title), body: plain(issue.body), comments: issue.comments.map((c) => ({ ...c, body: plain(c.body) })) };
}

export function createDashboard(state: FactoryState, github: GitHub, repo: string, autoApproveDefault = false, workspaces = workspacesDir(), fleet: Pick<FactoryConfig, "agents" | "stages"> = DEFAULT_CONFIG) {
export function createDashboard(state: FactoryState, github: GitHub, repo: string, autoApproveDefault = false, workspaces = workspacesDir(), fleet: Pick<FactoryConfig, "agents" | "stages"> = DEFAULT_CONFIG, probes: { which: typeof which; versionOf: typeof versionOf } = { which, versionOf }) {
const indexHtml = readFileSync(join(here, "public", "index.html"), "utf8");

const sessions = new Map<string, number>();
let boardCache: { at: number; issues: Awaited<ReturnType<GitHub["listOpenIssues"]>> } | null = null;
let boardInflight: Promise<Awaited<ReturnType<GitHub["listOpenIssues"]>>> | null = null;

// The Agents page shows each configured agent's installed version and doctor rows.
// Probing spawns `--version`, so it is cached for a minute and shared by concurrent requests.
let agentsCache: { at: number; rows: unknown[] } | null = null;
let agentsInflight: Promise<unknown[]> | null = null;
async function probedAgents(): Promise<unknown[]> {
if (agentsCache && Date.now() - agentsCache.at < 60_000) return agentsCache.rows;
agentsInflight ??= Promise.all(
agentCatalog(fleet.agents, fleet.stages).map(async (a) => {
const row = { ...a, name: plain(a.name), binary: plain(a.binary) };
if (!a.configured) return { ...row, installed: null, version: null, checks: [] };
const checks = await agentChecks(a.name, fleet.agents, probes);
const installed = await probes.which(a.binary);
const version = installed ? ((await probes.versionOf(a.binary)) ?? "").split("\n")[0]!.trim() || null : null;
return { ...row, installed, version: version === null ? null : plain(version), checks: checks.map((c) => ({ name: plain(c.name), ok: c.ok, warn: c.warn === true, detail: plain(c.detail) })) };
}),
)
.then((rows) => {
agentsCache = { at: Date.now(), rows };
return rows;
})
.finally(() => {
agentsInflight = null;
});
return agentsInflight;
}

// Single-flight + 10s cache in front of `gh issue list` (plan: "cached gh
// listing, single-flight, 10s") so a browser polling every few seconds,
// times any number of open tabs, doesn't turn into one `gh` call per poll.
Expand Down Expand Up @@ -303,7 +331,7 @@ export function createDashboard(state: FactoryState, github: GitHub, repo: strin
method: "GET",
pattern: /^\/api\/agents$/,
label: "GET /api/agents",
handler: () => json({ agents: agentCatalog(fleet.agents, fleet.stages).map((a) => ({ ...a, name: plain(a.name), binary: plain(a.binary) })) }),
handler: async () => json({ agents: await probedAgents() }),
},
{
method: "GET",
Expand Down
5 changes: 5 additions & 0 deletions docs/agent-research/codex/docs.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
# Codex CLI: prompt delivery

Source: `codex exec --help` as documented by OpenAI (not installed here, so this is unverified locally).

- `codex exec -` reads the prompt from stdin. The preset ends its argv with `-` and sends the prompt on stdin.
7 changes: 7 additions & 0 deletions docs/agent-research/cursor/docs.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
# Cursor Agent 2026.01.23-916f423: prompt delivery

Source: `cursor-agent --help` (help.txt), captured on this machine.

- The usage line is `agent [options] [command] [prompt...]` (help.txt:2): the prompt is a positional argument. No stdin prompt is documented, so the preset passes it as the last argument.
- `--mode plan` is the read-only mode. `--force` allows commands and is not needed for a plan-mode run.
- Skills: `.cursor/skills`. Context file: `AGENTS.md`.
66 changes: 66 additions & 0 deletions docs/agent-research/cursor/help.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,66 @@
2026.01.23-916f423
Usage: agent [options] [command] [prompt...]

Start the Cursor Agent

Arguments:
prompt Initial prompt for the agent

Options:
-v, --version Output the version number
--api-key <key> API key for authentication (can also use
CURSOR_API_KEY env var)
-H, --header <header> Add custom header to agent requests (format:
'Name: Value', can be used multiple times)
-p, --print Print responses to console (for scripts or
non-interactive use). Has access to all tools,
including write and bash. (default: false)
--output-format <format> Output format (only works with --print): text |
json | stream-json (default: "text")
--stream-partial-output Stream partial output as individual text deltas
(only works with --print and stream-json format)
(default: false)
-c, --cloud Start in cloud mode (open composer picker on
launch) (default: false)
--mode <mode> Start in the given execution mode. plan:
read-only/planning (analyze, propose plans, no
edits). ask: Q&A style for explanations and
questions (read-only). (choices: "plan", "ask")
--plan Start in plan mode (shorthand for --mode=plan).
Ignored if --cloud is passed. (default: false)
--resume [chatId] Resume a chat session. (default: false)
--continue Resume the last chat session (default: false)
--model <model> Model to use (e.g., gpt-5, sonnet-4,
sonnet-4-thinking)
--list-models List available models and exit (default: false)
-f, --force Force allow commands unless explicitly denied
(default: false)
--sandbox <mode> Explicitly enable or disable sandbox mode
(overrides config) (choices: "enabled",
"disabled")
--approve-mcps Automatically approve all MCP servers (only works
with --print/headless mode) (default: false)
--browser Enable browser automation support (default:
false)
--workspace <path> Workspace directory to use (defaults to current
working directory)
-h, --help Display help for command

Commands:
install-shell-integration Install shell integration to ~/.zshrc
uninstall-shell-integration Remove shell integration from ~/.zshrc
login Authenticate with Cursor. Set NO_OPEN_BROWSER to
disable browser opening.
logout Sign out and clear stored authentication
mcp Manage MCP servers
status|whoami View authentication status
models List available models for this account
about Display version, system, and account information
update|upgrade Update Cursor Agent to the latest version
create-chat Create a new empty chat and return its ID
generate-rule|rule Generate a new Cursor rule with interactive
prompts
agent [prompt...] Start the Cursor Agent
ls Resume a chat session
resume Resume the latest chat session
help [command] Display help for command
6 changes: 6 additions & 0 deletions docs/agent-research/gemini/docs.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,6 @@
# Gemini CLI 0.61.0 (from the bundled docs and dist, no live run)
- Skills: workspace `.gemini/skills/` or `.agents/skills/` (alias wins). Context file: `GEMINI.md`.
- Headless: `gemini -o stream-json -p ""` with the prompt on stdin. Exit 0 ok, 1 error, 42 input error, 53 turn limit.
- Events (JSONL): init{session_id,model}; message{role,content,delta?}; tool_use{tool_name,tool_id,parameters};
tool_result{tool_id,status,output,error?}; error; result{status:success|error,error?,stats{total_tokens,input_tokens,output_tokens,cached,input,duration_ms,tool_calls,models}}.
- Auth: GEMINI_API_KEY or GOOGLE_API_KEY.
Loading
Loading