Repository navigation
Conversation
The enqueuer trusted hand-written wave files completely: a typo in `after` stranded a card, a duplicate `id` silently lost one, and a reviewer on the same key pool as its coder made the review worthless. `swarm wave check` now reports those failures, plus missing required fields, cyclic `after` edges, unknown profiles, and hold cards, before anything is queued.
A lead cycle needed a handful of facts that only lived in the queue, the delivery receipts and the supervisor attempts, and the external lead-digest script re-read them outside the product. `plugbrain swarm report` answers the same questions from the Brain's own store — last delivery, next runnable task, waiting lanes and why, open reviews, integration backlog, automatic reruns and time since the last manual queue change — with a stable `--json` form.
Documents Supervisor, Runner contract, quota pools, Scout->Coder->Reviewer chain, FIX1, Owner-Stop, and model-free reports per DAUERMODUS-AUFTRAG-2026-10-03.
The NVIDIA lead lane needs the same five decisions every cycle but had no deterministic way to derive them without acting: which worker is starving, which is blocked and why, which review failed after FIX1, which PASS delivery is a merge candidate, and which wave cards come next. `swarm lead-tick` now computes that picture from the Brain store and writes nothing, so the lane can render it as JSON and decide itself. 🤖 Generated with Codebuff Co-Authored-By: Codebuff <noreply@codebuff.com>
Add 'nvidia' to ModelFamily type and register 'nemotron-' and 'nvidia/' prefixes. Includes tests for nemotron-3-super-120b-a12b and nvidia/nemotron-3-ultra. 🤖 Generated with Codebuff Co-Authored-By: Codebuff <noreply@codebuff.com>
swarm quota-pool show [<pool>] now surfaces the effective policy (max concurrency and RPM), every worker bound to the pool and each reservation that shares it, so a shared provider/project pool is visible to the whole fleet instead of only its counters. A pool without a configured row reports the concurrency=1 default a run would apply. 🤖 Generated with Codebuff Co-Authored-By: Codebuff <noreply@codebuff.com>
…succeeds as no-op - src/queue.ts: Added idempotency check in deliverTask() — if task is already delivered with matching sha256, return existing task without creating a second queue_deliveries row or invoking syncDependencies twice. Different evidence on delivered task is rejected. Missing receipt sha256 on delivered task is rejected. - src/coord/turn-delivery.ts: Added same idempotency check in deliverTaskWithEvidence() at the API layer, so HTTP /api/queue/deliver also returns 200 for duplicate deliveries with same evidence. - test/queue.test.ts: Four new tests covering idempotent duplicate delivery, different evidence rejection, missing sha256 rejection, and invalid task state rejection. - test/queue-e2e.test.ts: Updated E2E test to expect 200 (idempotent) for duplicate delivery instead of 403, and verify delivered_path and updated_at remain from first delivery. All tests pass: 18 unit tests + 1 E2E test = 19 tests exit 0.
When a worker process dies mid-turn, check if the agent has a current task (agents.task_id). If no task is assigned, mark the agent as needs-task instead of blocked, and don't notify the integrator. This prevents false blocked bookings when turn start --claim finds no work.
A stale `turn_state` from an unrelated turn released every task waiting on a predecessor the holder had just claimed, and returning a never-started task to the pool on the final launch error was a release in disguise. Require a turn end at or after the claim, leave launch-error tasks claimed and visibly blocked, and reset the turn state when the supervisor claims. Adds negative tests for both paths. Generated with Codebuff 🤖 Co-Authored-By: Codebuff <noreply@codebuff.com>
- Schema: agents.owner_stopped_at/by/note, watchdog_settings.owner_stop_global/note/stopped_by/stopped_at - swarm-ops.ts: ownerStopAgent, ownerStopGlobal, ownerResumeAgent, ownerResumeGlobal, isAgentOwnerStopped, getAgentOwnerStopInfo - watchdog.ts: silentWorkers excludes owner-stopped agents; watchdogSettings includes owner stop fields - recordTurn blocks turn start for owner-stopped agents (per-agent and global) - agentsBoard adds 'owner-stopped' attention flag - swarm-cli.ts: stop --owner <agent> --note <reason>, stop --global --note <reason>, resume <agent>, resume --global Abnahme: node --test test/swarm-ops.test.ts test/coordination.test.ts → Exit 0 (19 tests) 🤖 Generated with Codebuff Co-Authored-By: Codebuff <noreply@codebuff.com>
- skip ensureSwarmOpsSchema and assertActiveSupervisorAttempt for 'wave' step in runSwarmCli - add CLI test with temporary PLUGBRAIN_HOME that validates WELLE-01.json
… events - classifySupervisorFailure now parses JSONL log for provider error events (kind='error') - Word search for quota/auth patterns only applied to correlated provider error messages - Raw text '429' in agent_message or other non-error events no longer triggers quota - Added extractProviderErrorMessages helper to extract error event messages from JSONL - Updated quotaBackoffMs to also only check error events for retry-after - Removed unused readRunOutput function - Updated tests to verify new behavior: negative test (429 in message -> crash), positive tests (429/auth in error -> quota/auth)
- queue_deliveries: Composite PK (task_id, delivery_attempt), neue Spalten worktree, commit_hash - deliverTaskWithEvidence: Ermittelt attempt (max aus worker_task_attempts + queue_deliveries + 1), löst worktree aus agents.worktrees via Evidence-Pfad, liefert commit_hash (40-char SHA-1 aus sourceRevision) - agents: worktrees Spalte (JSON) für Worktree-Auflösung - Migration in ensureQueueSchema für Bestehende DBs Alle Delivery-Tests (Q-chain E2E, fleet-automation, swarm deliver) grün. 🤖 Generated with Codebuff Co-Authored-By: Codebuff <noreply@codebuff.com>
…ssed tasks
- Add plan_ref column to queue_tasks schema for mandate validation
- Modify claimNextTask() to validate plan_ref against ledger for unaddressed tasks
- Add check for existing claims (agent can only claim one task at a time)
- Update enqueueTask() to accept planRef parameter with M\d{2} validation
- Update API endpoint to accept planRef in POST /api/queue
- Add tests T1-T6 for acceptance criteria:
T1: Worker claims pending task with valid plan_ref in ledger
T2: Worker gets null for task without plan_ref
T3: Worker gets null for task with plan_ref not in ledger
T4: Worker claims explicitly addressed task regardless of plan_ref
T5: Worker with existing claim gets no second claim
T6: No second queue table introduced
- Update existing tests to use planRef
- Update e2e test with ledger setup
…inues - reconcileWorkerRuns: call ensureIntegrator before updating agent to needs-task when no current task exists (fixes 'actual undefined, expected 0' assertion) - supervisor.ts: ensureRunnerSchema before checking active supervisor attempt - isolated-home.ts: clean supervisor env vars between tests - resources.ts: anchor credential pattern to prevent false positives on task IDs - test/swarm-runner.test.ts: add debug logging for test diagnostics Abnahme: node --experimental-strip-types --test test/swarm-runner.test.ts test/watchdog.test.ts test/queue.test.ts exit 0 🤖 Generated with Codebuff Co-Authored-By: Codebuff <noreply@codebuff.com>
This reverts commit cee1d77.
# Conflicts: # src/cli.ts # src/swarm-cli.ts
- Add wave check command alongside existing report command - Preserve both report and wave check functionality in swarm-cli.ts - Update help text in cli.ts to include both subcommands
…61003 - Add artifact binding with attempt tracking and worktree resolution - Preserve idempotent delivery check from existing branch - Update schema to support commit hash and worktree fields
- Add failure classification logic for supervisor (quota, auth, crash, clean) - Update tests to use log file based classification - Preserve readRunOutput and setAgentTurnState functions from HEAD
This reverts commit 22b2bc5.
- Update usage comment in src/cli.ts to include 'deliver' right after 'enqueue' so the test regex matches - Add missing subcommands: retire, claim, release, wave-done - Update default case usage text to match - Remove [TEST] debug console.log statements from test/swarm-runner.test.ts
…1439dcc; Lead-Test volle Suite 531/585, nur M2 vorbestehend rot)
- Add ClaimNextTaskOptions with requireMandate and singleClaim flags (default false)
- Default behavior restored to pre-B15: unaddressed tasks claimable without plan_ref, agents can hold multiple claims
- B15 tests T1-T6 updated to use { requireMandate: true, singleClaim: true }
- B12 tests (3 delivery tests) unchanged - work with default behavior
…vor B15; Lead-Test volle Suite 522/576, nur M2 vorbestehend rot)
…o release/0.6.0-rc.1
Bumps the version to 0.6.0-rc.1, adds the changelog entry and release notes for the continuous-mode candidate, brings the README status up to date and drops a machine-specific path from a doc comment.
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: e78ece1761
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| // A long opaque token mixing cases and digits. Git hashes (lower-case hex) | ||
| // and ordinary words do not match. | ||
| /(?=[A-Za-z0-9_-]*[a-z])(?=[A-Za-z0-9_-]*[A-Z])(?=[A-Za-z0-9_-]*\d)[A-Za-z0-9_-]{32,}/, | ||
| /^(?=[A-Za-z0-9_-]*[a-z])(?=[A-Za-z0-9_-]*[A-Z])(?=[A-Za-z0-9_-]*\d)[A-Za-z0-9_-]{32,}$/, |
There was a problem hiding this comment.
Keep opaque credential matching inside surrounding text
Anchoring this fallback pattern means a mixed-case opaque secret is detected only when it is the entire field. Inputs such as runner arguments or notes containing --token AbCd...1234 now pass assertNotCredential, and redactCredentialText leaves the same embedded value visible, allowing otherwise-unrecognized provider credentials to be stored and printed.
Useful? React with 👍 / 👎.
| UPDATE agents | ||
| SET owner_stopped_at = ?, owner_stopped_by = ?, owner_stop_note = ?, | ||
| turn_state = 'paused', turn_state_at = ?, turn_summary = ? | ||
| WHERE id = ? AND workspace_id = ? | ||
| `).run(now, input.by, note, now, `owner-stopped: ${note}`, input.agentId, input.workspaceId) |
There was a problem hiding this comment.
Deactivate the supervisor when applying an owner stop
When the target has an active or idle supervisor, this update only pauses the agent row and leaves worker_supervisors.active = 1. runSupervisorLoop checks that active flag rather than the owner-stop columns, then sets the agent back to working and can claim another task, while startSupervisor can also reactivate a stopped agent; thus the advertised persistent stop does not prevent new work.
Useful? React with 👍 / 👎.
| const rows = db.prepare(`SELECT q.id, q.title, q.delivered_path, q.claimed_by, | ||
| d.delivered_by, d.review_judgment, d.source_revision | ||
| FROM queue_tasks q JOIN queue_deliveries d ON d.task_id = q.id | ||
| WHERE q.workspace_id = ? AND q.state = 'delivered' AND d.review_judgment IS NOT NULL |
There was a problem hiding this comment.
Associate PASS reviews with their source merge candidate
In the normal automation flow, the source task's delivery has no judgment, while the generated Review: task carries the PASS receipt. This query selects only tasks whose own receipt has a judgment, and the following title filter discards review tasks, so a real reviewed source never appears in mergeCandidates; the query needs to map review receipts back through fleet_automation_tasks or review_routes.
Useful? React with 👍 / 👎.
| // Check if the agent has a current task | ||
| const agentTask = db.prepare('SELECT task_id FROM agents WHERE id = ?').get(row.agent_id) as { task_id: string | null } | undefined | ||
| const hasCurrentTask = agentTask?.task_id !== null && agentTask?.task_id !== '' |
There was a problem hiding this comment.
Determine current work from the run or queue claim
For a supervisor-launched worker, the supervisor calls claimNextTask directly and passes the task to startWorkerRun, but never writes agents.task_id. If that process exits without a turn end, this check therefore treats it as taskless, marks it needs-task, and suppresses the lead notification even though worker_runs.task_id and the queue still show claimed work; use the current run's task or an actual claimed-task lookup.
Useful? React with 👍 / 👎.
| }) | ||
|
|
||
| test('CLI wave check works with temporary PLUGBRAIN_HOME', () => { | ||
| const wellenFile = join(process.cwd(), '..', '..', 'koordination', 'masterplan', 'inputs', 'NACHT-20261003', 'WELLE-01.json') |
There was a problem hiding this comment.
Use a repository-owned fixture for the wave CLI test
Under the test command's normal repository working directory, traversing two parents makes this point outside the checkout (for this checkout it becomes /koordination/...). A repository-wide search finds no NACHT-20261003/WELLE-01.json, so a clean CI checkout fails before exercising the CLI; create the synthetic wave in a temporary directory or commit a fixture under test/.
Useful? React with 👍 / 👎.
| // For unaddressed tasks (addressed_to IS NULL), only allow if plan_ref is valid in ledger (when requireMandate: true) | ||
| if (requireMandate && candidate.addressed_to === null && (candidate.plan_ref === null || !validPlanRefs.has(candidate.plan_ref))) { | ||
| db.exec('COMMIT') | ||
| return null |
There was a problem hiding this comment.
Filter mandate eligibility before limiting the candidate
With requireMandate enabled, a higher-priority unaddressed task lacking a valid plan_ref is selected first and then causes an immediate null return. Any lower-priority task with a valid ledger reference—or an explicitly addressed task—remains unclaimable until the invalid queue head is removed, so the opt-in scheduler can report an empty queue despite eligible work.
Useful? React with 👍 / 👎.
| const maxAttemptFromDeliveries = attemptResult2?.maxAttempt ?? 0 | ||
| const attempt = Math.max(maxAttemptFromWorkerTasks, maxAttemptFromDeliveries) + 1 |
There was a problem hiding this comment.
Record the active supervisor attempt on delivery
During supervised attempt N, worker_task_attempts already contains attempt N before the worker delivers. Taking that maximum and adding one writes the receipt as attempt N+1, so the artifact is bound to an attempt that never ran and cannot be reliably correlated with its attempt token, run, or outcome; only unsupervised deliveries correctly become attempt 1.
Useful? React with 👍 / 👎.
| } | ||
| case 'stop': { | ||
| const usage = 'plugbrain swarm stop <agent|--all> [--owner] [--note <reason>] [--global]' | ||
| const target = need(pos[0], usage) |
There was a problem hiding this comment.
Parse
--all before requiring a positional target
For the documented swarm stop --all --owner --note ... form, positionals discards --all, so pos[0] is undefined and this need throws before the global branch can run. The same ordering breaks resume --all; treat --all/--global as the target before validating a positional agent.
Useful? React with 👍 / 👎.
PlugBrain 0.6.0-rc.1 — continuous mode
This brings
mainup to date with the work done since 0.5.0-dev.1 (30 September): the parts of the overnight continuous mode loop (scout → coder → reviewer → rework → integration → next task) that are built and tested. Release notes:release/RELEASE-NOTES-v0.6.0-rc.1.md.Added
plugbrain swarm report [--since] [--json]— model-free fleet reportplugbrain swarm wave check <file>— validates a wave file before enqueueingplugbrain swarm lead-tick [--dry-run] [--json]— deterministic lead-cycle decisionsplugbrain swarm quota-pool show [<pool>]— effective pool policy and holdersswarm stop … --owner) that survives restarts and is never lifted by the watchdogclaimNextTask(default unchanged)bench/m09-parity.mjsFixed
needs-taskinstead ofblockedswarmsubcommand againHousekeeping
package.json, lockfile, Claude plugin manifests)Tests
Windows, full suite on this branch: 591 tests, 537 passed, 53 skipped, 1 failed —
notes.test.tsM2, which reads the author's real vault and fails locally when a referenced file is missing (pre-existing, unrelated). Every merged card was additionally tested on its own before integration. CI on this PR covers Windows, macOS and Linux.Not included
Notes for merging
maingets one clean commit.