Repository navigation
The binary owns process: the crate skeleton and the golden harness - #258
Closed
crenshawdev wants to merge 1078 commits into
Closed
crenshawdev wants to merge 1078 commits into
crenshawdev wants to merge 1078 commits into
Conversation
Codex read the six plans against HEAD and found ten blocking criteria: the close path never refuses an out-of-lease commit, only a staged path; native completion writes a Gate decision, not a boundary record; one truth cannot feed two allocated checks; the frozen why renderer labels a prune from its ARCHIVE.md heading, not a tag. Replaced by digest 0cbdca34; plans 2 and 5 carry their lease and wording fixes as dispatch settlements.
…red P14-1-T1 Pin the owner-visible progress answer across a roadmap conflict and restart so stale status or silent document repair cannot pass.
…P14-1-T1 Give the owner one durable status answer without hiding or repairing roadmap conflicts. Preserve checked memo publication and cursor retirement, including the native acceptance key, so repeated reads agree across restart. Include the compiled progress door and retired skill removals here because the unchanged T1 check requires them before it can turn green.
Keep the query-only progress skill under binary rendering and lease ownership so its installed bytes and retired alternatives cannot drift from the single progress answer.
Use the Refusal builder so progress responses satisfy the refusal shape contract while preserving codes, the progress slot, reasons, and structured error details. Update the protected rendered-file count to include the new cad-progress skill.
… red P14-2-T1 A legacy tree with two ticked phases the documents cannot derive complete, first touched through progress: the rows must spell the provenance word and the failed UAT count, the owner's adoption-declare must refuse the derivable, unticked and natively completed phases in the typed shape and write the post-import tick with declared-at-adoption, and a caller can never hand the binary a record. The binary refuses the first progress with inputs-changed today and has no adoption-declare, so the check is red.
D-137 owes the owner a way in after first touch has already passed. adoption-declare takes a phase and a request id and nothing else: it reads ROADMAP.md and the phase documents through the import's own guarded reader, computes the record itself under the declared-at-adoption provenance, and commits it through the one non-import intent the adoption namespace guard admits, so a caller can never hand the binary a record. It refuses a natively completed phase first, because verification never ticks a row by hand, then an unticked row, then a phase the documents already derive complete. First touch also had to see its own import. The artifacts and the acceptance overlay are read together before the session opens, and the import writes the declared completions between that read and the store view, so the artifacts are observed once more against the imported store before a difference counts as a changed input. A declared completion is complete with counts the legacy SUMMARY/UAT table can never produce: two passing rows and one failure, or no UAT file at all. The derived row now says whether the acceptance overlay or the legacy table decided it, and the memo's structural check asks its clean-UAT question only of the rows the legacy table decided. Without that, a tree whose second phase was declared at import refused every progress call with derivation-conflict on its own freshly derived answer.
…-2-T2 The owner submits three captures against a bound of two and sees each one come back as a typed item with its kind, its declared phase and the bound as it stood, the third landing over the bound rather than being refused by it. A replay answers from the record it already wrote, a phase on a seed, a phase ROADMAP.md does not declare and an empty text each name the slot that refused them, and the tree gains no commit and no CAPTURE.md. There is no capture operation at all today, so the first call is refused unknown-operation.
… P14-2-T2 D-144 puts the phase in the record instead of the sentence. ItemRecord carries `phase: Option<u32>`, the writer's item validation holds it as part of the identity a revision may not change, and recall's record provenance carries it so captures can be listed by phase without parsing prose back out. The capture operation takes a request id, a kind and the owner's own text, and computes everything else: the item's id is digested from the request, so the same request twice is the same item and a replay answers from the record it already wrote instead of appending a second one. The receipt's capture report is read over the journal up to and including that item, which is why a replay still reports the bound as it stood rather than as it stands now. The bound reports and never refuses; an item over it lands with exceeded true. Each refusal names the slot that decided it: blank or control-byte text, a kind outside the three, a phase on a seed or a note, and a phase ROADMAP.md does not declare, that last one naming the number it could not find. Nothing is committed and no CAPTURE.md is written. The compiled capture door lands here because the unchanged T2 check requires `cadence capture-instructions` to equal the installed skill before it can turn green; T3 registers and pins it.
Register skills/cad-capture/SKILL.md beside cad-progress in the rendered project files, so its installed bytes are checked against their renderer and cannot drift into a second authority. The pin reads the frontmatter, the one apply call, the phase rule and the reported bound, and holds the door to what the 3.x skill did not survive: no script, no CAPTURE.md, no commit, and --cadence named as parked rather than implied to work. Raise the tools/list budget to 9,800 bytes. The payload measured 9,303 with the derived property union and measures 9,578 now that adoption-declare and capture are in the apply enum; the assertion keeps the same headroom over the new measurement rather than the old one.
Two ItemRecord fixtures in the binary crate's own test modules gained the new phase field twice, which the narrow integration verify never compiled and clippy over all targets did. Each names it once now.
The guard test pins how many project files the binary renders so a new one cannot slip past protection unnoticed. Plan 2 added the compiled cad-capture door to that table, so the pin is one short and the suite fails on a number the change itself made true.
…3-T1 A refusal the log counts but cannot place in time or join to a line is a refusal the owner has to go looking for. The check pins all three: the code, a typed located object, and the second it happened, over a roadmap conflict, an out-of-lease staged path and an absent roadmap.
A refusal the log counts but cannot place in time is one nobody can line up against anything else that happened, and a sentence that fits every code cannot be joined to the line it came from. Every record the binary writes now carries the second it was written, an imported row carries none, and each refusal keeps its own detail beside a typed located object: the roadmap line and phase for a declaration conflict, the path for an input it could not read, the rule and slot otherwise. The stamp is an observation, not a proof, so every equality a validator takes over a whole record now leaves it out. P14-T6-C stays red here. Its second scenario is the close-path refusal P14-3-T2 records, and splitting the check to manufacture a green would be lying about what this commit does.
The native task-close family answers a refusal and walks away; the owner reading the log sees an operation that never happened. Every execution-* refusal now lands as a boundary decision beside its answer, the out-of-lease staged path included, with the rule that refused as its code because the plan diagnostic stamps the same family token on all of them. The record's operation id is the boundary's own identity, so a replay answers from what is already there. Recording is a log duty and never the answer: a store that cannot take the record still returns the refusal the caller asked about.
…-3-T3 The routing decision the writer records when a dispatch is issued has an outcome edge with nothing on its far end, so nothing in the log connects the choice that was made to what the plan it routed actually reached.
The writer recorded the route it chose and then never said what came of it, so the log held a choice with no consequence attached to it. A plan that completes natively, and a schema-1 patch that lands, now give that decision a second revision whose receipt names the record that answered the dispatch. The observed effort stays Missing: no host has ever reported one, and a receipt is not a measurement.
Both new payloads sit inside an enum variant that is returned by value everywhere, and clippy is right that carrying them inline makes every Result in the derivation layer that wide. The lease refusal beside them was already boxed for the same reason.
The operation fingerprint is the preimage a validator, a replay and a recovery all have to rebuild from the same evidence, and the write-time stamp is the one field none of them can reach: the second the writer observed is gone by the time anyone else derives the transaction. Task 1 took the stamp out of every whole-record equality and missed this one, so any claim that crossed a second boundary between the writer sealing its transaction and the validator rebuilding it refused itself with "verification claim intent differs from immutable complete transition". The validator still compares the operations map, the snapshot, the items, the decisions and the generation; only the observation is gone. A record that carries no stamp serializes exactly as it always did, so every fingerprint already written into a snapshot is still the one this computes.
Recording every native execution refusal put a boundary decision into the journal of stores the phase 12, 13, 29 and 33 regressions require untouched: a refused execution operation leaves every durable byte under .planning identical, and that invariant outranks the record. The narrowing the owner ruled on keeps exactly the case P14-T6-C pins, an execution-task-close the lease rule refused over a staged path outside the plan, and stands the rest of the breadth down. So every other execution-* refusal is still answered to the caller and never heard of by the log: the rest of the close family, admission, authorization and plan publication, and execution-task-close refused by any other rule. The recording site says so beside the test. It is a known gap, left open on purpose, for a later phase to home.
The evidence submit recognizes its own replay by finding the receipt already in the log and checking it against the one it just rebuilt, and that check was whole-record equality. A retry is a second process, so it stamped a second the first one never wrote, and a submit that should have answered from what was already there refused itself with "operation identity reused for different content" whenever the two crossed a second boundary. It took the whole-workspace suite to show it, because two children of the same test usually land inside one second. Same defect as the operation fingerprint, same rule: the stamp is an observation, so the equality is taken over the content.
Preserve the reported exit even when a late completion supersedes its interruption.
Retain the host observation independently of late returns so an exited worker remains distinguishable from one still running on either host.
… P14-4-T2 Require an owner continuation before redispatch while preserving late task and plan completion across a restart.
Parenthesize computed arrays so the fixture reaches the interruption assertions.
The native plan writer must own summary projection and retain the original exit receipt for replay.
The bounded task history omits allocation, so the completion fixture must copy check revisions from evidence-read.
Native completion needs the same explicit risk selection as the established process fixture.
A host exit is durable evidence that needs an owner answer while late completion remains authoritative. Keep the same dispatch and report its age without charging a read for its own memo repair.
Callers need to locate definitions across outline grammars without reading bodies. Request-bound ordinal cursors preserve distinct units that share a line, while named omissions expose the per-call budget.
Call discovery must distinguish a callee from definitions and text mentions, and apply the last-identifier rule consistently across the supported languages.
Callers need readable targets for actual call syntax while retaining separate calls on the same line. Explicit syntactic matching and omission counts expose the limits of discovery without retaining an index.
symbol-search sent its cursor back to the first file it couldn't read, whatever the reason, so a crossing or a failed outline in the middle of a scope served the same rows on every page. call-search held its cursor on a file whose parse failed, and a file that fails every time stopped every later page at 0 rows. Now only a spent budget retries the file it stopped on, and bound::resume_at holds that rule for both.
Pin the per-kind read and returned-byte requirements for both host record formats before implementing the tally.
Keep read classification over supplied observations so host episodes expose distinct query operations and returned bytes without file access during judgment.
…3-T2 Let the owner measure complete host episodes, including every called worker or none, and inspect each read kind with its returned bytes.
A match, a missing id, a duplicate id and a malformed id are four separate rules, so a break in the first no longer stops the other three from running.
Pin the compiled instruction surfaces before changing their guidance so missing locate-first text and retired measurement paragraphs remain observable.
Keep callers on bounded discovery and unit reads across every compiled carrier. Retire the phase-specific measurement handoff so shared instructions describe the current read contract.
…-232 through context-submit Approved 2026-09-23T23:15:36Z against 2e55fd1, digest 71dbbdb6, written by the binary from the held draft. Grounded on .codex-analysis/phase39-grounding.md: no planner text states the red-to-close test-file rule, a truth has no place for its running-program part, a blank check test file passes planning and fails only at launch, and nothing asks whether a fake supplies the decision under test.
Codex drafted them, a Claude reviewer checked each one, and the confirmed findings were fixed before approval through plan-submit.
Expose acceptance of blank test locators and the planner guidance that permits them before changing the shared content policy.
A check without a usable test locator cannot launch. Refuse blank locators in the shared planning content decision so the submitter can correct the exact slot before approval, with matching admission and repair guidance.
Pin the approved planner policy and phase executor disclosure before changing their carriers so the evidence detects missing obligations.
Planners and phase executors need explicit evidence limits and the existing file freeze requirements to produce meaningful tests that can close.
Context authors need the owned-decision rule and a separate live obligation. Pin the complete required section so an omission or weakened sentence is visible before the instruction changes.
Unit results cannot settle obligations that require a running program. Context authors need to keep that evidence in scope and assign it to the first phase that can observe it, without expanding the truth schema.
Pin the required review questions and completion judgments before changing their compiled carriers.
Give plan reviewers explicit requirement, assertion and fake-decision questions while preserving the required completion judgments.
Require both verifier surfaces to distinguish execution evidence, assertion adequacy and unsupported conclusions before adding that policy.
Keep declarations, passing runs and assertion adequacy distinct so unsupported conclusions remain explicit within the existing verdict protocol.
…ools Agents reached for their own tools anyway, and keeping them out cost seven tree-sitter grammars, a file walker and a rule every role had to carry. Cadence keeps document and document-search for its own records. Review material now reads a project file by path, and the agent definitions get Read, Grep and Glob back.
…rs exist Phase 18 starts a throwaway project from scratch on each host, and nothing can start one until phase 23 ships cad-new-project, so it now runs after phase 26.
0.23.44 accepted TLS 1.3 handshake messages across encryption-level boundaries. Cadence reaches it through reqwest for provider calls.
rustfmt joins after the one-time format commit; 242 files are not rustfmt-clean today.
… cargo deny.toml allows only the licenses Cargo.lock uses today and reports duplicate versions as warnings.
crenshawdev
force-pushed
the
cadence/binary-owns-process
branch
from
September 25, 2026 20:27
0785449 to
c2f9637
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The 4.0.0 cycle:
cadence-corebecomes a session-resident Rust server, onelong-running process per session that the main thread and every subagent share.
The design is
docs/rationale/architecture-v4.md; the frozen reference is thetag
v3.7.12.Two phases of nineteen are on the branch.
Phase 1: the crate skeleton
crates/cadenceok/refused/unknown/not-applicable)as one serde-tagged enum, serialized as MCP structured content
cadence serve: an rmcp stdio server declaring one tool,cadence_versioncargo-testjob beside the untouched Node matrix intest.ymlbehind the release workflow's existing
guardjobPhase 2: the golden harness
Parity against
v3.7.12has to be a test that runs, not a claim. This phasebuilds the instrument that will hold every later phase to it.
state, covering the archived, closed, deferred, incomplete, malformed and
multi-phase shapes
operations, each capturing argv, stdin, env, exit status, the parsed
envelope, stderr, and the bytes every write left behind. Forty-seven run
inside a throwaway git repository the recorder builds with pinned identity,
dates and config, so the SHAs are reproducible
normalization.json: six named clock rules carried as DATA, each naming thefrozen source line that writes the value, applied by the recorder and read
again by the Rust side. Never a blanket date scrub - a rule that normalizes
what an operation did not write hides the divergence the harness exists to
catch, so a rule nothing reaches is not written at all
golden-driftCI job pinned to the Node major every recording names, whichrefuses a foreign interpreter before comparing a byte
codebeside itsprose
reason, and an integration test loads every recording, applies thecommitted rules to an answer, compares field-by-field on the decision-bearing
keys, and accounts for every recording by name as compared or pending
The golden test reports
compared=0 pending=156, and that is the honestnumber. The binary implements none of the recorded operations yet, so the
harness is proven by its negative controls - a wrong field, a re-stamped line
that should have been left alone, a missing recording, an empty projection -
and never by agreement. Nothing here is a parity claim, and the test is named
so that reading it as one is hard.
cadence-core/is byte-identical tov3.7.12throughout, and the JavaScriptstill does all the work. Nothing is wired to a session yet: the checksum pin,
the SessionStart bootstrap and the plugin's
.mcp.jsoncome later, where arelease exists to point at.