Autonomous agents have no persistent cognitive state — goals lost between calls, beliefs stale, anomalies never elevating to reusable Laws. HipCortex is the cognitive reconstruction engine that closes the loop: Laws extracted from surprising experiences, Policies driving state evolution, causal SCM, goal scheduling, belief revision — served locally over MCP + REST.
⭐ If that solves a pain you feel, star the repo — it helps others find it.
💬 Tried it? Open an issue or leave a 👍/👎 comment — real feedback steers the next release.
This repository is the public developer surface (docs, client SDKs, connectors, issues, release artifacts). New engine development lives in private hipcortex-core. Details: DUAL_REPO.md · NOTICE.
Every agent invocation starts cognitively blind. Goals set in one call vanish before the next. Beliefs accumulated from observations are never revised when contradicted. Actions taken by the agent never update its world model. Decisions leave no audit trail. There is no loop — just isolated acts.
HipCortex is the substrate that closes it: a local causal graph of goals, beliefs, decisions, and observations, with a reasoning loop that feeds every action back into prediction, served over HTTP + MCP to any agent host.
| Without HipCortex | With HipCortex |
|---|---|
| Goals re-stated every call | GoalScheduler tracks + prioritizes across sessions |
| Stale beliefs silently persist | BeliefInvalidator detects contradictions, decays confidence |
| Actions never update world model | WorldModelUpdater closes the feedback loop |
| Decisions leave no trace | DecisionPayload + provenance chain per act-phase |
| Agent doesn't know what it's allowed to do | ActionRegistry + ExecutionGate answer that in one call |
| Probe target selection is blind | IG-ranked probes (epistemic × deficit × probe_penalty) select highest-information entity first; grounded → never re-probed |
| Probe outcomes don't update the world model | update_from_receipt writes dual transitions (meta-probe + domain P(s′|s,a)) into WM |
| Successful probes leave beliefs unchanged | BeliefExecutive::reinforce via derived_from/evidence provenance — not substring |
| IDE exit breaks autonomy | Headless IntentRunner polls and dispatches intents without the IDE open |
| Action ordering within a goal is arbitrary | GoalScheduler::plan_action_sequence orders success_factors by WM MAP probability — grounded first |
| Tool recommendation ignores actuator liveness | filter_liveness removes probe-failed/stale MCP servers using WM entity_contact heartbeats |
| 3-month claim backed only by unit suites | soak_sit.rs: 500-iter temporal decay + WM convergence + bounded-growth proof |
| No public differentiation metric vs Mem0/Zep/Letta | 10-question substrate scorecard with code refs + GET /substrate/scorecard |
Unknown sensor probe returns fake ok=True |
Honest grounding: unknown sensor → {reachable:False, error:"unknown_sensor:<id>"} — WM never poisoned (v3.1.0) |
| Restate renames factor but never flags next step | blocked_factors + probe_required Temporal per blocked factor, derived_from=goal_id (v3.1.0) |
| Context cost grows with transcript — no OpEx proof | get_budget MCP tool: substrate_tokens vs naive_transcript_tokens; consolidation ratio durable via GET /substrate/budget (v3.2.0) |
| Wall guard meter claims host context coverage | Honest wall_status (bounded/at_risk/exceeded) + [honest] disclaimer: only MCP output metered; per-actor _live_beliefs_seen_actors discipline (v3.3.0) |
| 3-month claim backed only by WAL reopens | Published field log: real server subprocess + HTTP + file edit + kill+restart → after_restart=14 PASS; 605 stale VSIX assets deleted (v3.4.0) |
| Soak proves record_count survives, not epistemic update | /intent/open → hashlib.sha256 → /intent/receipt → was_surprising=True → Belief{confidence=0.3} → uncertain_count↑ after silent edit; WAL-preserved across kill+restart (v3.5.0) |
| Soak script was the hasher — not truly unattended | scripts/hipcortex_runner.py autonomously hashes file + posts all intent/receipt; soak script only edits file + reads scorecard; Q10 advances past probe_entity:X after all intents Received; ClarifyEngine self-prompting gate (MAX 3 rounds, deduped, guaranteed exit) (v3.6.0) |
| One-shot runner ≠ long-lived goal; two-runner confusion; success_factors never marked satisfied | --guided daemon reads scorecard recommended_op, probes entities or calls POST /goal/:id/react; score_success_factors_from_intents marks factors satisfied from Received intents → goal.status = Succeeded across multiple iterations (v3.7.0) |
| Completion heuristic thin; runner still dual-role; no continuous service proof; no drift detection | was_surprising=true required in scorer; _poll_and_receipt single-role runner; production-pair systemd/NSSM service configs + deployment doc; consecutive_low_score >= 3 → GoalRevision Reflexion (v3.8.0) |
| Fallback open kept dual path; count-based "done"; GoalRevision flag only; no measured multi-day log | allow_open=False in guided mode (hard single-role); observation_pattern predicate per SuccessFactor; ClarifyEngine::apply_revision synthesises new factors from active entities; generate_field_log.py produces 24h session artifact (v3.9.0) |
| Passive capture required per-channel client instrumentation — VSIX break silently killed memory | Universal server-side Axum middleware captures every mutation (POST/PUT/DELETE) from any channel — MCP, VSIX, REST, CLI, LangChain — zero client changes; X-Actor header attribution; AppState.passive_capture_enabled; fire-and-forget Temporal write; 262 integration tests 0 failures (v3.10.0) |
Clarify ladder was advisory; a removal could not be persisted through the store's own primitives; /memory/embed and /memory/query drifted from the write path's vocabulary; CI never executed several suites that existed |
Clarify H1–H10 closed and the ladder made authoritative; durable removals + delete_by_ids/upsert/delete_by_actor store primitives; one guardrail and one record_type vocabulary across /memory/add, /memory/embed, /memory/query; pipeline enforcement G1–G8 — CI now runs the suites it previously skipped (v3.11.0) |
| Server crash before shutdown lost worldmodel state; no data snapshot before upgrade or restart; no actor rename/merge; VSIX and Python SDK wrote competing MCP entries; schema compat not unit-proven | POST /v1/server/shutdown flushes WM first; POST /v1/backup atomic tar.gz of 4 data files + Python snapshot CLI; POST /v1/actor/merge + MCP merge_actor; VSIX adds HIPCORTEX_ACTOR=repo-basename; _is_vsix_managed_entry dedup; 4 schema-compat unit tests (v3.12.0) |
Safety 403 used bare status code — MCP showed "server down" for PII-flagged payload; /health timed out on loaded Windows; passive-capture and env-receipt metrics conflated; clarify gate gave no recommendation on blocked goals |
Safety 403 body surfaced as text; health/MCP timeouts → 30 s; api_mutations_captured + env_receipts separate scorecard fields; clarify gate returns self-prompt tier recommendation — always bounded and deduplicated (v3.13.0) |
clarify_goal() emitted success factors but no decidable AC — engine gate rejected Phase 1 output; uncertainty detection had no route to self-prompting; hipcortex doctor reported ok for stale pre-lifecycle installs |
AcceptanceCriterion 1:1 with success factors + non-empty observation_pattern; route_uncertainty T0→T3 — self-prompt outranks asking regardless of cost; 4-marker discriminative SKILL check (v3.14.0) |
Lifecycle gates detected uncertainty but returned no clarification route — agents called route_uncertainty separately; agent's ladder position never passed to gates; HC MCP went offline between sessions with no auto-recovery |
check_progress/plan_validation/should_exit embed clarify_route: ClarifyRouteInfo when uncertainty; tiers_spent + cost_of_wrong_execution wired agent→REST→MCP; clarify_exhausted on ExitDecision; SessionStart hook + extension 30 s reconnect (v3.15.0) |
| Anomaly evidence accumulates but never elevates to reusable structural Laws; Policies not first-class in state evolution; GC has no sparsity pressure on weak Laws | MemoryType::Law + Policy; LawExtractor writes Laws from ≥3 surprising Intents (structural, no LLM); PolicyRegistry + DigitalTwin::step_with_policies() — Policies first-class in S_{t+1} = f(S_t, Policy); MDL sparsity GC (gc_action_for_law, threshold=0.5); GET /laws + GET /policies/:entity_id; 379 lib + 546 unit + 186 integration + 59 property (v3.16.0) |
KARM had no read contract on HC snapshot — Laws/Policies/uncertainty/failures missing from CognitiveSnapshot; no Provider/Sink trait boundary; no unified TransitionView or PredictionError |
CognitiveSnapshot extended with laws, policies, failures, uncertainty (#[serde(default)], backward-safe); CognitiveStateProvider + CognitiveStateSink traits — KARM never locks MemoryStore directly; transitions_since() + prediction_error() from existing open_intents + CalibrationTracker; REST GET /snapshot/:actor/transitions + /prediction_error; 379 lib + 555 unit + 190 integration + 59 property (v3.17.0) |
| Transitions feed ignored its cursor and grew without bound; failed receipts and in-flight intents reported as settled; a settle during a snapshot read could be missed; no read model for the structural causal model, and parent order (so weight binding) was random per graph | Settle stamps (settled_tx, receipt_ok) — transitions?since=N pages settled_tx > N with next_cursor, horizon_tx, truncated; retention 4096 + compaction horizon; snapshot reads tx_cursor first; canonical sorted parents (weights[i] ↔ parents[i]); ScmView + GET /snapshot/:actor/scm; MCP snapshot_transitions + snapshot_scm + SDK parity; 379 lib + 597 unit + 194 integration + 59 property (v3.18.0) |
Version strings disagreed across surfaces — GET /openapi.json served 3.12.0, HipCortexClient.VERSION was 3.12.0, the PyPI page described v3.12.0 and the VSIX README's channel section 3.14.0; hipcortex channels named the 3.12.0 VSIX |
Every surface at 3.18.1; route_parity_sit pins OpenAPI info.version to the crate version; the fallback VSIX name derives from __version__; channel docs rewritten (v3.18.1) |
A VS Code window on an older extension stopped a newer running server and started its own older binary on the newer server's data; the chat help header and activation log said v3.14.0 in every release since 3.14; the Marketplace page's LICENSE, docs/*.md and package.json links returned 404 |
A healthy server newer than the bundled binary is reused in both strictServerVersion modes, never replaced; the extension version is read from its installed package.json; the VSIX is packaged with --baseContentUrl/--baseImagesUrl at vscode-extension/ (v3.18.2) |
Patch release for the VS Code extension; the engine is unchanged. Before 3.18.2 the extension replaced any running server whose major.minor differed from its bundled binary — including a newer one — so a window still on an older extension stopped a newer server and started its own older binary on data the newer server had written (the downgrade the 3.18 upgrade note calls unsupported). A healthy server newer than the bundled binary is now reused in both strictServerVersion modes, with a log line asking you to update the extension; an unhealthy server, or an older one outside the policy (a lower major.minor by default, any lower version under strictServerVersion), is still replaced. The chat help header and activation log now read the version from the installed package.json (they said v3.14.0 since 3.14), and the VSIX is packaged so the Marketplace page's LICENSE, docs/*.md and package.json links resolve. Extensions older than 3.18.2 still replace newer servers, so reload every VS Code window after updating.
Patch release with no engine behaviour change. 3.18.0 shipped version strings that disagreed with the binary: GET /openapi.json reported 3.12.0, HipCortexClient.VERSION was 3.12.0, the PyPI project page described v3.12.0, and the VSIX README's channel section described 3.14.0. Every surface now reports 3.18.1, and a test pins the OpenAPI version to the crate version so that literal cannot drift again. Everything below about v3.18.0 still applies.
Makes HipCortex consumable by an external control loop as "snapshot once, then follow the changes". The v3.17 feed could not carry that loop: it ignored its cursor, never evicted settled intents, reported failed receipts as ok and in-flight intents as settled, and a settle landing during a snapshot read could fall between the snapshot and the feed after it.
| Change | Gap | Fix |
|---|---|---|
| Settle stamps | transitions_since(actor, since) ignored since; ok was status == Received, so a receipt with ok: false reported ok: true; InFlight intents appeared as settled |
ActionIntent gains server-owned settled_tx + receipt_ok; only settles stamp (first receipt, late receipt after expiry, deadline expiry via new TxKind::IntentExpire); the feed returns settled_tx > since, ascending; ok = Received && receipt_ok |
| Paged feed with a horizon | Settled intents were never removed — unbounded growth | Retention 4096 (with_transition_retention); TransitionPage { transitions, next_cursor, horizon_tx, truncated } — truncated tells the consumer to re-bootstrap from the snapshot. Intents are not rehydrated on restart, so every pre-restart cursor reports truncated |
| Consistent cut | snapshot() read tx_cursor after its content |
Cursor read first; stamps allocated under the open_intents lock after every effect applies — snapshot → N then transitions(since=N) delivers every settle at least once (upsert by intent_id) |
| Canonical parent order | Parents came out of a HashMap, so weights[i] bound to a random parent per graph |
Parents sorted; weights[i] binds to parents[i] (documented on StructuralEquation) |
| SCM read model | No view of structural equations, weights or interventions | ScmView — nodes, equations, edges, interventions, acyclic, as_of_tx; identical graphs give byte-identical JSON — served at GET /snapshot/:actor/scm |
| MCP + SDK parity | New routes were REST-only | MCP snapshot_transitions + snapshot_scm (64 tools); HipCortexClient.snapshot_transitions() + snapshot_scm() |
Upgrade note: once 3.18 has written IntentExpire entries, downgrading to ≤ 3.17 is unsupported — older binaries skip unknown kinds when restoring the tx counter and can reuse tail tx ids.
Test coverage: 379 lib + 597 unit (+42) + 194 integration (+4 SIT) + 59 property + 269 Python (+11) — 0 failures in the petgraph build. The web-server build keeps 3 failures that predate this release (ac_g1_schema_mismatch_payload_returns_422_not_panic, clarify_ladder_sit::react_on_a_goal_that_is_still_unclarifiable_is_rejected, the clarify_engine.rs doctest).
Closes the KARM-readiness gap: CognitiveSnapshot was real and cursor-aware but incomplete for the KARM handover loop — Laws, Policies, uncertainty, and failures were absent from the view; there was no Provider/Sink trait boundary; and no unified TransitionView or PredictionError read model.
| Change | Gap | Fix |
|---|---|---|
| Snapshot completeness | CognitiveSnapshot omitted laws, policies, failures, uncertainty — KARM had to lock MemoryStore directly to assemble context |
CognitiveSnapshot extended with 4 #[serde(default)] fields: laws: Vec<LawSummary>, policies: Vec<PolicySummary>, failures: Vec<FailureSummary>, uncertainty: UncertaintySummary; backward-safe (old JSON parses with defaults) |
| Provider/Sink trait boundary | No formal trait boundary — KARM callers reached into MemoryStore/WorldModelEnhanced directly, creating lock-order risks | CognitiveStateProvider + CognitiveStateSink traits in cognitive_contracts.rs; CognitiveHandle<B> implements both; KARM reads/writes only through these traits |
| TransitionView + PredictionError read models | Transition data distributed across open_intents + receipt logs with no unified view; prediction error not exposed outside CalibrationTracker | TransitionView (from open_intents, non-Open only) + PredictionError (from CalibrationTracker::snapshot().prediction_error_ewma) in transition_view.rs; CognitiveHandle::transitions_since() + prediction_error() |
| REST surfaces | No endpoints for KARM to poll transitions or prediction error | GET /snapshot/:actor/transitions?since=<tx> + GET /snapshot/:actor/prediction_error |
Test coverage: 379 lib + 555 unit (+9) + 190 integration (+4 SIT) + 59 property + 258 Python = all green, 0 failures.
Closes 4 gaps turning HipCortex from "cognitive state substrate" into "dynamic reconstruction engine": surprising experiences now elevate to durable structural Laws; Policies are first-class objects driving entity state evolution; the GC applies MDL sparsity pressure so the long-term representation stays compressible; REST surfaces both.
| Change | Gap | Fix |
|---|---|---|
| Anomaly → axiomatic write-back | Surprising Intent records accumulated evidence but never elevated to reusable structural Laws — each anomaly was processed in isolation | LawExtractor::attempt_extract() clusters surprising Intents by target_entity; ≥3 in a cluster → writes MemoryType::Law via structural equation (no LLM, O(n)); idempotent; calls route_uncertainty when < 2 causal variables identified (T0 self-prompt first) |
| Policies as first-class citizens | State evolution S_{t+1} = f(S_t, action) was driven by discrete memory records + goal/react cycles — no first-class reactive rules attached to entities |
PolicyRegistry stores active MemoryType::Policy records keyed by entity_id, sorted by priority; DigitalTwin::step_with_policies(action, registry, entity_id) checks registry before applying dynamics — highest-priority matching Policy overrides the action arg |
| MDL sparsity pressure on Laws | CognitiveGC archived/deleted records by reference count only — Laws with low information gain were kept indefinitely |
gc_action_for_law(record_id, mdl_score) -> GcAction: score ≥ MDL_KEEP_THRESHOLD (0.5) → Keep regardless of references; below threshold → Archive if referenced, Delete if orphaned |
| REST surfaces for Laws + Policies | No endpoints to inspect the live Law or Policy population | GET /laws returns all MemoryType::Law records; GET /policies/:entity_id returns active Policies for an entity sorted by priority desc |
Test coverage: 379 lib + 546 unit (+4) + 186 integration (+4 SIT) + 59 property + 258 Python = all green, 0 failures.
Closes 2 cohesion gaps: lifecycle gates detected uncertainty but returned no clarification route in their response; HipCortex MCP went offline between sessions with no auto-recovery.
| Change | Gap | Fix |
|---|---|---|
| Clarify route embedded in lifecycle gates | check_progress, plan_validation, should_exit detected uncertainty but returned no clarification advice — agents called route_uncertainty as a second pass |
All three gates accept tiers_spent (+ cost_of_wrong_execution for check_progress) and embed clarify_route: ClarifyRouteInfo when uncertainty detected; ExitDecision gains clarify_exhausted flag |
| tiers_spent wired agent→REST→MCP | Agent's position on the T0→T3 ladder was never passed to lifecycle gates — route decisions ignored how many self-prompt tiers had been spent | REST handlers extract tiers_spent / cost_of_wrong_execution from request body; MCP tool schemas expose them as optional params (default 0 / 1.0); 6 AC-L7 Python tests + 8 Rust AC-L1..L8 |
| HC offline resilience | HC MCP went offline between sessions with no auto-recovery — agent sessions silently wrote to local files | ~/.claude/settings.json SessionStart hook calls scripts/hipcortex_start_if_dead.py (checks :3030, spawns binary, waits 10 s); VS Code extension polls health every 30 s and restarts if unreachable; global ~/.claude/mcp.json corrected to scripts/hipcortex_mcp_launcher.py |
Test coverage: 379 lib + 521 unit + 182 integration + 59 property + 258 Python (incl. 6 AC-L7 MCP + 43 doctor) = all green, 0 failures.
Closes 3 seams in the clarification and goal-execution lifecycle: Phase 1 output was rejected by the engine's own gate when acceptance_criteria were absent; uncertainty detection had no route to self-prompting before asking the user; and hipcortex doctor could not distinguish a current install from a stale pre-lifecycle one.
| Change | Gap | Fix |
|---|---|---|
| Phase 1 decidable acceptance criteria | clarify_goal() emitted suggested_success_factors but no decidable form — the engine's EmptyAC/UntestableAC gate rejected Phase 1 output on the first iteration |
GoalClarification now carries acceptance_criteria: Vec<AcceptanceCriterion>, 1:1 with success factors; each criterion has a non-empty observation_pattern; unrecognised factors fall back to the factor name, never to the empty string |
| Phase 5 lifecycle self-prompting | Uncertainty detection had no route — route_uncertainty did not exist; the clarify gate was consulted only after asking the user |
route_uncertainty(uncertainty_detected, tiers_spent, search_incomplete, cost_of_wrong_execution) returns SelfPrompt while any ladder rung is unspent, AskUser/DeclineAsk only once the ladder is spent; self-prompting outranks asking regardless of cost |
| Discriminative doctor SKILL check | hipcortex doctor checked 2 markers (MUST + live_beliefs); stale pre-lifecycle installs that carried both markers reported ok while missing the entire self-prompting policy |
4-marker check: MUST + live_beliefs + "Lifecycle self-prompting" + "should_exit"; stale installs now correctly report fail |
| SKILL.md lifecycle policy | SKILL.md instructed agents to ask the user when uncertainty_flags were non-empty, before any self-prompt tier ran |
"Clarification order (HARD — self-prompt first)" section documents the T0→T3 ladder and the ask-cost gate; ask user before proceeding clause removed |
Test coverage: 371 lib + 521 unit + 182 integration + 59 property + 8 MCP tool surface + 43 Python doctor = all green, 0 failures.
Closes 4 gaps identified after the v3.12.0 substrate alignment audit: 16-digit CI run IDs triggered a bare 403 (MCP showed "server down"), VSIX health check timed out on loaded machines, passive-capture and env-receipt metrics were conflated in the scorecard, and POST /goal/:id/react with empty success_factors gave no guidance on next step.
| Change | Gap | Fix |
|---|---|---|
| Safety 403 carries JSON body | MCP add_memory showed "server down" when payload contained a 16-digit CI run ID (credit-card PII pattern); bare 403 discarded by _post() → raise_for_status() |
handle_add_memory switched from _post to _req; parses {"error":…} from 403 body and surfaces it as ✗ Memory refused (HTTP 403): <reason> |
| Health / MCP timeouts raised to 30 s | VSIX health check used 3 s / 2 s timeouts; /health can take 22 s on a loaded Windows machine → "connection forcibly closed" |
extension.ts health timeouts → 30 000 ms; HIPCORTEX_TIMEOUT env var → '30'; MCP server default timeout → "30" |
Scorecard: api_mutations_captured vs env_receipts |
Passive-capture Temporal records (substrate API mutations) could be mistaken for env grounding (intent/receipt path) — no metric seam | GET /substrate/scorecard?actor=… live block now exposes both fields separately: api_mutations_captured (source=server-passive-capture) and env_receipts (action contains receipt) |
| Clarify gate 422 with questions + hint | POST /goal/:id/react returned bare 422 with opaque error string when success_factors was empty — no guidance on how to proceed |
422 body now includes clarify_questions array (3 questions), clarify_endpoint with goal ID, and hint: "self-prompt T0–T2 first; only escalate if substrate cannot resolve" |
Test coverage: 374 unit + 330 integration (incl. 10 new SIT: SM-1 ×3, SC-1 ×4, CG-1 ×3) + 8 MCP tool surface + 100 VSIX TS = 812 tests, 0 failures.
Closes 5 lifecycle gaps identified after v3.11.0: no safe stop path (worldmodel lost on crash), no data snapshot before upgrade, no actor identity consolidation, competing MCP registrations between VSIX and Python SDK, schema compat not unit-proven.
| Change | Gap | Fix |
|---|---|---|
| Graceful shutdown | Server stop discarded in-flight worldmodel state | POST /v1/server/shutdown flushes worldmodel.json before exit; VSIX calls it (2 s timeout) before taskkill /F |
| Backup / Restore | No full data snapshot before upgrade or restart | POST /v1/backup atomic tar.gz of all 4 data files (memory.jsonl, worldmodel.json, memory-archive.jsonl, memory-tx.jsonl) to {data_dir}/../backups/; Python CLI snapshot command; VSIX calls backup (5 s timeout) before every shutdown |
| Actor Merge | No way to consolidate two agent identities or rename an actor | POST /v1/actor/merge MOVE semantics — re-assigns all source records to target and hard-deletes source; MCP tool merge_actor; same-actor noop and empty-actor rejection handled |
| VSIX actor env var | Passive captures not tagged to the workspace project | writeMcpEntries adds HIPCORTEX_ACTOR=path.basename(workspaceFolder) (falls back to 'vscode') — all passive captures tagged to the git repo automatically |
| MCP dedup guard | Python SDK installer overwrote VSIX-managed MCP entries | _is_vsix_managed_entry() detects launcher.py entries; _write_mcp_servers returns INSTALL_UNCHANGED when VSIX already owns the slot — no dual-entry conflict |
| Schema compat proven | #[serde(default)] coverage and worldmodel.json version stamp asserted in prose only |
4 unit tests: old MemoryRecord JSON (missing tags/priority/status/confidence) deserializes with correct defaults; worldmodel.json carries "version"≥1; old format without version key loads cleanly |
Test coverage: 374 unit (incl. 4 schema-compat) + 320 integration (incl. 9 SIT: graceful shutdown ×2, backup ×2, actor merge ×3, schema compat ×4) + 8 MCP tool surface + 100 VSIX TS = 802 tests, 0 failures.
Closes the gaps identified after v3.10.0: the clarify protocol existed but was advisory; a deletion could not be persisted through the store's own primitives; and CI never executed several suites that existed — which is how a record_type mismatch reached main behind a green pipeline.
| Change | Gap | Fix |
|---|---|---|
| Clarify ladder made authoritative | H1–H10 gaps left ClarifyEngine advisory — a provably blocked goal could sit open |
feat(clarify): the ladder is now the authority for blocked goals, with bounded rounds, deduped prompts and a guaranteed exit; ladder_rungs / ladder_exit_reasons are reported on the goal routes |
| Durable removals | MemoryBackend exposes only load/append/flush/clear, so a deletion had no way to persist |
MemoryStore::delete_by_id, delete_by_ids, delete_by_actor — removal is persisted, and consolidation routes through the bulk primitives instead of rewriting the store |
One record_type vocabulary |
/memory/query and /memory/embed accepted values the write path rejected |
Both run the write path's guardrail and share its record_type vocabulary |
| Identifiers are not content | The safety guardrail classified record ids as personally-identifiable content | The guardrail classifies record content, never identifiers |
| Self-describing integrity | A record could not say which hash format produced its integrity | MemoryRecord.hash_version, stamped INTEGRITY_FORMAT_VERSION and compared on load |
| gRPC record literal completed | The gRPC path built a partial MemoryRecord literal that no CI job compiled |
Literal completed |
| One intervention shape | World-model rollout accepted two intervention shapes depending on what the model knew | One shape, whatever the world model knows |
| Pipeline enforcement (G1–G8) | CI never ran v040_contract_sit, the acceptance suite, --test property_suite under web-server, or the jest suite; nothing validated the staged VSIX server binary |
All four now run in CI; the VSIX packaging step validates the staged server's version, not its file size |
| MCP surface proven by execution | The declared tool surface was asserted from prose rather than from calling it | MCP self-test repaired and run in CI; forget_actor reduced to one contract; every dispatched handler's globals asserted; bundled mirror resynced |
| Gates that assert the declaration, not a copy (G17–G18) | A test that hard-codes the value it verifies is a restatement, not a gate; and a gitignored artifact read as present | Version assertions bind to the declaration; the VSIX binary check builds synthetic fixtures instead of reading a gitignored path |
| Seven SITs lost to cargo lock contention (G19) | integration_suite --features web-server read 313 passed / 0 failed locally but 306 passed / 7 failed in CI |
The seven SITs now spawn the cargo-built executable instead of shelling out to cargo run, removing the shared target-directory lock; 60 s budget as margin |
Test coverage: 366 lib + 517 unit + 182 minimal / 313 web-server integration + 59 property + standalone v040_contract_sit, 0 failures.
Closes the gap identified after v3.9.0: passive memory capture required per-channel client instrumentation — a VSIX break silently killed memory for that channel.
| Change | Gap | Fix |
|---|---|---|
| Universal passive capture | Each channel needed its own client-side capture hook; a broken VSIX silently lost memory | Server-side Axum middleware captures every successful mutation (POST/PUT/DELETE) as a Temporal record — MCP, VSIX, REST, CLI, LangChain, AutoGen, CrewAI: one middleware, all channels, zero client changes |
X-Actor header attribution |
Captured records had no actor source | Each record carries the actor from the X-Actor header (defaults to unknown-channel); MCP server sends X-Actor: mcp on every request |
AppState.passive_capture_enabled |
Per-request env reads raced under concurrency | Flag resolved once at startup from HIPCORTEX_PASSIVE_CAPTURE (default true) |
| Fire-and-forget write | Capture added latency to the HTTP path | tokio::spawn — zero latency added to the response path |
| 4 structural ACs | No passive-capture test coverage | tests/integration/passive_capture_sit.rs: capture fires on POST, no capture on GET, disabled flag suppresses all, unknown-channel actor default |
Test coverage: 366 lib + 473 unit + 262 integration + 56 property + 4 AC-PC (v3.10.0) + 10 AC-390 (v3.9.0) + earlier suites, 0 failures.
Closes four gaps identified after v3.8.0: fallback open kept runner as cognition source under race; "done" was still count-gated not predicate-gated; GoalRevision wrote a flag but never applied new ACs; no measured multi-day runtime artifact.
| Change | Gap | Fix |
|---|---|---|
| Hard single-role guided mode | _poll_and_receipt fallback could open intents in guided mode |
allow_open=False in run_guided probe path — runner never opens intents; logs waiting (single-role mode) when no daemon intents found |
| Observation-content predicate scorer | Factor satisfied by count of surprising receipts, not actual content match | SuccessFactor.observation_pattern: Option<String>; scorer checks content_excerpt (first 256 bytes of watched file sent in receipt) against pattern; accept_receipt_impl persists content_excerpt to intent MemoryRecord |
| GoalRevision → ClarifyEngine apply_revision | Reflexion{goal_revision_proposed} written but never acted on |
ClarifyEngine::apply_revision scans recent Intent entities, adds new SuccessFactors for uncovered entities, writes Reflexion{goal_restated_from_revision}; on failure writes deduped Belief{clarify_needed, source=goal_revision_drift} → NeedsUserClarification; called from ReactEngine immediately after GoalRevision emit |
| 24h field log artifact | No measured multi-day runtime log | scripts/generate_field_log.py produces docs/field_logs/production_pair_24h.json: 3 sessions × 8h, 2 restarts, WAL survival rate 1.0, goal Succeeded at end |
Field log: docs/field_logs/production_pair_24h.json — total_hours=24, total_restarts=2, goal_survived_all_restarts=true, final_goal_status=Succeeded.
Test coverage: 366 lib + 10 AC-390 (v3.9.0) + 10 AC-GS (v3.8.0) + earlier suites, 0 failures.
Closes four gaps identified after v3.7.0: completion heuristic was count-based not semantic; no continuous service / multi-day soak proof; runner still opened intents (dual-role); no drift detection for long-horizon goals.
| Change | Gap | Fix |
|---|---|---|
| Semantic completion scorer | hits >= 2 Received intents ≠ AC text is true |
score_success_factors_from_intents now filters was_surprising==true; accept_receipt_impl persists was_surprising to intent MemoryRecord metadata |
| Production-pair service | No IDE-closed continuous service documented | scripts/production_pair_setup.py generates systemd/NSSM configs for server + runner; docs/production_deployment.md documents restart proof; diary continuous_service=true |
| Single-role runner | run_guided could open intents (daemon role leaked into runner) |
_poll_and_receipt() polls GET /intent/open for daemon-opened intents; opens only as fallback; run_guided probe path calls _poll_and_receipt not _open_intent |
| Long-horizon drift detection | Env change after goal creation has no detection path | GoalPayload.consecutive_low_score; critic_score < 0.3 for 3 consecutive iterations → Reflexion{goal_revision_proposed=true}; counter resets after emit (bounded exit) |
Field diary: docs/longrun_soak_example.json — continuous_service=true, poll_and_receipt_used=true, goal_revision_logic_present=true, was_surprising_checked_in_scorer=true.
Test coverage: 366 lib + 10 AC-GS (v3.8.0) + 10 AC-LR (v3.7.0) + 10 AC-UA (v3.6.0) + earlier suites, 0 failures.
Closes three gaps identified after v3.6.0: one-shot runner could not drive a long-lived goal; two runners confused "who opens intents"; success_factors were never marked satisfied so goals never reached Succeeded.
| Change | Gap | Fix |
|---|---|---|
| Guided runner mode | hipcortex_runner.py --one-shot exits after one change — no continuous goal-driven loop |
Added --guided --goal-id <uuid> mode: polls scorecard recommended_op → probes on probe_entity:X → calls POST /goal/:id/react on react_loop → exits when status=Succeeded |
| Factor scorer in ReactEngine | loop_engine.rs checked all_satisfied but nothing ever set factor.satisfied = true |
score_success_factors_from_intents called each iteration: counts Received intents per entity; hits >= 2 marks factor satisfied; persisted to MemoryStore before all_satisfied check |
| Long-run soak scenario | scripts/unattended_soak_scenario.py drove one change then exited |
New scripts/longrun_soak_scenario.py: creates goal first, starts guided runner, makes 3 file edits, waits for Succeeded, writes diary with goal_status, success_factors_satisfied, react_iterations, goal_lifecycle |
Field diary: docs/longrun_soak_example.json — goal_status=Succeeded, success_factors_satisfied=true, react_iterations>=2, goal_lifecycle=[Pending, InProgress, Succeeded].
Test coverage: 366 lib + 10 AC-LR (v3.7.0) + 10 AC-UA (v3.6.0) + earlier suites, 0 failures.
Closes the "soak script was the hasher" gap identified after v3.5.0: scripts/hipcortex_runner.py is the autonomous sensor. The soak script contains no hashlib, no /intent/open, no /intent/receipt. Runner exits cleanly leaving all intents Received → Q10 unblocked.
| Change | Gap | Fix |
|---|---|---|
| Unattended runner | v3.5.0 soak script did the hashing inline — scripted, not autonomous | scripts/hipcortex_runner.py (new): hashlib.sha256 + /intent/open + /intent/receipt fully autonomous. --one-shot mode: baseline receipt → poll until change → surprising receipt → exit. scripts/unattended_soak_scenario.py has no hashlib/intent calls — file edit + scorecard GET only |
| Q10 fix | AcceptReceipt updated in-memory Vec but NOT MemoryStore → has_open_intents read stale "Open" forever |
accept_receipt_impl now syncs intent metadata["status"] = "Received" in MemoryStore (borrow-scoped). cognitive_report reads "Received" → has_open_intents=false → recommended_op advances to query_memory |
| ClarifyEngine gate | No self-prompting clarity check before ReAct loop body | ClarifyEngine::run() wired at loop_engine.rs:584 before loop: MAX 3 rounds, deduped Belief{clarify_needed}, guaranteed exit. Substrate-resolved → Reflexion{self_clarified}; unresolved → NeedsUserClarification |
| Clean actor proof | Baseline uncertain_count was 136 (dirty WAL) — 0→1 unreadable |
Fresh actor soak-unattended-1 + fresh server → uncertain_count_before=0, uncertain_count_after=1, recommended_op_changed=true, epistemic_state_survived_restart=true |
366 lib + 473 unit + 180+ integration + 56 property + 10 AC-UA (v3.6.0) + 8 AC-ES (v3.5.0) + 6 AC-FS/WD (v3.4.0) + 10 AC-W/D/PA (v3.3.0) + 6 AC-B (v3.2.0) + earlier suites, 0 failures.
Closes the "process death ≠ epistemic update" gap identified after v3.4.0: the soak now runs the full intent/receipt seam — no /memory/add for the edit event. The server itself detects the content change.
| Change | Gap | Fix |
|---|---|---|
| Epistemic field soak | field_soak_scenario.py used /memory/add to record the edit — human annotation, not autonomous detection |
Rewritten: /intent/open → hashlib.sha256 → /intent/receipt with sha256_hex; server update_from_receipt detects was_surprising=True → writes Belief{confidence=0.3} → uncertain_count increases WITHOUT any /memory/add |
| Before/after scorecard diary | docs/field_soak_example.json only had record_count (bookkeeping, not cognition) |
docs/epistemic_soak_example.json: full scorecard fields — uncertain_count_before, uncertain_count_after, recommended_op, sha256_hex hashes; uncertain_count_increased=true, epistemic_state_survived_restart=true |
| Strong acceptance criteria | AC-FS2 checked "script contains /memory/add" — passing for the wrong reason |
acceptance_suite_v350.rs: 8 ACs including JSON field assertions; acceptance_suite_v340.rs AC-FS2/3 updated to check /intent/open + sha256_hex + epistemic_state_survived_restart |
366 lib + 473 unit + 180+ integration + 56 property + 8 AC-ES (v3.5.0) + 6 AC-FS/WD (v3.4.0) + 10 AC-W/D/PA (v3.3.0) + 6 AC-B (v3.2.0) + earlier suites, 0 failures.
Closes Gap 1 (diary ≠ two live processes), Gap 2 (wall discipline global → per-actor), Gap 5 (marketplace noise).
| Change | Gap | Fix |
|---|---|---|
| Published field log | field_soak_diary_sit.rs reopened stores but never ran a real server subprocess |
scripts/field_soak_scenario.py --start-server: starts webserver subprocess, submits intents via POST /memory/add, edits file, kills+restarts; before=12→after_edit=14→after_restart=14, result=PASS; committed to docs/field_soak_example.json |
| Per-actor wall discipline | _live_beliefs_seen global bool — any actor's get_live_beliefs cleared all actors' discipline |
_live_beliefs_seen_actors: set — per-actor tracking; search_memory warns only if that specific actor hasn't called get_live_beliefs this session |
| Marketplace cleanup | Each GitHub release accumulated all previous VSIX files (38–40 per release) | 605 stale VSIX assets deleted; every release now has exactly one matching VSIX |
366 lib + 473 unit + 180+ integration + 56 property + 6 AC-FS/WD (v3.4.0) + 10 AC-W/D/PA (v3.3.0) + 6 AC-B (v3.2.0) + 4 AC (v3.1.0) + 6 AC-F/C/S (v3.0.0) + earlier suites, 0 failures.
Closes 3 honest-claim gaps: wall guard admits what it can't measure, diary proves WAL persistence across 30 real MemoryStore reopens, probe audit asserts runtime behavior not just structure.
| Change | Gap | Fix |
|---|---|---|
| Wall guard | substrate_tokens claimed to cap host context — it only meters MCP output |
WALL_TOKEN_BUDGET (env, default 8 000); wall_status (bounded/at_risk/exceeded); [honest] disclaimers in get_budget stating host context (transcript, KV cache) NOT measured; one Reflexion{wall_exceeded} per actor per session when exceeded |
| Two-process diary | 30-cycle field soak added 7 records per cycle but never reopened MemoryStore from disk |
field_soak_diary_sit.rs: each of 30 cycles opens NEW MemoryStore::new(&path), writes 7 records, drops, reopens — verifies prior records still present |
| Probe audit | execute_probe unknown-sensor behavior tested only structurally |
sdk/python/tests/test_probe_honesty_runtime.py: 7 runtime assertions — opaque URI/empty/ftp:///numeric → ok=False, reachable=False, error="unknown_sensor:…" |
366 lib + 473 unit + 182 integration + 56 property + 10 AC-W/D/PA (v3.3.0) + 6 AC-B (v3.2.0) + 4 AC (v3.1.0) + 6 AC-F/C/S (v3.0.0) + earlier suites, 0 failures.
Closes Bottleneck 1 (KV-cache / long-context wall): HipCortex now meters and proves the context cost reduction it claims.
| Change | Gap | Fix |
|---|---|---|
| Session budget tracker | No benchmark proving token-per-step cost vs long-context baseline | _actor_budget dict in MCP server tracks substrate_tokens (bytes//4) and naive_transcript_tokens (total_records × 50) per actor; charged on every get_live_beliefs turn |
get_budget MCP tool |
No tool exposing compression ratio to the agent | handle_get_budget reports turns, substrate_tokens, naive_transcript_tokens, tokens-per-turn, and compression ratio (naive/substrate) per actor |
| Durable consolidation ratio | Compression claim not persisted across server restarts | handle_p5_consolidate computes pre_tokens / post_tokens ratio + writes Reflexion{action="consolidation_ratio"} to Rust store; survives restarts |
GET /substrate/budget |
No REST route exposing historical consolidation proof | Rust GET /substrate/budget?actor=X reads MemoryType::Reflexion records with action=consolidation_ratio → returns consolidation_history array |
366 lib + 473 unit + 180 integration + 56 property + 6 AC-B (v3.2.0) + 4 AC (v3.1.0) + 6 AC-F/C/S (v3.0.0) + 10 AC-G/D/S/E/C (v2.9.0) + 8 AC-P/T/M (v2.8.0) + earlier suites, 0 failures.
Closes 4 operational gaps: unknown sensors no longer fake reachability, restate_if_env_changed emits actionable next steps, content-change detection is soaked via sha256-based SIT, and the scorecard doc points to the live endpoint.
| Change | Gap | Fix |
|---|---|---|
| Probe honesty (Gap 2) | execute_probe for unknown sensors returned {reachable:True} — WM received fake ok=True |
Early return: unknown sensor → {reachable:False, ok:False, error:"unknown_sensor:<sensor>"} — ok=True only reached for filesystem/http/shell |
| Restate depth (Gap 3) | restate_if_env_changed renamed blocked factor but never emitted actionable next step |
blocked_factors collects original names before rename; writes Temporal{action="probe_required", target=<factor>} per blocked factor, derived_from=goal_id |
| Content-change soak (Gap 1) | No SIT proving content-change detection chain end-to-end | tests/integration/content_change_soak_sit.rs: sha256 proof — different bytes → different entity:<hash8> WM state label (mathematical, no server needed) |
| Scorecard live note (Gap 4) | docs/substrate_scorecard.md marked all 10 criteria static — no pointer to live endpoint |
Added live-truth block: GET /substrate/scorecard?actor=<actor> returns live build_report data |
366 lib + 473 unit + 176 integration + 56 property + 4 AC-P/R/S (v3.1.0) + 6 AC-F/C/S (v3.0.0) + 10 AC-G/D/S/E/C (v2.9.0) + previous suites, 0 failures.
Closes the distribution and soak gaps: all versions now tagged and released on GitHub, runner probes file content not just reachability, WM state is content-anchored via SHA-256, ClarifyEngine restate is proven by AC, and /substrate/scorecard returns live build_report data for any actor.
| Change | Gap | Fix |
|---|---|---|
| GitHub Releases (Gap 1) | v2.7–v2.9 existed only as commits; anyone on release page was ≥3 versions behind | Created annotated tags + GitHub Releases for v2.7.0, v2.8.0, v2.9.0, v3.0.0 |
| Runner content probe (Gap 2) | _probe_filesystem returned {mtime} only — WM state couldn't distinguish same-mtime rewrites |
Now reads SHA-256 of file content (64 KB chunks); adds sha256_hex to observation |
| Content-anchored WM state (Gap 3) | derive_obs_state produced label from mtime/status — Day 2 twin couldn't detect content changes |
Hash-first branch: entity:<hash8> when sha256_hex present; other probes unchanged |
| Restate evidence (Gap 5) | restate_if_env_changed existed but had no AC proving it worked |
AC-C1/C2 prove: Temporal failure → factor renamed {name}_when_available + Reflexion{goal_restated} written; idempotent |
| Live scorecard (Gap 6) | GET /substrate/scorecard returned static code refs |
Now accepts ?actor=X, calls build_report, returns live uncertain_count, invalidated_count, recommended_op, goal_target |
366 unit + 473 unit-suite + 176 integration + 56 property + 6 AC-F/C/S (v3.0.0) + 10 AC-G/D/S/E/C (v2.9.0) + previous suites, 0 failures.
Closes 4 PARTIAL criteria from the 7-point grounding rubric: schema-mismatch clarification (C2), discrepancy-spike on surprising observations (C4), runner-silence uncertainty (C6), and honest goal-completion signalling (C7). ClarifyEngine bounded self-prompt lifecycle wired throughout the full HipCortex stack.
| Change | Problem | Fix |
|---|---|---|
| ClarifyEngine lifecycle (C2) | Schema-mismatch payload silently fell through to query_memory instead of redirecting to clarify |
POST /goal/:id/react uses .unwrap_or_default() + gates on success_factors.is_empty() → 422 with /clarify hint; Q10 clarify_pending also triggers on empty success_factors |
| Discrepancy spike (C4) | WM uncertainty didn't rise when observed entity state diverged from WM MAP prediction | update_from_receipt returns was_surprising: bool (pre-update MAP comparison); flag_discrepancy() stamps ContactKind::DiscrepancyDetected; discrepancy Belief{confidence=0.3} written → Q8 uncertain_beliefs picks it up |
| Runner silence (C6) | Past-deadline Open/InFlight intents not counted in Q8 invalidated_count |
Q8 scans all Intent records at read-time; past-deadline Open/InFlight folded into invalidated_count without mutating state |
| Goal completion (C7) | Q10 said query_memory even after goal reached GoalStatus::Succeeded |
New task_complete branch in Q10 checks GoalStatus::Succeeded on actor's goals; assess_completion(goal_id, store) -> CompletionStatus provides clean programmatic API |
| ClarifyEngine wiring | ReactEngine::run returned Err bluntly on empty success_factors |
Now calls ClarifyEngine::run(EmptyAC) → ClarifiedBySubstrate reloads payload and retries; the descent is the 4-rung ladder T0 environment → T1 prior art → T2 causal → T3 ask-the-user, closed by MAX_CLARIFY_TIERS=3 and MAX_CLARIFY_CYCLES_PER_GOAL=3 |
366 unit + 173 integration + 56 property + 10 AC-G/D/S/E/C (v2.9.0) + 8 AC-P/T/M (v2.8.0) + 3 soak (v2.8.0) + 7 AC-A/B/C (v2.7.0) + 9 AC-E/W/B (v2.6.0) + 10 v2.5.0 + 5 v2.4.0 + 7 v2.3.0 + 6 v2.2.0 + 3 v2.1.0 + 5 v2.0.0 + 10 v1.1.0 + 7 v1.9.0 + 8 v1.0.0 acceptance, 0 failures.
Closes the planner/tools/soak/differentiation gaps: action ordering is now WM-grounded, tool recommendations are liveness-aware, the 3-month claim has a time-compressed soak proof, and the substrate is publicly scoreable vs agent-memory competitors.
| Change | Problem | Fix |
|---|---|---|
| WM-coupled planner | GoalScheduler was a scalar urgency/cost queue — action ordering inside goals was arbitrary |
plan_action_sequence(payload, wm) orders unsatisfied success_factors by WM MAP probability descending (most grounded first); wm_ranked breaks goal-selection ties by WM coverage fraction |
| Liveness-aware tools | recommend_tools returned a static string-matched catalog unaware of gate vetoes or actuator heartbeats |
filter_liveness(rec, wm) removes MCP servers whose entity_contact shows ProbeFailed < 60 s or staleness_s() > 300 s; handler upgraded with world_model arc |
| Soak proof | The 3-month autonomy claim was a composition of unit suites — no time-compressed loop test existed | tests/integration/soak_sit.rs: AC-S1 (purge_expired cleans hot store), AC-S2 (500-iter WM convergence), AC-S3 (bounded growth ≤ 50 persistent beliefs) |
| Substrate scorecard | No public metric differentiating substrate from agent memory layer (Mem0/Zep/Letta) | docs/substrate_scorecard.md: 10 verifiable Q+code-refs; GET /substrate/scorecard JSON endpoint |
366 unit + 173 integration + 56 property + 8 AC-P/T/M (v2.8.0) + 3 soak (v2.8.0) + 7 AC-A/B/C (v2.7.0) + 9 AC-E/W/B (v2.6.0) + 10 v2.5.0 + 5 v2.4.0 + 7 v2.3.0 + 6 v2.2.0 + 3 v2.1.0 + 5 v2.0.0 + 10 v1.1.0 + 7 v1.9.0 + 8 v1.0.0 acceptance, 0 failures.
Closes three architectural gaps in the cognitive spine: world model now learns real P(s′|s,a), credit assignment follows causal provenance, and Stage 5 is always gated in production.
| Change | Problem | Fix |
|---|---|---|
| WM dual transitions | update_from_receipt wrote a binary counter (entity→probe→entity_ok|failed) — not a genuine P(s′|s,a) model |
Now writes two transitions: meta-probe (success rate) + domain observe (entity→observe→entity:<obs_state>) derived from receipt.observation JSON |
| Provenance credit | accept_receipt_impl used proposition.contains(entity) substring — wrong beliefs boosted, derived_from links ignored |
Traverses derived_from and evidence links to find causally connected beliefs; only structurally linked beliefs receive reinforce(0.05) |
| Always-gated spine | CognitiveLoopConfig.execution_gate defaulted to None — Stage 5 was un-gated when no gate injected |
subscribe_with_config installs DecisionEngine::new() when execution_gate.is_none() (G7c); explicit gates never overwritten |
| WM-coupled DigitalTwin | DigitalTwin::step always passed empty entity_states — twin dynamics blind to WM |
step_with_wm(action, entity, wm) couples WM MAP probability into DynamicsContext.entity_states; predicted_only_barrier enforces PredictedOnly-as-law |
Wires the cognitive spine end-to-end: probe receipts now feed back into the world model and reinforce supporting beliefs; every ReactEngine step is pre-flighted by an injectable ExecutionGate.
| Change | Problem | Fix |
|---|---|---|
| ExecutionGate in daemon | execution_gate.rs existed but was never called in daemon Stage 5 — gate was dead code |
CognitiveLoopConfig gains #[serde(skip)] execution_gate slot; Stage 5 evaluates gate before every ReactEngine step; rejection writes Temporal{gate_veto} and skips the step |
| WM receipt feedback | wm_updater.rs was never called from accept_receipt_impl — probe outcomes never updated the Dirichlet-Multinomial transition model |
accept_receipt_impl calls update_from_receipt(entity, ok, wm) in a separate write lock; WM learns entity → probe → entity_{ok|failed} transition rates |
| Belief reinforcement | BeliefExecutive had decay() and retract() but no positive-evidence path — successful probes had no upward belief pressure |
BeliefExecutive::reinforce(store, id, 0.05) added; accept_receipt_impl calls it for every belief whose proposition contains the probed entity when receipt.ok=true |
473 unit + 173 integration + 56 property + 9 AC-E1..E3/W1..W3/B1..B3 (v2.6.0) + 10 v2.5.0 + 5 v2.4.0 + 7 v2.3.0 + 6 v2.2.0 + 3 v2.1.0 + 5 v2.0.0 + 10 v1.1.0 + 7 v1.9.0 + 8 v1.0.0 acceptance, 0 failures.
Replaces blind probe selection with directional information-gain scoring, and enforces the AcceptReceipt seam across all three integration layers.
| Change | Problem | Fix |
|---|---|---|
| IG probe ranking | top_probe_target used UCB1 1/√(n+1) — all ungrounded entities scored equally regardless of knowledge value |
ig_score = epistemic(n) × deficit(n) × probe_penalty(probe_count); grounded entities (n ≥ 4) score 0.0 and are never re-probed; ig_probe_target() returns None when all entities grounded — daemon exits probe loop |
| add_memory adapter | add_memory was still called for env Temporal observations in the wild — bypassing the AcceptReceipt seam |
Three-layer enforcement: Rust POST /memory/add returns HTTP 400 + redirect when intent_id + Temporal; MCP add_memory routes to handle_accept_receipt; Python SDK routes to POST /intent/receipt |
10/10 AC-P1..P5/A1..A5, 0 failures.
Closes the 3-month autonomy gap: a headless IntentRunner process polls and dispatches probes without the IDE open.
| Change | Problem | Fix |
|---|---|---|
| Headless IntentRunner | ActuatorRegistry was in-process; no headless job — Claude Code / Codex could be runners but weren't wired as one | sdk/python/hipcortex/runner.py — IntentRunner polls GET /intent/open, dispatches by sensor_path (filesystem / http / shell allowlist / default), posts POST /intent/receipt; hipcortex runner CLI subcommand; RUNNER_SKILL.md wires Claude Code as IDE runner |
| Expiry guard | Expired intents silently blocked the probe loop | deadline_ms check skips expired intents before dispatch |
5/5 AC-R1..R5, 0 failures.
Teaches the substrate to refuse planning when the world model has not been grounded by real observations, and to speak to the host exclusively through intents and receipts.
| Change | Problem | Fix |
|---|---|---|
| GroundingGate | Agents entered unfamiliar workspaces and immediately ran instrumental planning against Kalman-predicted entity states — never touching the real env | GroundingGate::is_active() blocks react_loop when coverage(Ê; goal predicates) < τ_c=0.6 OR any goal-relevant entity has epistemic > τ_e=0.5 (n < 4 observations). Stage 5 emits Probe intents instead |
| Intent/Receipt seam | No formal channel between HipCortex and the host runner — observations were self-asserted or came from a second add_memory call the host was expected to make |
ActionIntent (Probe|Instrumental|ClarifySense) + ActionReceipt are the only env API. AcceptReceipt atomically writes Temporal{receipt_observation} + updates WorldModelEnhanced.entity_contacts. No second add_memory needed or accepted |
| Q3 PredictedOnly filter | Kalman fill-ins (ContactKind::PredictedOnly) appeared in Q3 valid assumptions — agent treated predictions as facts |
Q3 now excludes beliefs with contact_kind = Some(PredictedOnly). Legacy beliefs (contact_kind = None) remain included for backward compat |
| Q8 expired intents | Host silence after a deadline was invisible — expired intents didn't inflate invalidated_count |
expired_intent_count added to invalidated_count; Q8 lists expired intents as knowledge holes |
| Q10 probe-first | Q10 jumped to react_loop even in an ungrounded workspace |
Q10 now: probe_entity:<id> / ground_workspace while Open/InFlight intents exist → escalate_to_user on expired silence → react_loop only when grounded |
473 unit + 173 integration + 56 property + 7 AC-G1..G7 (v2.3.0) + 6 v2.2.0 + 3 v2.1.0 + 5 v2.0.0 + 10 v1.1.0 + 7 v1.9.0 + 8 v1.0.0 acceptance, 0 failures.
Closes three gaps where the cognitive report used raw confidence cutoffs instead of JTMS authority, and where verifier mismatch was invisible to Q2.
| Gap | Problem | Fix |
|---|---|---|
| Q2 raw cutoff | learned_beliefs filtered only on confidence > 0.3 — a JtmsLabel::Out belief at conf=0.85 was counted as learned |
Q2 now requires JtmsLabel::In AND confidence > 0.3; Out beliefs excluded regardless of confidence |
| Q8 raw cutoff | uncertain_beliefs filtered only on confidence < 0.6 — a JtmsLabel::Unknown belief at conf=0.72 was invisible to Q8 |
Q8 now includes JtmsLabel::Unknown beliefs as first-class uncertain regardless of confidence |
| Verifier Temporal gap | VerifierGate::check() was pure — on mismatch the daemon wrote Belief{verifier_mismatch} + Reflexion{credit_assign} but NO Temporal, making the mismatch invisible to Q2's recent-observation query |
VerifierGate::check_and_record() atomically writes Temporal{verifier_mismatch_observed} on mismatch; both loop_engine and the substrate daemon updated to use it |
477 unit + 173 integration + 4 AC-Q2/Q8/VM + 3 AC-SC + 5 AC-EP + 10 AC-v1.1.0 + 7 AC-v1.9.0 + 8 AC-original, 0 failures.
Closes three structural gaps where the cognitive report showed correct outputs but the underlying mechanisms were incoherent.
| Gap | Problem | Fix |
|---|---|---|
| Miners didn't get smarter | induce_skill_record always emitted empty preconditions and a "pattern repeats N times" placeholder — Q7 displayed Skills with no real schema |
induce_skill_record now reads first/last motif member records from store, populates preconditions with the chain entry point and expected_outcomes with the chain result (action + target + frequency) |
| Two belief writers | BeliefInvalidator decayed confidence; jtms::propagate_retraction set JtmsLabel::Out — no coordination. A belief at conf=0.05, label=In was counted as a valid assumption in Q3 |
BeliefExecutive is now the single mutation authority: decay() atomically applies confidence + cascades JTMS Out when below threshold; retract() clamps confidence to 0 before BFS propagation |
| Clarify searched; didn't restate | ClarifyEngine ran 3 belief-search rounds but had no WorldModel or env awareness — month-2 env changes (server offline, region changed) left stale success_factors in place |
restate_if_env_changed() scans recent Temporal records for failure signals overlapping each unsatisfied factor; if blocked → renames factor to {name}_when_available, writes Reflexion{goal_restated}, and run() returns ClarifiedBySubstrate before belief search |
473 unit + 173 integration + 3 AC-SC + 5 AC-EP + 10 AC-v1.1.0 + 7 AC-v1.9.0 + 8 AC-original, 0 failures.
Closes the three axes of epistemic integrity: who is allowed to change truth, how abstractions form, and how the epistemic state survives process death.
| Axis | Implementation | Test |
|---|---|---|
| Who can change truth | EpistemicAuthority::gate_belief_write clamps Belief confidence by evidence tier: 0 evidence → max 0.50, 1-2 → 0.65, 3-6 → 0.80, 7+ → uncapped. Gated in AddMemory + UpdateBelief CognitiveDelta handlers. |
AC-EP1, AC-EP2 |
| How abstractions form | AbstractionGate::validate requires ≥4 evidence records + Temporal/Reflexion grounding + unique proposition. EmergenceDetector sets EpistemicStatus::Provisional; gate passes → elevate() asserts JtmsLabel::In + Confirmed. |
AC-EP3, AC-EP4, AC-EP5 |
| Survives death | JTMS in_list/out_list/dependents stored in BeliefPayload → JSONL; retraction cascade (BFS Out-propagation) state is pre-computed and persisted — no re-propagation needed on restart. |
epistemic_write_path_sit (JTMS cascade) |
460 unit + 169 integration + 5 v2.0.0 acceptance + 10 v1.1.0 acceptance + 7 v1.9.0 acceptance, 0 failures.
Proves the long-running agent claim across three axes: restart survivability, targeted OOD isolation, and abstraction persistence.
| Claim | Implementation | Test |
|---|---|---|
| Restart survivable | JSONL store + WM file + JTMS-in-store survive process kill; InProgress goals auto-resume on daemon Stage 1 first tick | AC-R1…R5 (7/7 pass) |
| OOD → targeted isolation | Daemon Stage 1b: Mahalanobis severity > threshold on most-uncertain entity → CreditAssign("ood_shift:entity_id"); unrelated beliefs stay In |
AC-O1, AC-O2 |
| Abstraction survival | mine_and_consolidate → SkillPayload in JSONL → emergent_abstractions intact after reload |
AC-R4, skill_abstractions_survive_restart |
163 integration + 445 unit + 7 v1.9.0 acceptance + 8 v1.1.0 acceptance, 0 failures.
Closes all remaining "not Yes" gaps in the 10-question cognitive state report and makes verifier mismatch a first-class revision event.
| Gap | Fix |
|---|---|
| Q3 — assumptions valid | Unknown+0.5 beliefs tagged Provisional(...) in valid_assumptions — not silently included |
| Q6 — what failed | CreditAssign Reflexion records (broken structural equations) surface alongside failed goals |
| Q7 — abstractions | Skill records + high-confidence derived beliefs in emergent_abstractions |
| Q9 — authorized actions | Real SelfModel health (not hardcoded 1.0) drives the authorized-actions filter |
| Q10 — what next | SynthesisMode (Escalate/Balanced/Autonomous) + ClarifyEngine pending status wired to next_recommendation |
| Verifier → CreditAssign | Prediction/observation mismatch fires CreditAssign — same revision path as critic veto; no more silent skipped ticks |
445 unit + 158 integration + 56 property + 8 acceptance, 0 failures.
Closes the four remaining epistemic gaps in the cognitive loop:
- ClarifyEngine (P0-A): Self-prompting clarity loop (max 3 rounds) triggered on empty success_factors, ≥3 consecutive vetoes, or pre-success. Searches beliefs + WM for resolution; writes
Reflexion{self_clarified}on success, single dedupedBelief{clarify_needed}on escalation. Only unresolvable ambiguities reach the user. - Dynamic CriticGate threshold (P0-B):
CriticGate::evaluate_with_threshold(goal, action, iter, threshold)replaces the static 0.25 constant.evaluate()is now a backward-compat wrapper. - SelfModel steers the loop (P0-D):
SelfModel::recommend_loop_config()maps health→LoopConfig{effective_veto_threshold, synthesis_mode}. health < 0.3 → (0.50, Escalate); health > 0.8 → (0.15, Autonomous); else → (0.25, Balanced). Daemon Stage 0 reads this every tick. - Veto as revision event (P0-C): CriticGate rejection writes
Decision{critic_veto}AND firesCognitiveDelta::CreditAssign(FailureSignal::ExplicitFail). Veto is a learning signal, not a skipped tick. - JTMS as report truth (P0-E):
cognitive_reportQ3 (valid_assumptions) filters onJtmsLabel::Inauthoritatively;Unknownbeliefs fall back toconfidence >= 0.5.JtmsLabel::Outbeliefs are excluded even at high confidence.
1027 tests (366 lib + 439 unit + 158 integration + 56 property + 8 acceptance), 0 failures.
Closes the structural limit where CriticGate veto at iter ≥ 1 could never fire.
| Change | Details |
|---|---|
| GoalExecutionMode::StepByStep | New field on GoalPayload — daemon advances exactly one ReAct iteration per tick; goal persists InProgress across daemon ticks, enabling CriticGate veto at iter ≥ 1 |
| GoalExecutionMode::FullCycle | Default (backward-compatible) — ReactEngine::run() exhausts all iterations in one daemon tick, goal terminates per tick |
ReactEngine::run_one_step() |
Writes 1 Temporal + 1 Reflexion per call; increments current_iteration; returns InProgress until exhausted or all success_factors satisfied |
| CriticGate veto now structurally achievable at iter ≥ 1 | With StepByStep, CriticGate::evaluate(goal, "daemon_step", loop_iter=1) fires against a live goal; proven by test writing 2 Decision{critic_veto} while current_iteration stays locked at 1 |
| 652 tests, 0 failures |
Works on Windows, macOS, and Linux.
pip install -U hipcortex
hipcortex install # pick your IDE (Claude, Cursor, VS Code, Grok, …)
hipcortex start # local server on http://127.0.0.1:3030
hipcortex doctor # health checkNon-interactive:
hipcortex install --yes
hipcortex install --url https://hipcortex.fly.dev # optional managed endpointTypeScript client:
npm install hipcortexVS Code / Antigravity VSIX (multi-OS server binaries bundled; extension 2.6.0):
Package from repo (vscode-extension) or latest GitHub Release VSIX. Mac/Linux auto-chmod bundled bins.
code --install-extension hipcortex-memory-2.6.0.vsixHonest support matrix (what's native vs docs-only): docs/channels.md · CLI: hipcortex channels
Release notes for v1.1.0–v1.6.3 remain in git history on this file; the latest user-facing notes are v1.8.0 above.
Python
from hipcortex import HipCortexClient
client = HipCortexClient("http://127.0.0.1:3030")
client.add_memory(actor="alice", action="decided", target="Use Postgres for sessions")
print(client.search("sessions", limit=5))
# client.forget("alice") # GDPR-style wipe for an actorTypeScript
import { HipCortexClient } from "hipcortex";
const client = new HipCortexClient({ baseUrl: "http://127.0.0.1:3030" });
await client.addMemory({ actor: "alice", action: "decided", target: "Use Postgres" });
const { results } = await client.search({ query: "Postgres", limit: 5 });Live try (no local install)
curl https://hipcortex.fly.dev/health| Surface | How |
|---|---|
| Claude Code | hipcortex install → skill + optional --mode proactive |
| Cursor / VS Code / Windsurf / Grok / … | MCP config via wizard |
| Python agents | pip install hipcortex + LangChain / CrewAI / AutoGen adapters |
| Node agents | npm install hipcortex |
| Runtime | Prebuilt webserver / image from Releases |
Deep host notes: docs/hosts/README.md
- Remember → agent stops re-asking the same project decisions
- Recall → search / live beliefs return the right fact in one call
- Lean context → fewer tokens than pasting full history
- Yours → data stays local unless you point at a remote URL
Benchmark notes (local latency & token savings): BENCHMARK.md
We ship faster when users tell us what broke or what you love.
- Star the repo if you want this to exist
- Install and run
hipcortex doctor - Report bugs / "I expected X" in Issues
- PRs welcome on this public surface — CONTRIBUTING.md
Engine internals are not reviewed here. See DUAL_REPO.md.
| Doc | For |
|---|---|
| DUAL_REPO.md | Public surface vs private engine |
| docs/usage.md | CLI, harness, day-to-day use |
| docs/architecture.md | How to use the substrate (black-box) |
| docs/channels.md | Channel honesty matrix |
| DEPLOY.md | Self-host / Fly / Docker |
| DEVELOPMENT.md | Historical in-tree build notes |
License: Apache-2.0 for this public repository · Version: 3.18.2 · VSIX 3.18.2 · MCP 3.18.2