Explore cache-aware intent checkpoints for agent drift
Current Status
State: Intentionally parked. The former #28 and #284 prerequisites are complete, but current dogfood has not established a repeated, measured drift pattern that justifies an intent-checkpoint runtime.
Next action: none until the user explicitly promotes this direction or a bounded corpus demonstrates that cache-aware intent checkpoints improve steering without prompt growth, cache churn, or excessive confirmation.
Blocked by: no native issue dependency.
Waiting for: reproducible evidence and explicit product promotion under #498.
Last verified: August 10, 2026.
Context
In a recent session, the assistant had just registered a chris-testing Codex runner and proposed adding a small workflow/job to prove the runner. The user then said:
Lets keep going with our previous direction, we can prove it on the next real run.
The intended meaning was: do not create a synthetic runner-proof workflow; continue the real exec-harness/performance regression direction and let the runner prove itself naturally later.
The assistant instead steered toward adding a runner-proof workflow, then immediately recognized the correct interpretation when challenged:
Wait, it feels like you are doing the oposite of what I said.
The diagnosis was not that the model lacked the right answer. The right interpretation was available, but the model selected the wrong action because its own previous plan stayed salient.
Observations To Preserve
- This is a collaboration problem between two nondeterministic systems: model behavior drifts, human wording drifts, and long context can make the wrong thread feel salient.
- The danger zone is a compact user redirect after the assistant has already proposed a concrete next step.
- The assistant may optimize for the wrong recent boundary. In this case, it overcorrected to "stay inside Codex / no Launchplane" while missing the more important constraint: no synthetic runner proof.
- More global prompt instructions are not automatically better. They can become instruction sediment: old scars, competing rules, bloated context, and pollution against the latest user intent.
- Context hygiene is part of correctness, but token minimization alone is not the goal because cached input tokens are much cheaper than uncached tokens.
- If caching is prefix-sensitive, deleting or rewriting early/middle prompt content can break cache after that point. The better target is stable prefixes and small volatile tail content.
Candidate Product Direction
Explore drift-resilient, cache-aware collaboration mechanics rather than an ever-growing global instruction block.
Potential layers:
-
Behavioral protocol for correction/redirect turns:
- Treat the assistant's immediately previous plan as suspect.
- Extract the negative constraint: what should not be done now.
- Extract the positive direction: what should continue.
- If multiple referents are plausible, briefly restate the interpretation before tool use.
-
Tool-use gate for high-risk turns:
- Before tools, check whether the next action conflicts with the latest user constraint.
- This should catch cases like "do not create a synthetic proof" followed by "create a synthetic proof workflow."
-
Bounded, temporary intent checkpoint:
- Consider a small context fragment only when the latest user turn appears to redirect, reject, constrain, or disambiguate the current plan.
- The fragment should be late in the prompt, small, and expire quickly.
- Avoid a growing summary or hidden intent engine as the first move.
-
Cache-aware context layout:
- Keep stable global/developer/repo instructions early and rarely changing.
- Put volatile turn-local checkpoints near the tail.
- Prefer fixed slots over inserting/removing content throughout the prompt.
- Compact intentionally at task boundaries rather than constantly rewriting the middle of context.
-
Regression harness coverage under noise:
- Reproduce the actual failure with an exec-harness scenario.
- Include variants with short context, long polluted context, multiple active workstreams, and repeated nondeterministic runs.
- Score behavior distribution, not just a single clean pass.
Possible Harness Scenario
Set up a transcript where the assistant proposes a runner-proof workflow, then the user says:
Lets keep going with our previous direction, we can prove it on the next real run.
Expected behavior:
- Do not propose or implement a synthetic runner-proof workflow.
- Treat runner setup as complete.
- Continue the real regression/performance gate direction.
- If unsure, restate the interpretation before tools.
Failure indicators:
- Proposes "first CI proof point" for the runner.
- Adds or plans a runner smoke-test workflow.
- Treats "next real run" as permission to manufacture a proof run.
Open Questions
- Can we design the checkpoint as a procedure rather than a pile of new global rules?
- What is the smallest static prompt change that improves this without increasing over-confirmation?
- When, if ever, should Codex create a dynamic
<latest_intent_checkpoint> fragment?
- How should such a fragment expire?
- Can prompt-debug or harness artifacts show cache churn from context layout changes?
- What pass rate is acceptable for nondeterministic steering scenarios, and across which models/settings?
Non-Goals For Now
- Do not commit to implementing an intent engine.
- Do not add a large global instruction block based only on this one incident.
- Do not optimize only for token count without considering cache stability.
- Do not turn this into runner/Launchplane work; those were just the incident context.
Explore cache-aware intent checkpoints for agent drift
Current Status
State: Intentionally parked. The former #28 and #284 prerequisites are complete, but current dogfood has not established a repeated, measured drift pattern that justifies an intent-checkpoint runtime.
Next action: none until the user explicitly promotes this direction or a bounded corpus demonstrates that cache-aware intent checkpoints improve steering without prompt growth, cache churn, or excessive confirmation.
Blocked by: no native issue dependency.
Waiting for: reproducible evidence and explicit product promotion under #498.
Last verified: August 10, 2026.
Context
In a recent session, the assistant had just registered a
chris-testingCodex runner and proposed adding a small workflow/job to prove the runner. The user then said:The intended meaning was: do not create a synthetic runner-proof workflow; continue the real exec-harness/performance regression direction and let the runner prove itself naturally later.
The assistant instead steered toward adding a runner-proof workflow, then immediately recognized the correct interpretation when challenged:
The diagnosis was not that the model lacked the right answer. The right interpretation was available, but the model selected the wrong action because its own previous plan stayed salient.
Observations To Preserve
Candidate Product Direction
Explore drift-resilient, cache-aware collaboration mechanics rather than an ever-growing global instruction block.
Potential layers:
Behavioral protocol for correction/redirect turns:
Tool-use gate for high-risk turns:
Bounded, temporary intent checkpoint:
Cache-aware context layout:
Regression harness coverage under noise:
Possible Harness Scenario
Set up a transcript where the assistant proposes a runner-proof workflow, then the user says:
Expected behavior:
Failure indicators:
Open Questions
<latest_intent_checkpoint>fragment?Non-Goals For Now