Skip to content

feat(ai): bounded agent context-assembly policy (ADR-0028 Option A) - #94

Merged
gustavobertoi merged 1 commit into
feat/llm-usage-metricsfrom
feat/agent-context-policy
Jun 4, 2026
Merged

gustavobertoi merged 1 commit into
feat/llm-usage-metricsfrom
feat/agent-context-policy

Conversation

@gustavobertoi

Copy link
Copy Markdown
Collaborator

Second "finish the agents" leaf PR. ADR-0028 Option A — bounded context-assembly.

Stacked on #93 (shares agent.go). Merge #93 first, then this retargets to main.

What

ai/agent gains two optional inputs (backward compatible — absent = today's behavior):

  • maxContextTokens — approximate token budget for the running transcript.
  • contextStrategy — drop-oldest (default, deterministic) | summarize.

Before each completion, trimContext (internal/packages/functions/ai/context.go) bounds the
transcript: it preserves the leading system turns + the first user task and the most recent turns,
dropping the oldest middle turns. summarize replaces the dropped span with one LLM-generated
summary turn (an opt-in extra model call, recorded as a steps entry; falls back to drop-oldest on
error). Token estimation is a coarse ~4-chars/token heuristic, centralized so a real tokenizer can
replace it later.

Scope

Persisted cross-run session store (Option B) and vector/RAG memory (Option C) remain deferred.
ADR-0028 Proposed → Accepted (Option A).

Verify

make lint 0 · make build ok · make test 706 pass · make swagger ok. trimContext is
pure and table-tested (under/over budget, head preservation, recent-turns-exceed-budget).

Next leaf PR: 0030 structured output.

🤖 Generated with Claude Code

ai/agent gains optional maxContextTokens (approximate budget) + contextStrategy
(drop-oldest default | summarize). Before each completion the running transcript
is bounded by trimContext: it preserves the leading system turns + the first
user task and the most recent turns, dropping the oldest middle turns;
'summarize' replaces them with one LLM-generated summary turn (opt-in extra call,
recorded as a steps entry). Token estimate is a coarse ~4-chars/token heuristic,
centralized for a future tokenizer. Absent budget = unchanged behavior.

Persisted cross-run session store (Option B) and vector/RAG (Option C) deferred.
ADR-0028 Proposed -> Accepted (Option A shipped).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@gustavobertoi
gustavobertoi merged commit 3e29a5e into feat/llm-usage-metrics Jun 4, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant