Skip to content

feat(ai): LLM token-usage metrics for ai/chat & ai/agent (ADR-0029 Phase A) - #93

Merged
gustavobertoi merged 1 commit into
mainfrom
feat/llm-usage-metrics
Jun 4, 2026
Merged

gustavobertoi merged 1 commit into
mainfrom
feat/llm-usage-metrics

Conversation

@gustavobertoi

Copy link
Copy Markdown
Collaborator

First of the "finish the agents" leaf PRs (quick wins). ADR-0029 Phase A — usage visibility.

What

Surface per-call LLM token usage as Prometheus counters, so cost is aggregatable across runs
(previously usage was only embedded in each node's usage output):

  • fuse_llm_tokens_total{function,provider,model,type=prompt|completion}
  • fuse_llm_calls_total{function,provider,model,status}

How

  • New ai.UsageRecorder port (internal/packages/functions/ai/usage.go) keeps the ai package
    free of the prometheus dependency; a metrics-backed adapter
    (internal/packages/usage_recorder.go) is injected through packages.NewInternal (fx provides
    *FuseMetrics).
  • ai/chat records once per completion; ai/agent records per reasoning iteration. Emitted from
    the async goroutine (node span already ended — fine for counters).
  • Backward compatible: usage is still also returned in the node output; no behavior change.

Scope

Budget enforcement (Option B, needs a pricing table) and external metering (Option C) remain
deferred. ADR-0029 Proposed → Accepted (Phase A).

Verify

make lint 0 · make build ok · make test 701 pass. End-to-end:
curl localhost:9090/metrics | grep fuse_llm_ after running an ai node.

Next leaf PRs: 0028 (context-assembly policy) then 0030 (structured output).

🤖 Generated with Claude Code

…ase A)

Surface per-call token usage as Prometheus counters so cost is aggregatable
across runs (previously usage was only in each node's output):
- fuse_llm_tokens_total{function,provider,model,type=prompt|completion}
- fuse_llm_calls_total{function,provider,model,status}

A narrow ai.UsageRecorder port keeps the ai package free of the prometheus
dependency; a metrics-backed adapter is injected via packages.NewInternal.
ai/chat records once per completion; ai/agent records per reasoning iteration.

Budget enforcement (Option B) and external metering (Option C) remain deferred.
ADR-0029 Proposed -> Accepted (Phase A shipped).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

This branch was previously deployed

1 inactive deployment
prod — d0f4ba63 Deployed Jun 4, 2026 by gustavobertoi via pr-image #318
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant