Skip to content

Report Grok usage with recorded and estimated spend - #3135

Open
olddonkey wants to merge 39 commits into
steipete:mainfrom
olddonkey:feat/grok-real-token-usage
Open

olddonkey wants to merge 39 commits into
steipete:mainfrom
olddonkey:feat/grok-real-token-usage

Conversation

@olddonkey

@olddonkey olddonkey commented Aug 22, 2026

Copy link
Copy Markdown
Contributor

Problem and resulting behavior

Grok's local fallback reads ending context occupancy from signals.json, which is not completed-turn usage, and publishes no local cost. This PR reads completed turns from bounded native CLI logs and carries recorded-versus-estimated cost through the menu and Usage & Spend.

  • Read both session/update and _x.ai/session/update records in updates.jsonl, with daily, model, and request breakdowns.
  • Prefer a positive outer usage.costUsdTicks / 1e10 as the authoritative turn total. Count it once. Show nested model dollars only when all nested ticks exist and reconcile to that total; otherwise retain model tokens with unknown model dollars. Missing recorded cost falls back to disclosed public xAI list prices.
  • Keep useful remote data during billing failures while rescanning local sessions. Consumers select the newer publication for the current configuration and preserve account override isolation.
  • Retain cost provenance after menu and dashboard window/day filtering. Populated surfaces disclose Grok CLI-recorded spend, list price where unrecorded · not a bill.
  • Refresh the xAI models.dev catalog for Grok-only installs and keep its pricing fingerprint independent from native Codex pricing.

OpenCodex integration

With the existing Include OpenCodex usage logs switch enabled (off by default), Usage & Spend also includes reported physical Grok OAuth attempts. The producer contract originally proposed in OpenCodex #3642 has now landed in dev through the attributed carry #3762, merged as f00f2bcaea251ebe7ad4de9e38337b4be0ccee47. It writes attempts[].credentialSource from the resolved upstream transport. Compatibility is verified against the pinned producer commit in the evidence below; availability in a released OpenCodex version is not claimed.

The landed carry's immutable source contract has additionally been inspected directly, including resolved-adapter stamping, persistence normalization, and OAuth/API-key regression assertions. The executable capture remains pinned to its documented producer revision; this source audit does not claim a new capture or released-version coverage.

Only provider: "xai", credentialSource: "grok-oauth" attempts with an upstream send and reported token usage qualify. API-key traffic, historic rows, locally answered requests, unknown sources, and estimated/unreported token counts stay excluded. Current configuration and top-level credential metadata cannot retroactively classify usage.

Combo requests contribute each qualifying attempt's own token counts, never the parent aggregate. Duplicate request IDs are resolved before grouping; duplicate attempt ordinals are rejected. SQLite cache schema 3 retains attempt metadata and rebuilds older derived caches from the log. OpenCodex dollars use list prices and remain estimates, including when combined with native CLI-recorded spend. Missing token classes and unknown prices retain tokens without inventing a dollar value. The Grok menu continues to use native CLI logs.

Producer-to-dashboard evidence and review fixes

The committed evidence report and raw ledger fixture contain output from OpenCodex's unmodified production handlers and durable usage writer at 146ed679c9633e5d68726217fcadc8e0b107339b. Two localhost HTTP requests exercised OAuth 401 replay through Responses and native Chat API-key dispatch. Upstream responses and credentials used isolated fixtures; this is production-path capture evidence, not live vendor authentication or billing evidence. The capture helper rejects unexpected external fetches and is reproducible against that pinned checkout.

CodexBar imports the exact captured bytes through the production disk loader, with no injected entries or loader closure, and reopens the persisted cache:

producer_capture_sha256=ef6d8758b40910f6e5993d5b5a105a2ad2834c6c1bd0565ab87b61cf091c4978
producer_log_rows=2 total_reported_tokens=10 grok_oauth_tokens=5
producer_import_dashboard_tokens=5 cache_reopen_bytes=0
producer_api_key_only_subscription_rows=0

Standalone reports retain explicit xAI custom-price estimates from either the caller overlay or the application overlay, including known zero. Raw xAI records without an explicit price remain token-only, and the subscription fan-out still excludes API-key and historic records. opencode and opencode-free retain their existing catalog and custom prices.

Application overlays first match the original model name before the provider-qualified catalog name. Bare keys retain precedence when both keys exist; incomplete rates stay unknown, cached-input accounting is preserved, and caller-supplied custom pricing still takes precedence over the application overlay. New regressions cover these cases through the application-overlay parameter and the standalone disk/cache loader. Coverage includes known-zero overrides, incomplete rates, cache accounting, and both snapshot and application overlays.

Both async native Grok scan entry points propagate the executor's cancellation callback through discovery, JSONL reads, and aggregation. An in-flight regression cancels after parsing begins, observes CancellationError, and confirms the next queued scan runs within one second. Cancelled partial parses are uncacheable and cannot establish complete history. The earlier c3919a224 serial proof stopped at 4,874/40,000 decoded records and released the queue after 0.001864583 seconds. The evidence report also retains the initial measurement.

Native numeric safety

Native token conversion rejects booleans, fractions, negative values, and out-of-range numbers without trusting NSNumber's clamping integer bridge. An absent token class retains the established zero default; a malformed count remains unknown. Explicit valid totals retain precedence, so decoding never adds input and output unnecessarily.

Checked addition carries unknown or overflowed values through model, day, and window aggregation, including after later valid records and cached-log rereads. Valid neighboring token classes and CLI-recorded spend remain available. Incomplete token accounting does not establish full history coverage; estimated pricing requires representable inputs. Production-scanner regressions cover Int.max plus one within a turn, sums across turns and days, nested models, invalid JSON number types, explicit totals, snapshots, and cache reuse.

The same checked, complete-count aggregation now covers remote-backed Grok menu projection, rolling-window requests, comparison summaries, and menu/Widget fallback totals. Unknown daily counts cannot be dropped into a plausible partial sum. An end-to-end scanner regression covers remote-backed and fallback projections for cross-day overflow and an unknown day followed by a valid day; menu, Widget, and dashboard totals stay unavailable, while narrowing to a valid one-day window restores its count.

Bounds and compatibility

Native scans run on the dedicated executor with limits of 64 MiB / 20,000 turns per file, 1 MiB per record, 256 sessions / 256 MiB / 100,000 turns per scan, and 4,096 discovery entries. The process cache retains at most 64 files or 50,000 turns. Cancellation and truncated history cannot publish complete coverage.

Merged main 8b254dbec11ddd5c5547878d9640e4e965306c71, retaining its checked token aggregation, daily spend ledger, native parser-revision migration, and Grok terminal-billing-failure work avoidance. The local-summary injection seam now invokes the PR's pricing-aware 365-day scanner only after the upstream billing/identity gate permits a snapshot. Failed billing still refreshes local spend through the app's existing independent fallback path.

OpenCodex token counts retain upstream's safe numeric conversion and truncation policy. Attempt ordinals and send counts remain exact integers: booleans, fractional counts, and out-of-range values cannot qualify an OAuth attempt. Aggregation retains overflow as unavailable while keeping valid neighboring token classes and Grok estimate coverage. The upstream cursor parser version invalidates legacy numeric caches; schema 3 still preserves request-time attempt provenance.

Regenerated the native parser hash from the merged source: a8559238a5fc0480. Current-main hash f5fdba377006d7be and prior-PR hash c3a879df4eff7187 join the compatible predecessors. Stored rows and checkpoints are retained, while upstream's per-file parser revisions schedule bounded reparsing where needed. SQLite adoption tests cover the current-main and prior-PR hashes, and upstream whitespace/subagent migration tests remain in the validation set. The architecture gate keeps its exact provider-reference checks at their updated source locations. Grok release notes are under 0.59.1 — Unreleased; all published sections match main.

Validation

Head: a41736133e7a2d922db7bbc5b0e18e7e729e855d.

  • make check passed: zero violations in 2,199 Swift files.
  • Focused Grok, menu, Widget, dashboard, window, and architecture regressions passed: 888 tests in 110 suites, including all four native-log-to-menu overflow scenarios.
  • Fresh serial native-window, numeric safety, menu projection, producer-import, provenance, and cancellation proof passed: 40 tests in 7 suites. The native 30-day window retains 98,631,812 tokens / $12.94957366, with vendorMetered provenance and populated menu/dashboard disclosures. Empty 1-day/7-day windows retain unknown. The pinned producer ledger imports five OAuth tokens, excludes API-key traffic, and reopens its cache without log reads.
  • Current-head make test passed: 1,081 selections, 91/91 groups successful on the first attempt, zero failures/retries/timeouts (632.5 seconds execution; 634.5 seconds total). No code changed after validation.
  • Current-head GitHub CI: lint, changes, GitGuardian, and all three Linux builds/tests passed. Both macOS CI shards are queued waiting for runners; the full CI gate is not yet complete. The local macOS full-suite pass is recorded separately above.
  • git diff --check passed.

The supplemental proof and local-corpus measurements below were captured at c3919a224. They remain historical, source-linked evidence; current-head validation is recorded separately above. The supplemental proof uses the repository’s serial execution mode because those suites share scanner-cache counters.

The September 6 local-corpus proof retained vendorMetered for populated windows: 7 days = 95,891,889 tokens / $12.57875572; 30 days = 98,631,812 tokens / $12.94957366. The empty 1-day window retained unknown. Controlled JSONL fixtures also proved recorded-only and estimate-only windows after narrowing mixed source history.

CODEXBAR_LIVE_GROK_CATALOG_PROOF=1 CODEXBAR_SUPPRESS_TEST_KEYCHAIN_ACCESS=1 CODEXBAR_ALLOW_TEST_KEYCHAIN_ACCESS=0 CODEXBAR_TEST_CODEX_FILE_ISOLATION=1 swift test --skip-build --no-parallel --filter 'GrokWindowProvenanceProofTests|GrokXAISpendCatalogTests|GrokTokenSnapshotProjectionTests|GrokCostUsagePricingTests|GrokLocalSessionScannerTests|ProviderArchitectureGatekeeperTests|CostProvenanceTests|SpendDashboardGrokFreshnessTests|GrokOpenCodexUsageTests|CostUsageScanExecutorTests|OpenCodexUsagePricingTests|CostUsageStoreTests|SpendStackedBarChartTests'

The OpenCodex regressions cover mixed OAuth/API-key/provider attempts, historical and malformed records, duplicate suppression, cache reopen/incremental append/schema upgrade, missing-price behavior, recorded-plus-estimated date filtering, and the dashboard's opt-in switch. Existing native regressions cover bounded parsing, outer/nested cost reconciliation, local fallback freshness, account isolation, and menu/dashboard provenance.

All ordinary tests suppress Keychain access and isolate provider files. The optional native proof reads local Grok logs and the cached pricing catalog without authentication, browser-cookie import, or live provider requests. Presentation evidence uses production menu/dashboard models; no app-bundle screenshot is claimed.

Maintainer decision

Please revisit the 2026-08-21 Grok cost ruling before merge. It preferred public-card pricing based partly on my incorrect 1e9 divisor; #3345 established 1e10, and the corrected measurement explains the difference. This branch uses recorded spend with public-card fallback.

Whether existing Grok users should receive this disclosed dollar surface by default remains an owner decision. The accounting corrections and source labeling do not override that decision. Maintainer approval is still required.

@clawsweeper

clawsweeper Bot commented Aug 22, 2026

Copy link
Copy Markdown

🦞👀
ClawSweeper picked this up.

Pull request received. I will update this pull request when review starts.

ClawSweeper review complete

ClawSweeper finished reviewing this revision. The review result is being finalized.

View the workflow run.

olddonkey added a commit to olddonkey/CodexBar that referenced this pull request Aug 22, 2026

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 09cf7edb0f

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread Sources/CodexBar/UsageStore+Refresh.swift Outdated
Comment thread Sources/CodexBarCore/Providers/Grok/GrokLocalSessionScanner.swift
@clawsweeper clawsweeper Bot added merge-risk: 🚨 compatibility 🚨 Merging this PR could break existing users, config, migrations, defaults, or upgrades. P2 Normal priority bug or improvement with limited blast radius. rating: 🧂 unranked krab Not merge-ready due to missing proof or serious correctness/safety concerns. status: 📣 needs proof The PR needs real behavior proof before ClawSweeper can clear the contributor ask. labels Aug 22, 2026
@clawsweeper

clawsweeper Bot commented Aug 22, 2026

Copy link
Copy Markdown

Codex review: blocked before merge. Reviewed September 12, 2026, 1:51 AM ET / 05:51 UTC (Revision 60).

ClawSweeper review

What this changes

Read completed Grok CLI turns, distinguish recorded spend from estimates, and include eligible OpenCodex OAuth usage in the spend dashboard.

Merge readiness

Blocked before merge - 3 items remain

This remains useful work absent from main and v0.59.0. The prior overflow findings are resolved, and the supplied real behavior proof is sufficient; product acceptance remains outstanding.

Priority: P2
Reviewed head: a41736133e7a2d922db7bbc5b0e18e7e729e855d
Owner decision: Required. See Decision needed.

Review scores

Measure Result What it means
Overall readiness 🐚 platinum hermit (4/6) Useful implementation with strong source-linked runtime evidence and no remaining actionable code finding; owner acceptance is a separate merge gate.
Proof confidence 🦞 diamond lobster (5/6) Sufficient (terminal): Current-head local Grok logs exercise the production scanner and menu/dashboard projections with recorded totals and disclosures. The pinned OpenCodex production-handler capture exercises disk import, API-key exclusion, and cache reopen; inspected screenshots are historical supplements, not proof of the final attribution logic.
Patch quality 🐚 platinum hermit (4/6) No actionable review findings were identified.

Verification

Check Result Evidence
Real behavior Verified Sufficient (terminal): Current-head local Grok logs exercise the production scanner and menu/dashboard projections with recorded totals and disclosures. The pinned OpenCodex production-handler capture exercises disk import, API-key exclusion, and cache reopen; inspected screenshots are historical supplements, not proof of the final attribution logic.
Evidence reviewed 11 items Repository policy: Read the complete root AGENTS.md; no nested AGENTS.md or maintainer-note files were found. Applied provider isolation, bounded background work, and safe validation guidance without executing tests or provider probes.
Still needed on main: The fetched main scanner still reads context occupancy from signals.json and supplies nil local costs. It does not implement this branch's completed-turn accounting.
Release comparison: The v0.59.0 scanner likewise uses signals.json and publishes no local dollars; the requested behavior is not present in the supplied latest release.
Findings None None.
Security None None.

How this fits together

CodexBar combines local usage logs and remote provider snapshots into menu, widget, and spend-dashboard summaries. This change updates Grok accounting and the optional OpenCodex log import feeding those summaries.

flowchart TD
  A[Grok CLI completed-turn logs] --> B[Bounded local scanner]
  C[Recorded spend and price catalog] --> B
  D[Optional OpenCodex logs] --> E[Reported OAuth attempt filter]
  B --> F[Windowed usage and cost provenance]
  E --> F
  F --> G[Menu and spend dashboard]
Loading

Decision needed

Question Recommendation
Should existing Grok users receive CLI-recorded dollars with disclosed public-price fallback by default? Approve the disclosed dollar display: Accept recorded spend with clearly labeled fallback estimates and resolve the outstanding review after confirming its repairs.

Why: The owner supported completed-turn accounting, but the current proposal explicitly leaves acceptance of the revised dollar-display policy unresolved.

Before merge

  • Resolve merge risk (P2) - Existing Grok users with cost tracking enabled would receive dollar figures automatically; CLI-recorded spend and public-price fallback can differ materially from subscription billing, so the default display still needs owner acceptance.
  • Complete next step (P2) - Have steipete approve or narrow the default Grok dollar-display policy and resolve the outstanding changes-requested review.
  • Resolve maintainer decision - Resolve the maintainer decision shown above before merge.
Agent review details

Security

None.

Review metrics

Metric Value Why it matters
Production and test delta Production/scripts +1,880 −261; tests/fixtures +3,307 −158 The production growth supports bounded parsing, pricing provenance, and optional log attribution, with substantial regression coverage.

Merge-risk options

Maintainer options:

  1. Accept the new default explicitly (recommended)
    Approve the revised recorded-versus-estimated display contract for existing cost-tracking users.
  2. Preserve token-only defaults
    If automatic dollars are unwanted, retain token-only presentation until the user opts in.

Technical review

Best possible solution:

Adopt accurate completed-turn accounting with source-specific, non-bill disclosures under an explicitly approved default-display policy.

Do we have a high-confidence way to reproduce the issue?

Yes, source establishes that main sums context occupancy rather than completed-turn consumption; independent corpus reports corroborate the mismatch. This read-only review did not execute a reproduction.

Is this the best way to solve the issue?

Yes for the accounting implementation: bounded completed-turn parsing and recorded-cost precedence fit the evidence. Whether dollars should appear automatically remains a product choice.

AGENTS.md: found and applied where relevant.

Codex review notes: model internal, reasoning medium; reviewed against a998f677dd1f.

Labels

Label justifications:

  • P2: This corrects Grok usage accounting and adds spend reporting with a limited provider-specific blast radius.
  • merge-risk: 🚨 compatibility: Existing users would automatically gain a dollar display whose subscription-billing meaning requires explicit product acceptance.
  • rating: 🐚 platinum hermit: Overall readiness is 🐚 platinum hermit; proof is 🦞 diamond lobster and patch quality is 🐚 platinum hermit.
  • status: 👀 ready for maintainer look: ClawSweeper has no concrete contributor-facing blocker left for this PR. Sufficient (terminal): Current-head local Grok logs exercise the production scanner and menu/dashboard projections with recorded totals and disclosures. The pinned OpenCodex production-handler capture exercises disk import, API-key exclusion, and cache reopen; inspected screenshots are historical supplements, not proof of the final attribution logic.
  • proof: sufficient: Contributor real behavior proof is sufficient. Current-head local Grok logs exercise the production scanner and menu/dashboard projections with recorded totals and disclosures. The pinned OpenCodex production-handler capture exercises disk import, API-key exclusion, and cache reopen; inspected screenshots are historical supplements, not proof of the final attribution logic.

Evidence

What I checked:

  • Repository policy: Read the complete root AGENTS.md; no nested AGENTS.md or maintainer-note files were found. Applied provider isolation, bounded background work, and safe validation guidance without executing tests or provider probes. (AGENTS.md, a41736133e7a)
  • Still needed on main: The fetched main scanner still reads context occupancy from signals.json and supplies nil local costs. It does not implement this branch's completed-turn accounting. (Sources/CodexBarCore/Providers/Grok/GrokLocalSessionScanner.swift:48, a998f677dd1f)
  • Release comparison: The v0.59.0 scanner likewise uses signals.json and publishes no local dollars; the requested behavior is not present in the supplied latest release. (Sources/CodexBarCore/Providers/Grok/GrokLocalSessionScanner.swift:48, dca9c5f4b7a2)
  • Review continuity and numeric repairs: The checkout matches the previous completed review's SHA. Native conversion rejects invalid numbers, checked aggregation preserves unknown totals, and four projection scenarios cover both remote-backed and fallback consumers through menu, widget, and dashboard models. (Tests/CodexBarTests/GrokTokenSnapshotProjectionTests.swift:107, a41736133e7a)
  • Complete captured body and current proof: Retrieved the complete 13,125-unit body and verified SHA-256 df3cfe3efbe0c58679da85cb80c2b4c5d73f913f06920d67cbe86a17f22414f5 against the supplied snapshot. It records current-head real local-log results of 98,631,812 tokens and $12.94957366, preserved recorded provenance and disclosures, plus 91 successful full-suite groups. These are contributor-reported runs, not executions by this reviewer. (a41736133e7a)
  • Production-path import and persistence proof: The committed report records production OpenCodex handler output imported through CodexBar's disk loader: five OAuth tokens, zero API-key subscription rows, and cache reopen without log reads. Source inspection confirms the test imports the captured bytes without an injected loader; separate coverage exercises schema-version rebuild and incremental append. (docs/evidence/grok-opencodex-producer-2026-09-05.md:29, a41736133e7a)

Likely related people:

  • steipete: Suggested for follow-up; no historical authorship or introduction is verified. (role: unverified routing candidate; confidence: low)
  • olddonkey: Suggested for follow-up; no historical authorship or introduction is verified. (role: unverified routing candidate; confidence: low)

Rating scale

Score Internal tier Crab rank Meaning
6/6 S 🦀 challenger crab Exceptional readiness
5/6 A 🦞 diamond lobster Very strong readiness
4/6 B 🐚 platinum hermit Good normal PR; ordinary maintainer review
3/6 C 🦐 gold shrimp Useful, but confidence is limited
2/6 D 🦪 silver shellfish Proof or implementation needs work
1/6 F 🧂 unranked krab Not merge-ready
N/A NA 🌊 off-meta tidepool Rating does not apply

Overall follows the weaker of proof and patch quality.
Shiny media proof means a screenshot, video, or linked artifact directly shows the changed behavior. Runtime, network, CSP, and security claims still need visible diagnostics.

Workflow

  • ClawSweeper keeps one durable marker-backed review comment per issue or PR.
  • Re-runs edit this comment so the latest verdict, findings, and automation markers stay together instead of adding duplicate bot comments.
  • A fresh review can be triggered by eligible @clawsweeper re-review comments, exact-item GitHub events, scheduled/background review runs, or manual workflow dispatch.
  • PR/issue authors and users with repository write access can comment @clawsweeper re-review or @clawsweeper re-run on an open PR or issue to request a fresh review only.
  • Maintainers can also comment @clawsweeper review to request a fresh review only.
  • Fresh-review commands do not start repair, autofix, rebase, CI repair, or automerge.
  • Maintainer-only repair and merge flows require explicit commands such as @clawsweeper autofix, @clawsweeper automerge, @clawsweeper fix ci, or @clawsweeper address review.
  • Maintainers can comment @clawsweeper explain to ask for more context, or @clawsweeper stop to stop active automation.

History

Review history (59 earlier review cycles; latest 8 shown)
  • reviewed 2026-09-08T00:51:48.532Z sha c35734a :: blocked before merge. :: none
  • reviewed 2026-09-08T01:02:38.334Z sha c35734a :: blocked before merge. :: none
  • reviewed 2026-09-08T01:07:50.811Z sha c35734a :: blocked before merge. :: none
  • reviewed 2026-09-08T01:32:25.385Z sha c35734a :: blocked before merge. :: none
  • reviewed 2026-09-12T04:57:16.810Z sha 40f9e47 :: blocked before merge. :: [P2] Guard native Grok token decoding and accumulation against overflow
  • reviewed 2026-09-12T05:06:33.398Z sha 40f9e47 :: blocked before merge. :: [P2] Guard native Grok token decoding and accumulation against overflow
  • reviewed 2026-09-12T05:29:15.595Z sha 6d3d5af :: blocked before merge. :: [P2] Preserve overflow safety through the Grok menu projection
  • reviewed 2026-09-12T05:39:46.066Z sha a417361 :: blocked before merge. :: none

@olddonkey
olddonkey force-pushed the feat/grok-real-token-usage branch from 08360b5 to e3cd3b9 Compare August 22, 2026 06:21
@olddonkey olddonkey changed the title Report real Grok token usage and list-price cost from CLI session logs Report real Grok token usage and list-price cost, from the CLI logs and OpenCodex alike Aug 22, 2026
@clawsweeper clawsweeper Bot added merge-risk: 🚨 auth-provider 🚨 Merging this PR could break OAuth, tokens, provider routing, model choice, or credentials. rating: 🦐 gold shrimp Decent PR readiness signal, but merge confidence is limited. and removed rating: 🧂 unranked krab Not merge-ready due to missing proof or serious correctness/safety concerns. labels Aug 22, 2026
@olddonkey

Copy link
Copy Markdown
Contributor Author

Both automated findings are addressed, plus the review's other checklist items. The inline comments were left against 09cf7edb0, which no longer exists — the branch has since been rebased onto 27c7f334e and the head is now 03e5a25dc, so I'm summarising here rather than replying in a stale diff.

P1 — Preserve the Grok fallback on repeated probe failures

Fixed in 923193ec0, Sources/CodexBar/UsageStore+Refresh.swift. The guard had been hoisted into the if provider == .grok, publication == nil condition, so a Grok failure with a publication fell through to the generic else if tokenCostRequiresProviderSnapshot { clearTokenSnapshot } branch. Grok now owns its branch outright and can never reach the clear:

if provider == .grok {
    if self.tokenSnapshotPublicationForCurrentProviderConfig(for: provider) == nil {
        Task { @MainActor [weak self] in
            await self?.scanAndPublishGrokLocalTokenSnapshot(...)
        }
    }
} else if Self.tokenCostRequiresProviderSnapshot(provider) {
    self.clearTokenSnapshot(for: provider)
}

Regression coverage is in missing remote snapshot scans and publishes local tokens then clears empty data. Per the review's request it now drives two consecutive failing refreshes (03e5a25dc) rather than one — which matters here, because the first failure is what publishes through the fallback scan and only the second arrives with a publication in place, i.e. the failure that used to wipe the row. Both iterations assert the row still reads 77 tokens and that no redundant rescan ran.

P2 — Refresh pricing before scanning Grok sessions

Correct, and thank you — this was a genuine gap and not one the local tests would have surfaced. refreshPricingIfAllowed is gated to Codex and Claude, and Grok never reaches it at all because its snapshot comes from the provider probe rather than CostUsageFetcher.loadTokenSnapshot. On a machine with Codex or Claude also enabled the shared cache is already populated, so the failure is invisible there; enable only Grok and the catalog never appears and the Cost row shows tokens with no money, permanently.

Fixed in 744677e68. The Grok scan paths now request ModelsDevPricingPipeline.refreshIfNeeded through a summarizeRequestingPricingRefresh wrapper, called from all four scan sites (GrokStatusProbe, both branches in GrokProviderDescriptor, and UsageStore.scanAndPublishGrokLocalTokenSnapshot). It is detached rather than awaited, matching how the Codex and Claude paths already treat it — pricing availability must not delay or fail a local scan — and it is safe to call repeatedly, since it returns immediately unless the cache is stale and serialises through its own coordinator. summarize itself stays synchronous and side-effect free.

Note the inline comment still points at GrokLocalSessionScanner.swift:662; that line is the unchanged pricing lookup, and the fix is upstream of it in the new wrapper, so the anchor looks live even though it is addressed.

Coverage: absent models dev cache requests a background refresh, stale models dev cache requests a background refresh, and fresh models dev cache skips the background refresh. All three assert whether a refresh was requested through an injected transport — no test touches the network.

Real-session evidence

CODEXBAR_LIVE_GROK_CATALOG_PROOF=1 swift test --filter GrokXAISpendCatalogTests, against real local Grok CLI sessions, through the shipped code path:

catalog_source=grok
today_tokens=5043749
last_30_days_tokens=52696354
today_cost_usd=3.3471699999999993
window_cost_usd=49.353424
cost_provenance=listPriceEstimate
history_days=365
priced_days=4
token_days=4
daily_buckets=4
available_sources=grok

The same corpus on main reports 653K tokens and no cost. history_days=365 shows the requested window is honoured (it was pinned to 30). priced_days == token_days shows no day was silently left unpriced. The gated proof was extended in e3cd3b9ce to print cost, provenance and priced-day coverage, since tokens alone cannot evidence the half of this change that is about money.

Those figures were cross-checked against an independent reimplementation of the pricing formula over the same logs; the two agree to the cent.

Merge risk / branch state

Rebased onto current main (27c7f334e); the branch reports clean. Full suite on the head: 77/77 groups, 922 selections, 0 failures. swiftformat --lint and swiftlint --strict clean. Upstream CI green on the previous head including all three Linux builds.

One thing deliberately left undone: no CHANGELOG.md entry. 0.54.1 was finalized and there is no open Unreleased section, so I did not invent a version heading — happy to add one wherever you prefer.

@clawsweeper clawsweeper Bot added rating: 🧂 unranked krab Not merge-ready due to missing proof or serious correctness/safety concerns. rating: 🦐 gold shrimp Decent PR readiness signal, but merge confidence is limited. and removed rating: 🦐 gold shrimp Decent PR readiness signal, but merge confidence is limited. rating: 🧂 unranked krab Not merge-ready due to missing proof or serious correctness/safety concerns. labels Aug 22, 2026
@olddonkey

olddonkey commented Aug 22, 2026

Copy link
Copy Markdown
Contributor Author

Both new findings addressed at 44d79a95a.

P1 — Do not map every xAI log record to the Grok subscription

Agreed, and taken as specified rather than argued down. Routing on the prefix alone is right for the case that motivated this — traffic authenticated with the user's Grok account, which is what makes it consume SuperGrok quota — but it silently folds an API-key user's pay-as-you-go xAI spend into the subscription row. CodexBar already models the developer platform as its own xai provider precisely to keep those apart, so the old behaviour crossed a boundary the app deliberately maintains.

The usage log carries no per-record credential evidence; I checked every field emitted for xai rows (requestId, timestamp, provider, model, requestedModel, resolvedModel, usage, usageStatus, status, routeDecision, …) and there is nothing about auth, account or key. The signal that does exist is ~/.opencodex/config.json, which records authMode per provider.

So attribution now requires positive OAuth evidence:

  • xai routes to .subscription(.grok) only when its configured authMode is OAuth. Anything else returns .tokenOnly — the spend is real, it just belongs to no tracked subscription — rather than .unknown, which would read as "unrecognised provider".
  • Fail closed. A missing or malformed config, a providers block without xai, or an entry without authMode all count as no evidence and keep the records off the Grok row.
  • OpenCodexRouteDispatcher stays a pure function. The set of OAuth-backed provider ids is threaded in from the caller (OpenCodexUsageFanOutSpendDashboardSource), so the routing site never touches the filesystem and every existing caller and test that does not care about auth keeps working.
  • The gate applies only to xai. openai, kimi-coding, deepseek and opencode-go are untouched — changing them would be an unreviewed behaviour change for other providers — and a test pins that they ignore xAI auth state entirely.

Coverage: xai OAuth config routes to Grok, xai non OAuth config stays token only (parameterised over several non-OAuth values), xai routing fails closed without readable complete OAuth config, non xai subscription routes ignore xai auth state, plus fan-out cases proving the same entries land on the Grok row under an OAuth config and are absent under an API-key one. No test reads the developer's real ~/.opencodex; the home directory is injected.

docs/grok.md no longer claims this path cannot distinguish OAuth from API-key traffic, because it now can.

P2 — Republish the Grok snapshot after a missing catalog refreshes

I looked at this closely and am deliberately not adding a republish path. Reasoning, so you can overrule it if you disagree:

The refresh is fire-and-forget, so the scan that requests it returns whatever the cache currently holds — that part is accurate. But the parse cache stores parsed turns, not prices, so aggregation and pricing re-run on every summarize. The next Grok scan therefore prices against the refreshed catalog with no extra machinery, bounding the unpriced window to a single refresh cycle. That is the same behaviour Codex and Claude already have: refreshPricingIfAllowed dispatches into Task.detached and their current scan does not wait for it either.

The alternative — plumbing a completion signal back across the actor boundary into the @MainActor publication path — buys one refresh cycle of latency on first run, at the cost of a new cross-actor completion path in code that publishes user-visible spend. That trade looked disproportionate, and inconsistent with how the two established providers behave. I have recorded the reasoning as a comment at the call site rather than leaving it implicit, so the next reader does not have to re-derive it.

Happy to build it if you would rather have it.

Evidence

The attribution itself only becomes visible in the app: SpendDashboardSource.mergingOpenCodexInputs is what merges the fan-out into provider rows, and the CLI's cost command reports OpenCodex as its own source rather than routing it, so terminal output cannot show this path. The figures below are read off the freshly packaged build running against real local data, on a machine whose ~/.opencodex/config.json has "xai": { "authMode": "oauth" }; screenshots of both panes follow.

The two halves stay distinguishable in the UI, which makes the attribution legible rather than something you have to take on trust: the CLI goes through the responses API so its SKU is grok-4.6-build, while OpenCodex's records resolve to the bare grok-4.6 / grok-4.5 / grok-4.3. Both sit under the Grok provider.

model row source shown
grok-4.6-build Grok CLI session logs $50.52 · 54M
grok-4.6 OpenCodex $161.21 · 182M
grok-4.5 OpenCodex $4.54 · 5.5M
grok-4.3 OpenCodex $0.50 · 201K

Independently recomputing the same corpus agrees to the cent on both halves: 54,121,501 tokens / $50.52 for the CLI logs, and $166.25 across 1,520 OpenCodex xai records. The CLI half is reproducible by anyone on their own machine through the gated proof test (CODEXBAR_LIVE_GROK_CATALOG_PROOF=1), whose output is in the PR body.

The negative direction — API-key traffic staying off the Grok row — is covered by tests rather than a screenshot, since demonstrating it live would mean rewriting the machine's OpenCodex config.

State

Full suite on 44d79a95a: 77/77 groups, 922 selections, 0 failures. swiftformat --lint and swiftlint --strict clean. Rebased on 27c7f334e.

image image

@clawsweeper clawsweeper Bot added proof: 📸 screenshot Contributor real behavior proof includes screenshot evidence. proof: sufficient Contributor real behavior proof is sufficient. status: ⏳ waiting on author ClawSweeper has contributor-facing work open and is waiting for author action. and removed status: 📣 needs proof The PR needs real behavior proof before ClawSweeper can clear the contributor ask. labels Aug 22, 2026
@olddonkey olddonkey changed the title Report real Grok token usage and list-price cost, from the CLI logs and OpenCodex alike Report real Grok token usage and list-price cost from CLI logs Aug 23, 2026
@olddonkey

Copy link
Copy Markdown
Contributor Author

Addressed both current findings in 211e1977d and resolved the two review threads.

  • P1 / historical xAI attribution: removed current-config-based xAI → Grok routing. usage.jsonl has no request-time credential provenance, so xAI records now remain token-only until the producer can persist that evidence. Removed the config reader/plumbing and added dispatcher/fan-out regressions.
  • P2 / first pricing publication: when no models.dev artifact exists, the first Grok scan now awaits the initial best-effort refresh attempt before summarizing. A successful refresh prices the first returned snapshot; stale catalogs still price immediately and refresh in the background. Added a regression that writes the catalog during refresh and asserts the first summary is priced.
  • Updated the PR title/body and docs/grok.md so they no longer claim OpenCodex xAI traffic is merged into the Grok subscription row.

Validation on the exact pushed head:

  • focused Grok/OpenCodex suites: 32 tests passed
  • make check: passed
  • make test: 922 selections, 77/77 groups, 0 failed groups, 0 retries
  • branch is based on current main (27c7f334e) and the merge-tree is clean

@clawsweeper re-review

@clawsweeper

clawsweeper Bot commented Aug 23, 2026

Copy link
Copy Markdown

🦞🧹
ClawSweeper re-review requested.

I asked ClawSweeper to review this item again.
Action: item re-review queued (workflow sweep.yml, event exact_review_queue).
Result: when the review finishes, ClawSweeper will create the durable review comment if needed or update the existing comment in place.

@clawsweeper clawsweeper Bot added merge-risk: 🚨 availability 🚨 Merging this PR could cause crashes, hangs, restart loops, stalls, or process outages. and removed merge-risk: 🚨 compatibility 🚨 Merging this PR could break existing users, config, migrations, defaults, or upgrades. proof: sufficient Contributor real behavior proof is sufficient. proof: 📸 screenshot Contributor real behavior proof includes screenshot evidence. labels Aug 23, 2026
@clawsweeper

clawsweeper Bot commented Sep 5, 2026

Copy link
Copy Markdown

🦞🧹
ClawSweeper re-review requested.

I asked ClawSweeper to review this item again.
Action: item re-review queued (workflow sweep.yml, event exact_review_queue).
Result: when the review finishes, ClawSweeper will create the durable review comment if needed or update the existing comment in place.

@olddonkey

Copy link
Copy Markdown
Contributor Author

Synced #3135 with main 0be7714904c311b6349a250407ad88f0ba73c524 in final head c3919a224eeb50298728921290efa784ba987b21.

The generated parser-hash conflict is resolved by regeneration from the merged source (1a4afd74939160fd). Main's atomic catalog replacement does not change native parsed rows or checkpoints, so the released-main hash 9547dc9d7b7675f6 and prior-PR hash 0bd6588c70196700 are retained as compatible predecessors. The actual SQLite adoption tests verify that both preserve stored state without rebuilding. Existing predecessors remain intact.

Main's Linux cache fixes, stacked-chart rendering, provider/widget changes, and release metadata are preserved. The two Grok changelog entries now live under 0.56.8 Unreleased; published release sections match main. The evidence report also records that the producer contract landed through the attributed OpenCodex #3762 carry; its original captured bytes remain pinned and are not represented as a new capture of that carry.

Validation on this final head:

  • make check: zero violations in 2,137 files.
  • make test: 1,031 selections, 86/86 groups passed on the first attempt; zero failures, retries, or timeouts.
  • Focused cache adoption, OpenCodex, architecture, and chart tests: 136 passed.
  • Serial native Grok/producer import/pricing/cache adoption proof: 206 tests in 14 suites passed. Native recorded-cost totals remain unchanged; the producer fixture still contributes five OAuth tokens, excludes API-key subscription rows, and reopens its cache without reading log bytes.

The PR description now contains the final source, commands, and evidence. The default Grok dollar display remains an owner decision. Please review the final head and description.

@clawsweeper re-review

@clawsweeper

clawsweeper Bot commented Sep 6, 2026

Copy link
Copy Markdown

🦞🧹
ClawSweeper re-review requested.

I asked ClawSweeper to review this item again.
Action: item re-review queued (workflow sweep.yml, event exact_review_queue).
Result: when the review finishes, ClawSweeper will create the durable review comment if needed or update the existing comment in place.

@olddonkey

Copy link
Copy Markdown
Contributor Author

Updated to 08744509f9691781b67ee47b242cb57df065012f on current main 7353c717502de72ed10dd1507a3f91d5e9ccb716. The only merge conflict was in CHANGELOG.md; both Grok entries and the new upstream Poe fix are retained.

Current-head validation: make check passed with zero SwiftLint violations; make test passed all 1,031 selections in 86 groups on the first attempt, with zero failures, retries, or timeouts. Grok accounting and OpenCodex import source is unchanged. The PR body distinguishes the current full-suite validation from the earlier supplemental proof.

The default Grok dollar display remains an owner decision, and the prior maintainer change-request review still needs reconsideration.

@clawsweeper re-review

@clawsweeper

clawsweeper Bot commented Sep 6, 2026

Copy link
Copy Markdown

🦞🧹
ClawSweeper re-review requested.

I asked ClawSweeper to review this item again.
Action: item re-review queued (workflow sweep.yml, event exact_review_queue).
Result: when the review finishes, ClawSweeper will create the durable review comment if needed or update the existing comment in place.

@olddonkey

Copy link
Copy Markdown
Contributor Author

@steipete The latest ClawSweeper review of 08744509f reports no actionable findings, and local make check plus all 1,031 test selections passed. Could you confirm whether recorded-first Grok costs with disclosed list-price fallback should be shown to existing users who already enabled cost tracking, and revisit the earlier changes-requested review? The current PR body records the accounting evidence and leaves this display policy decision with you. Cross-platform CI is still running.

@olddonkey

Copy link
Copy Markdown
Contributor Author

Updated to 6f4a7d1058de2180f1f293824074cbbcd15ec5cc after main advanced to 31131374b. The latest integration preserves the upstream Claude warning-state and MiMo fixes. The sole conflict was the architecture gate's source locations and Grok bridge record; the merged gate retains the existing Grok reference fingerprints and correct source lines.

make check passed with zero violations in 2,138 files; 46 focused architecture/Claude/Grok/dashboard tests passed. The required full suite and new CI run are in progress. The previous head passed the full local suite, all nine CI checks, and an exact-head review with no actionable findings. The Grok display-policy decision still belongs to the owner.

@clawsweeper re-review

@clawsweeper

clawsweeper Bot commented Sep 6, 2026

Copy link
Copy Markdown

🦞🧹
ClawSweeper re-review requested.

I asked ClawSweeper to review this item again.
Action: item re-review queued (workflow sweep.yml, event exact_review_queue).
Result: when the review finishes, ClawSweeper will create the durable review comment if needed or update the existing comment in place.

@olddonkey

Copy link
Copy Markdown
Contributor Author

Final validation for 6f4a7d1058de2180f1f293824074cbbcd15ec5cc:

  • All nine GitHub CI checks passed: https://github.com/steipete/CodexBar/actions/runs/34060587552
  • Local make check passed with zero violations in 2,138 files.
  • Local make test passed all 1,032 selections in 86 groups on the first attempt, with no failures, retries, or timeouts.
  • The 46 focused architecture, Claude credential-warning, Grok fallback, and dashboard checks passed.
  • ClawSweeper reviewed this exact head and reports no actionable findings; there are no unresolved inline review threads.

The PR body is updated with the current source and validation. The default Grok dollar-display policy and the earlier maintainer changes-requested review still require the owner's decision.

@olddonkey

Copy link
Copy Markdown
Contributor Author

Synced with main 912eac223 and pushed c35734ad5a747a83c46b3844c8ce33fd46dbefe1. The merge preserves the upstream shared report/pricing optimizations and the Grok/OpenCodex accounting behavior.

The parser hash is regenerated as c3a879df4eff7187; SQLite adoption coverage includes the current-main, prior-PR, and intermediate upstream hashes without rebuilding. Existing compatible predecessors remain. The model-target mapping moved unchanged into the existing resolver extension to keep the pricing type within the repository limit; the architecture gate retains its exact reference fingerprints at the new source locations. Grok notes now live under 0.56.9 Unreleased, with published sections matching main.

make check passed with zero violations in 2,143 files. Cache/pricing/provenance regressions and all 41 architecture tests passed. The required full suite and current-head CI are running. The PR body distinguishes current validation from the earlier pinned supplemental proof. The default dollar-display policy remains an owner decision.

@clawsweeper re-review

@clawsweeper

clawsweeper Bot commented Sep 8, 2026

Copy link
Copy Markdown

🦞🧹
ClawSweeper re-review requested.

I asked ClawSweeper to review this item again.
Action: item re-review queued (workflow sweep.yml, event exact_review_queue).
Result: when the review finishes, ClawSweeper will create the durable review comment if needed or update the existing comment in place.

Re-review progress:

@olddonkey

Copy link
Copy Markdown
Contributor Author

I retrieved the landed OpenCodex carry directly through the GitHub API to clarify the dependency-inspection limitation in the latest review.

#3762 is merged, with immutable merge commit f00f2bcaea251ebe7ad4de9e38337b4be0ccee47. I inspected that commit's source, not a floating branch:

For content identity, the inspected schema/persistence blob is 257cb86d3ef8419bc2fbb8ad1a4afa765742f1b0, and the stamping-helper blob is a4c942bd883bb81f3cd08e392c6afe2cc8c8d3ec.

This verifies the landed source contract. The executable capture remains pinned to 146ed679, as the evidence report explicitly states; I am not claiming a new capture of the carry, released-version coverage, or live vendor billing evidence. Please distinguish the review environment's failed dependency-page retrieval from a demonstrated compatibility defect. The owner display-policy decision remains open.

@clawsweeper re-review

@clawsweeper

clawsweeper Bot commented Sep 8, 2026

Copy link
Copy Markdown

🦞🧹
ClawSweeper re-review requested.

I asked ClawSweeper to review this item again.
Action: item re-review queued (workflow sweep.yml, event exact_review_queue).
Result: when the review finishes, ClawSweeper will create the durable review comment if needed or update the existing comment in place.

Re-review progress:

@olddonkey

Copy link
Copy Markdown
Contributor Author

Final validation for c35734ad5a747a83c46b3844c8ce33fd46dbefe1:

  • All nine GitHub CI checks passed, including both macOS shards and all Linux builds.
  • Local make check passed with zero violations in 2,143 files.
  • Local make test passed all 1,036 selections in 87 groups on the first attempt, with zero failures, retries, or timeouts.
  • Cache-adoption, pricing/provenance regressions and the 41-test architecture gate passed. The generated parser hash is c3a879df4eff7187, with current-main and previous-PR cache compatibility retained.
  • GitHub reports this head mergeable/clean; there are no unresolved inline threads. The latest ClawSweeper review reports no actionable code findings.

The PR body now has the final validation results and the immutable landed-producer source audit. The owner display-policy decision and older changes-requested review remain unresolved. The review environment's inability to retrieve the later producer carry is also recorded; the supplied source audit does not expand the pinned executable capture into a claim about a released producer version.

@olddonkey

Copy link
Copy Markdown
Contributor Author

Merged latest main 8b254dbec11ddd5c5547878d9640e4e965306c71 into this branch and pushed 40f9e47f7373d69f7738c582a9dc288231f329fd. GitHub now reports the branch mergeable.

The conflict resolution preserves upstream's checked numeric aggregation, parser-revision migrations, and Grok terminal-billing work gate while retaining completed-turn accounting, recorded-versus-estimated disclosure, custom-price precedence, and per-attempt OAuth attribution. The generated parser hash is a8559238a5fc0480; current-main and prior-PR SQLite adoption are covered. Release notes now sit under 0.59.1 Unreleased, with published sections identical to main.

Validation on this head:

  • make check: zero violations in 2,198 Swift files.
  • Focused tests: 433 tests in 35 suites passed, including nine upstream billing-failure scenarios and the added OAuth overflow/estimate-coverage regression.
  • make test: 1,080 selections in 90 groups passed on the first attempt, with zero failures, retries, or timeouts.
  • The pinned producer ledger still imports five OAuth tokens, excludes API-key traffic, and reopens the cache without rereading the log.
  • GitHub CI is running: https://github.com/steipete/CodexBar/actions/runs/34674108255

The PR body records the current results and distinguishes them from the historical native-corpus proof. The owner display-policy decision remains open.

@clawsweeper re-review

@clawsweeper

clawsweeper Bot commented Sep 12, 2026

Copy link
Copy Markdown

🦞🧹
ClawSweeper re-review requested.

I asked ClawSweeper to review this item again.
Action: item re-review queued (workflow sweep.yml, event exact_review_queue).
Result: when the review finishes, ClawSweeper will create the durable review comment if needed or update the existing comment in place.

@olddonkey

Copy link
Copy Markdown
Contributor Author

Addressed the newly reported native overflow finding in 6d3d5af1bb28aeaeb2bdb13cfcc70cdd4d0b2cc7.

The native parser now validates NSNumber conversions, preserves unknown counts, and uses checked addition across model/day/window totals. Later valid records and cache rereads cannot clear an unknown aggregate. Explicit valid totals retain precedence, valid neighboring token classes and recorded spend remain available, and incomplete token accounting does not establish full coverage. The added production-scanner regression suite covers per-record and cross-record overflow, multiple days, nested models, malformed number types, explicit totals, snapshot projection, and cache reuse.

Validation:

  • Final make check: zero violations in 2,199 files.
  • Focused regression run: 326 tests in 30 suites passed.
  • make test: 1,081 selections in 91 groups passed on the first attempt, zero failures/retries/timeouts. Only test-call formatting changed afterward; it was rebuilt and rechecked by the final proof run.
  • Final serial proof: 31 tests in 6 suites passed, including fresh local native logs through menu/dashboard projections. The 30-day window retains 98,631,812 tokens and $12.94957366 with recorded provenance; empty 1-day/7-day windows remain unknown. No provider authentication or live billing is involved.
  • The pinned producer ledger again imports five OAuth tokens, excludes API-key traffic, and reopens the cache without log reads; cancellation releases the scan queue within one second.

The PR body has the current results. The new CI run is pending; the owner dollar-display decision remains open.

@clawsweeper re-review

@clawsweeper

clawsweeper Bot commented Sep 12, 2026

Copy link
Copy Markdown

🦞🧹
ClawSweeper re-review requested.

I asked ClawSweeper to review this item again.
Action: item re-review queued (workflow sweep.yml, event exact_review_queue).
Result: when the review finishes, ClawSweeper will create the durable review comment if needed or update the existing comment in place.

@olddonkey

Copy link
Copy Markdown
Contributor Author

Addressed the remaining menu-projection overflow finding in a41736133e7a2d922db7bbc5b0e18e7e729e855d.

A checked, complete-count helper now covers the remote-backed Grok window projection, window request totals, comparison summaries, and menu/Widget fallback totals. Unknown daily values are retained as unknown instead of being dropped into a partial sum. This complements the native scanner repair in the preceding commit.

The new end-to-end regression writes native completed-turn logs and exercises both remote-backed and fallback live consumers. It covers individually representable Int.max/1 totals across days and an already-unknown day followed by a valid day. The menu, Widget, and dashboard keep the aggregate unavailable; narrowing to the unaffected one-day window restores its known count. All four scenarios passed.

Current-head validation: make check has zero violations in 2,199 files; 888 focused tests in 110 suites passed; 40 serial native-window/producer-import/projection/cancellation proof tests passed. The fresh 30-day native window still reports 98,631,812 tokens and $12.94957366 with recorded provenance and disclosures. Empty short windows remain unknown. The full suite and new GitHub CI are running, with current results distinguished in the PR body.

The owner's default dollar-display decision remains open.

@clawsweeper re-review

@clawsweeper

clawsweeper Bot commented Sep 12, 2026

Copy link
Copy Markdown

🦞🧹
ClawSweeper re-review requested.

I asked ClawSweeper to review this item again.
Action: item re-review queued (workflow sweep.yml, event exact_review_queue).
Result: when the review finishes, ClawSweeper will create the durable review comment if needed or update the existing comment in place.

@olddonkey

Copy link
Copy Markdown
Contributor Author

Final author-side validation for a41736133e7a2d922db7bbc5b0e18e7e729e855d is complete:

  • make check: zero violations in 2,199 files.
  • make test: all 1,081 selections in 91 groups passed on the first attempt, with zero failures, retries, or timeouts. No source changed afterward.
  • 888 focused tests and 40 current-source native/producer/projection/cancellation proof tests passed.
  • GitHub lint, changes, GitGuardian, and Linux x64/ARM64/musl builds/tests passed: https://github.com/steipete/CodexBar/actions/runs/34675986313
  • Both macOS CI shards remain queued waiting for runners, so I am not claiming a complete GitHub CI pass.
  • GitHub reports the head mergeable; there are no unresolved inline threads. The latest exact-head review reports Platinum 4/6 and no actionable findings.

The PR body now has the final local validation and exact CI limitation. The default-dollar display policy and the owner's outstanding changes-requested review remain maintainer decisions.

@clawsweeper re-review

@clawsweeper

clawsweeper Bot commented Sep 12, 2026

Copy link
Copy Markdown

🦞🧹
ClawSweeper re-review requested.

I asked ClawSweeper to review this item again.
Action: item re-review queued (workflow sweep.yml, event exact_review_queue).
Result: when the review finishes, ClawSweeper will create the durable review comment if needed or update the existing comment in place.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

merge-risk: 🚨 compatibility 🚨 Merging this PR could break existing users, config, migrations, defaults, or upgrades. P2 Normal priority bug or improvement with limited blast radius. proof: sufficient Contributor real behavior proof is sufficient. rating: 🐚 platinum hermit Good normal PR readiness with ordinary maintainer review expected. status: 👀 ready for maintainer look ClawSweeper has no concrete contributor-facing blocker left for this PR.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants