Conversation
…vements Implements four major feature requests for enhanced API observability, security, and transaction simulation capabilities: ASTROIDX556#247 - Prometheus Metrics Interceptor - Create MetricsInterceptor for HTTP request metrics collection - Add request duration histograms, active request gauges, and request counters - Categorize metrics by route, method, and status code - Exclude /metrics endpoint from self-instrumentation - Add comprehensive unit tests (9 tests) ASTROIDX556#246 - Stellar Transaction Simulation Service - Create StellarSimulationService for transaction simulation - Integrate with Soroban RPC for XDR validation and simulation - Include risk assessment and fee estimation - Add circuit breaker protection for RPC failures - Add XDR validation helper method - Add comprehensive unit tests with mocked Stellar RPC (24 tests) ASTROIDX556#245 - Cryptographic API Key Hashing Upgrade - Upgrade from SHA-256 to Argon2id for enhanced security - Implement memory-hard algorithm resistant to GPU/ASIC attacks - Add timing-attack resistant comparison via Argon2 verification - Maintain SHA-256 fallback for backward compatibility - Update ApiKeyService with dual-algorithm verification - Add comprehensive unit tests (32 crypto tests, 15 API key tests) ASTROIDX556#244 - Webhook Retry and Dead Letter Queue - Verify existing implementation meets all requirements - Confirm exponential backoff with jitter (2000ms base, 20% jitter) - Confirm 5 max attempts and non-transient error detection - Verify dead-letter handler via DeadLetterService - All existing tests passing Closes ASTROIDX556#247, ASTROIDX556#246, ASTROIDX556#245, ASTROIDX556#244 Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
… logging Add runWorkerJob, a wrapper every background worker now routes its handler through. It times the job via WorkerMetricsService when available, logs a structured completion record, and classifies failures before rethrowing: transient failures log a job.retrying warning, while a failure on the final attempt or an UnrecoverableError logs a job.dead-lettered error with the scrubbed payload and stack. The original error is always rethrown untouched so BullMQ retry semantics are preserved, and logging can never mask it. Add scrubForLog/scrubString, which redact sensitive keys (sharing the audit sanitizer's key list) and secret-shaped substrings such as Stellar seeds, bearer tokens and URL credentials, and coerce cycles, bigints and errors into JSON-safe values. QueueFailureListener now scrubs its log line as well; previously the raw webhook payload, including its signing secret, was written to the error log. The dead-letter copy keeps the raw payload so re-drive still works.
Extend the shared list query contract used by every list endpoint with an offset parameter alongside the existing page parameter, and bound the page size. - limit defaults to 50 and is capped at 200 - offset defaults to 0; page remains supported as an alternative, and supplying both is rejected - negative, non-integer, non-numeric or out-of-range values are rejected by the Zod pipe with 400 Bad Request - Prisma queries use bound skip/take values and an allow-listed sort column, so no pagination input reaches SQL as text - meta now includes offset, and hasNext/hasPrev are computed from the offset so unaligned slices report correctly - paginated responses set an X-Total-Count header, exposed via CORS - replace duplicated Swagger page/limit docs with ApiPaginationQuery - add unit tests and an HTTP-level integration test
|
@Deb-Auth Great news! 🎉 Based on an automated assessment of this PR, the linked Wave issue(s) no longer count against your application limits. You can now already apply to more issues while waiting for a review of this PR. Keep up the great work! 🚀 |
Merge Conflict — Action NeededThis pull request has merge conflicts with the base branch ( What to do: update your branch by merging or rebasing against |
Merge Conflict — Action NeededThis pull request has merge conflicts with the base branch ( What to do: update your branch by merging or rebasing against |
…ASTROIDX556#338, ASTROIDX556#339) (ASTROIDX556#369) * perf(analytics): batch dashboard overview queries into one transaction Closes ASTROIDX556#338 AnalyticsService.overview() issued 7 independent queries (counts, spend aggregates, status/risk group-bys) via Promise.all, each its own roundtrip. Batches the 3 counts and 2 aggregates into a single $transaction([...]) call; the 2 groupBy calls stay outside the batch since Prisma's groupBy return type doesn't infer correctly inside a $transaction array. Response shape is unchanged. Existing indexes on organizationId/status/createdAt already cover these queries. * test(audit): cover pagination and sorting edge cases for activity log Closes ASTROIDX556#339 The audit log (this repo's activity log) already supported page/limit/sort/order/filter query params and returned pagination metadata (total, totalPages, hasNext, hasPrev), with limit capped at 100. Adds unit test coverage for the previously-untested list() method: normal pagination, empty results, out-of-bounds pages, invalid sort field fallback, ascending order, and entity filtering. * fix(migrations): resolve colliding timestamp between two merged migrations Migrations 20260928120000_add_agent_contribution_stats_index and 20260928120000_add_notifications_user_created_at_index landed with the same 14-digit timestamp prefix from two separately merged PRs (ASTROIDX556#374, ASTROIDX556#357), which scripts/verify-migrations.sh rejects as a conflict. Bumps the notifications index migration to 20260928120001; both migrations are independent, additive CREATE INDEX statements with no ordering dependency between them, so the rename is safe. * fix(ci): document missing env vars and fix flaky retry.util test Two pre-existing, unrelated-to-this-PR CI failures fixed while unblocking this branch: - docs/configuration.md was missing 7 env vars added by recent merges (DATABASE_SLOW_QUERY_THRESHOLD_MS, DATABASE_CONNECT_RETRY_ATTEMPTS, DATABASE_CONNECT_RETRY_DELAY_MS, PUBLIC_RATE_LIMIT_*), which env.validation.spec.ts asserts against. Documented all 7. - retry.util.spec.ts had 3 tests that create a rejecting promise, advance fake timers with vi.runAllTimersAsync(), then attach the rejection assertion afterward — a race that surfaces as an unhandled rejection under full-suite load (deterministic once >100 files run together). Attaching a no-op .catch() immediately after creating the promise prevents the unhandled state without changing what each test asserts.
…r Transaction Risk Scoring (ASTROIDX556#381)
…utating Operations (ASTROIDX556#384)
…ensitive API Endpoints (ASTROIDX556#383)
Wires the existing migration status checker into app bootstrap so the process halts before accepting traffic when prisma/migrations has pending or failed migrations (DATABASE_MIGRATION_CHECK_MODE=halt, the default; 'warn' logs and continues). Gated behind DATABASE_MIGRATION_CHECK_ENABLED. Adds a db_pool_connections Prometheus gauge (active/idle/waiting) sourced from pg_stat_activity, since Prisma's Rust query engine doesn't expose pool internals through the Node client. Closes ASTROIDX556#335 Closes ASTROIDX556#332
- docs/configuration.md was missing entries for DATABASE_SLOW_QUERY_THRESHOLD_MS, DATABASE_CONNECT_RETRY_ATTEMPTS, DATABASE_CONNECT_RETRY_DELAY_MS, and the PUBLIC_RATE_LIMIT_* vars, failing the configuration-documentation test. - retry.util.spec.ts left three rejected promises unhandled between `runAllTimersAsync()` and the `expect(...).rejects` assertion that attaches the handler; attach a no-op .catch() immediately after creating each promise so fake-timer-driven rejections don't fire as unhandled rejections mid-test-run.
main's build/typecheck/lint/test were already broken before this branch touched anything (confirmed by checking out upstream/main directly). CI enforces these repo-wide, so they block this PR too. Fixed each: - event-names.ts: duplicate object key (TransactionRiskScoringRequested) - throttler.guard.ts: read AuthenticatedUser.sub, a field that doesn't exist on that type (JWT payload field name leaked into the wrong type) - sliding-window-throttler.guard.ts: removed a user-tier rate-limit multiplier keyed on AuthenticatedUser.tier, a field never present anywhere in the auth/user model — dead, unbacked logic. Removed its now-orphaned tests too. - agent.controller.ts: AstroidThrottlerGuard used but never imported; dropped an unused SlidingWindowThrottlerGuard import block - risk.service.ts: event-driven risk scoring built a RiskFactorsInput with fields (destination/velocityCount/isNewRecipient) that don't exist on the current type; mapped to the real shape instead - risk.service.spec.ts: removed orphaned unused fixtures - stellar.service.ts (src/modules/stellar/services, unused elsewhere in the app but still typechecked/tested): getTransactionInfo called a client method that doesn't exist (real method is getTransaction); simulateTransaction passed a bare string where the client expects an options object - stellar.service.spec.ts: rewritten against the real SorobanSimulationResult shape; fixed mockResolvedValueOnce/ mockRejectedValueOnce being consumed by the test's own first assertion, leaving the second call unmocked - transaction.service.spec.ts: rewritten against TransactionService's actual create() contract (it doesn't call Soroban simulation at all; the previous spec tested a flow that was never implemented) and a real Ed25519 checksum address - sensitive-rate-limit.integration.spec.ts: app.inject() doesn't exist on this Express-platform app; switched to app.listen + fetch, matching the sibling public-rate-limit.integration.spec.ts pattern, and named the test throttler 'api' so AstroidThrottlerGuard's tier-matching actually engages it
CI runs lint and test repo-wide, so these also blocked the PR: - 4 pre-existing no-explicit-any lint errors in throttler guard code and specs, typed properly instead of suppressed - 4 THROTTLE_* env vars (WEBHOOK_LIMIT, API_BURST, AUTH_BURST, WEBHOOK_BURST) were validated by the env schema but missing from docs/configuration.md, failing the docs-sync test
Every authenticated request performed a Redis round trip against the token blacklist to answer "is this session still revoked?". This adds a short-TTL caching layer in front of the blacklist so repeated verifications within one window skip the Redis query entirely. - Add CacheService: a small get/set/delete cache over the shared REDIS_CLIENT with TTL-bounded entries, JSON payloads, and SCAN-based prefix invalidation. Every operation degrades to a no-op/miss on Redis failure so caching can never break the request path. - Add TokenVerificationCacheService: caches per-session revocation answers for TOKEN_CACHE_TTL seconds (default 30, well below the 15-minute access-token lifetime) and exposes invalidation hooks. - Wire the cache into JwtStrategy.validate (cache-first, source of truth on miss, fail-open unchanged on Redis outages). - Hook invalidation into every revocation path: TokenBlacklistService drops the cached answer after each blacklist write (including on Redis-outage fallback), AuthService invalidates on logout and on refresh rotation, so revocations are observed immediately instead of after the TTL window. Revocation reliability is preserved because no cached answer outlives its TTL, and explicit logout/rotation clears the entry at once. Tests: unit suites for CacheService and TokenVerificationCacheService (hits, misses, resolver fallback, invalidation hooks), an integration suite proving repeated authentications trigger a single blacklist lookup and that logout flips a cached-valid session to 401 immediately, plus updated JwtStrategy/TokenBlacklistService/api-key integration suites for the new wiring. Closes ASTROIDX556#341 🤖 Generated with Codebuff Co-Authored-By: Codebuff <noreply@codebuff.com>
…ic routes Public endpoints are the first thing abusive traffic hits, but their limits were a single fixed IP budget: every public route shared PUBLIC_RATE_LIMIT_MAX_REQUESTS per window, and all clients behind one shared address (NAT, office egress, CI runners) exhausted one bucket together. This makes the public rate limiter configurable per route and per client identifier, on the existing Redis sliding-window counter. - Add @PublicRateLimit(max, windowSeconds) decorator: per-route (or per-controller) budget overrides resolved by the guard through Reflector; the global PUBLIC_RATE_LIMIT_* settings remain the default. - Add PUBLIC_RATE_LIMIT_CLIENT_IDENTIFIERS (optional, comma-separated, currently 'apiKey'): when enabled, a presented x-api-key or ApiKey/Bearer ak_... Authorization header is folded into the bucket key so distinct key-holding clients behind one IP get their own budgets. The IP always participates; keyless callers share the plain-IP bucket as before. - Extract a testable PublicRateLimitGuard.check() returning the full decision (allowed, limit, windowSeconds, count, resetAt); canActivate keeps its existing 429 + X-RateLimit-Limit/Remaining/Reset + Retry-After contract and in-memory fallback on Redis outage. Tests: unit suites for per-route rule resolution, identifier bucketing (with/without identifiers configured), header correctness on allowed and limited requests, and @SkipPublicRateLimit() interaction with rules; an HTTP-level integration suite simulating bursts that proves the 429-with-headers behaviour at the global limit, per-route overrides (next to unaffected sibling routes), and per-key budget isolation. Existing public-rate-limit suites pass unchanged (bucket keys keep the ip: prefix, so stored counters stay compatible). Closes ASTROIDX556#342 🤖 Generated with Codebuff Co-Authored-By: Codebuff <noreply@codebuff.com>
…-spending-limit-guard Feat/agent spending limit guard
…ues-63-64-283 feat: stream audit exports and harden payment validation
…-handling Add centralized error handling and scrubbed structured logging for background workers
…60-draft feat: address requested issues - closes ASTROIDX556#260, closes ASTROIDX556#259, closes ASTROIDX556#258, closes ASTROIDX556#257
…atch-recovery-validation-retry Add query metrics, batch retry recovery, input sanitization, and HTTP retry with backoff
feat: Add comprehensive observability, security, and simulation improvements
…57-61 feat: correlate requests, paginate audit, sign webhooks
…-264-draft feat: address requested issues - closes ASTROIDX556#264, closes ASTROIDX556#263, closes ASTROIDX556#262, closes ASTROIDX556#261
CI Checks Failed — Action NeededThe current CI run has failing checks, so this PR remains unmergeable. Please inspect the failed jobs, fix the underlying issues, and push the correction to this PR branch. Failing checks:
I will not merge while these checks are failing. |
Summary
Adds
offset/limitpagination with bounded page sizes to the shared list query contract. Every list endpoint now acceptsoffset, rejects invalid bounds with 400, and reports the total row count inmetaand in anX-Total-Countheader.Why
This repo has no
GET /api/resources. Instead, all 12 list endpoints (agents, wallets, transactions, policies, budgets, approvals, audit, API keys, memory, notifications, organization members, webhooks) parse the samepaginationQuerySchemaand build Prisma arguments throughtoPrismaPagination. Changing that shared contract applies the issue's requirements uniformly, rather than to one endpoint.What changed
src/common/helpers/pagination.tslimitnow defaults to 50 and is capped at 200 (DEFAULT_PAGE_LIMIT,MAX_PAGE_LIMIT). It was previously 20, capped at 100.offsetparameter (default 0). The existingpageparameter still works, so current clients are not broken. Sending both is rejected.offsetandpageare always both populated, so services don't need to know which one the client sent.limit/page), fractional, non-numeric or over-cap values fail Zod validation. They surface as400 Bad Requestthrough the existingZodValidationPipe.toPrismaPaginationmapsoffset/limitto Prismaskip/take, which are bound as parameters.sortremains restricted to each endpoint's allow-list, so no pagination input reaches SQL as text.buildPaginationMeta(total, query)now includesoffset.hasNext/hasPrevare computed from the offset, so offsets that aren't page-aligned report correctly.buildPaginationMeta. This is a one-line change each.ResponseInterceptor: setsX-Total-Counton everyPaginatedresponse, andmain.tsexposes that header via CORS so browser clients can read it.@ApiPaginationQuery()decorator documentsoffset,page,limit,sort,order, the header and the 400 response. It replaces the duplicated@ApiQuerypairs on the 12 controllers.API_DOCUMENTATION.mdpagination sections are updated.Type of change
Testing
src/common/helpers/pagination.spec.ts(new) covers:src/common/helpers/pagination.integration.spec.ts(new) boots a Nest app with the real pipe, interceptor and exception filter and sends real HTTP requests. It checks:X-Total-Countlimit=200acceptedBehaviour change to note
Clients that relied on the implicit default of 20 rows now get 50. Requests with
limitbetween 101 and 200 now succeed instead of returning 400.Related issue
Closes #344
Checklist
npm run buildpassesnpm testpassesnpm run lintpassesnpm run typecheckpasses