Skip to content

feat(shutdown): graceful shutdown for HTTP, BullMQ workers, queues, Redis and Prisma - #365

Open
AdaBliss wants to merge 4 commits into
ASTROIDX556:mainfrom
AdaBliss:feat/60-graceful-shutdown
Open

AdaBliss wants to merge 4 commits into
ASTROIDX556:mainfrom
AdaBliss:feat/60-graceful-shutdown

Conversation

@AdaBliss

Copy link
Copy Markdown
Contributor

closes #60

Summary

Adds coordinated graceful shutdown so deployments no longer abandon in-flight HTTP requests or BullMQ jobs, and all queue, Redis and Prisma connections close in a deterministic order.

Previously, SIGTERM killed the process immediately: Nest shutdown hooks were not enabled, and the only hook was a beforeExit listener that never fires on a signal.

Shutdown order

ShutdownCoordinator (src/common/shutdown) is installed from src/main.ts after app.listen(). On SIGTERM or SIGINT it runs:

Step What happens Bound
1. HTTP Stop accepting connections, close idle keep-alive sockets, send Connection: close on in-flight responses, wait for them Grace period
2. Workers worker.pause() on every BullMQ worker (no new jobs; waits for active ones), then worker.close() Grace period (shared with step 1)
3. Lifecycle hooks app.close() runs every Nest lifecycle hook, so each provider releases what it owns 15s
4. Queues BullMQ queues closed 3s per resource
5. Redis Shared and auth Redis clients closed with QUIT 3s per resource
6. Database Both Prisma pools disconnected 3s per resource

Steps 4-6 run from the coordinator's beforeApplicationShutdown, so they also run when a module is closed outside a signal (tests, scripts).

Default grace period: SHUTDOWN_GRACE_PERIOD_MS=20000. This leaves about 10s for steps 3-6 inside Kubernetes' default 30s terminationGracePeriodSeconds. It is validated in src/config/env.validation.ts and exposed as the shutdown config namespace (src/config/shutdown.config.ts).

Behavior

  • Clean drain: exits 0.
  • Grace period expires: remaining HTTP connections are destroyed, and busy workers are force-closed with worker.close(true). Their job locks lapse and BullMQ's stalled-job recovery re-queues them. A clear error is logged, cleanup still completes, and the process exits 1.
  • A resource fails or hangs while closing: the failure is logged, later phases still close, and the process exits 1.
  • Repeated signals: logged and ignored. Cleanup runs once, and every step is bounded, so shutdown cannot hang.
  • Logs: resource names, phases and timings only. No job payloads, no connection strings.

Why not app.enableShutdownHooks()

Nest's implementation:

  • runs onModuleDestroy, which disconnects Prisma, before it closes the HTTP server;
  • has no grace period;
  • exits by re-raising the signal, which gives a non-zero status.

The coordinator installs its own handlers and still runs every Nest lifecycle hook through app.close().

Ownership (dependency injection and config conventions)

Providers keep ownership of their connections and register how to close them. The coordinator only decides when.

  • PrismaService registers in the database phase. Under the coordinator, its onModuleDestroy no longer disconnects early; without a coordinator it behaves as before.
  • The LocksModule factory that creates the shared REDIS_CLIENT registers it. RedisLock no longer disconnects a client it does not own; previously that happened in onModuleDestroy, before queues were closed.
  • The auth module's Redis provider registers too; previously it was never closed.
  • Workers (@Processor) and queues (BullModule.registerQueue) created by @nestjs/bullmq are discovered at bootstrap. No domain module listens to process signals.

Testing

  • shutdown-coordinator.service.spec.ts (17 cases) uses mocked HTTP server, workers, queues, Redis and Prisma, and invokes the lifecycle directly (no OS signals). It covers:
    • clean drain and exact close order;
    • in-flight requests and jobs holding back later phases;
    • worker grace timeout leading to a force close and exit 1;
    • HTTP grace timeout;
    • the grace period shared across steps 1-2;
    • a resource close that fails (later phases still close, and payloads are not logged);
    • a resource close that hangs;
    • a worker close failure;
    • hung lifecycle hooks;
    • repeated signals leading to a single cleanup and a single exit;
    • reverse-order resource closure without a signal;
    • Connection: close during a drain.
  • Tests for closeRedisClient and for PrismaService's coordinated and uncoordinated disconnect.
  • npm run typecheck, npm run lint, npm test (1105 passing), npm run build.

Manual validation

Ran the built server (dist/main.js) against real PostgreSQL and Redis 5, then delivered SIGTERM twice. On Windows, process.emit fires the same listeners a real signal would. Log output:

Received SIGTERM; shutting down (grace period 20000ms)
HTTP: stopped accepting connections and drained in-flight requests (0ms)
Received SIGTERM while shutdown is already in progress; ignoring
Worker "webhooks": drained and closed (128ms)
Worker "webhooks": drained and closed (133ms)
Worker "audit-cleanup": drained and closed (132ms)
Worker "dead-letter": drained and closed (132ms)
Closed queues resource "queue:webhooks" (0ms)
Closed queues resource "queue:audit-cleanup" (0ms)
Closed queues resource "queue:dead-letter" (0ms)
Closed redis resource "redis:auth" (0ms)
Closed redis resource "redis:shared" (0ms)
Closed database resource "prisma" (37ms)
Shutdown complete in 188ms

The process exited with code 0. A handle inspection 200ms after the exit call found 0 open non-stdio handles, so there were no open-handle warnings. Startup and npm run start:dev are unchanged.

Follow-ups, out of scope

  • SlidingWindowThrottlerGuard, AgentRateLimiterGuard and BalanceCacheService each construct their own ioredis client outside DI and never close it. Moving them onto the shared client (or registering them) is a separate refactor. The final process.exit still releases them.
  • Implement secure environment variable validation at startup #350 (environment validation, opened alongside this PR) adds a unified environmentSchema. Once both are merged, add shutdownEnvSchema to it and a row to docs/configuration.md. The branches merge without textual conflicts (verified).

…and Prisma in order

Add a ShutdownCoordinator that owns SIGTERM/SIGINT handling and runs a
fixed, bounded sequence:

1. HTTP: stop accepting connections, close idle keep-alive sockets, mark
   in-flight responses `Connection: close` and wait for them.
2. Workers: pause every BullMQ worker (no new jobs) and wait for active
   jobs, then close it. Steps 1-2 share SHUTDOWN_GRACE_PERIOD_MS (default
   20000). On expiry, remaining connections are destroyed and busy workers
   are force-closed, so their jobs are retried via stalled-job recovery.
3. app.close(): run every Nest lifecycle hook. From
   beforeApplicationShutdown the coordinator closes registered resources
   phase by phase: queues -> redis -> database, each bounded.

The process exits 0 after a clean drain and 1 if the grace period expired
or any step failed. Repeated signals are ignored while shutdown runs once.
Logs carry resource names and timings only, never job data or URLs.

Nest's app.enableShutdownHooks() is not used: it disconnects Prisma in
onModuleDestroy before the HTTP server stops, has no grace period and exits
by re-raising the signal.

Providers keep ownership of their connections and register how to close
them: PrismaService (database), the shared LocksModule Redis client and
the auth Redis client (redis). Queues and workers created by
@nestjs/bullmq are discovered at bootstrap. RedisLock no longer
disconnects the shared client it does not own, and PrismaService's
beforeExit hook is replaced by the coordinator.

Closes ASTROIDX556#60
@drips-wave

drips-wave Bot commented Sep 28, 2026

Copy link
Copy Markdown

@AdaBliss Great news! 🎉 Based on an automated assessment of this PR, the linked Wave issue(s) no longer count against your application limits.

You can now already apply to more issues while waiting for a review of this PR. Keep up the great work! 🚀

Learn more about application limits

@Cjay-Cyber-2

Copy link
Copy Markdown
Contributor

Merge Conflict — Action Needed

This pull request has merge conflicts with the base branch (main).

What to do: update your branch by merging or rebasing against main, resolve any conflicts locally, and push the result.

@Cjay-Cyber-2

Copy link
Copy Markdown
Contributor

Merge Conflict — Action Needed

This pull request has merge conflicts with the base branch (main).

What to do: update your branch by merging or rebasing against main, resolve any conflicts locally, and push the result.

@Cjay-Cyber-2

Copy link
Copy Markdown
Contributor

CI fails in migration verification for the same duplicate Prisma migration timestamp: 20260928120000_add_agent_contribution_stats_index and 20260928120000_add_notifications_user_created_at_index. Rename one migration to a unique 14-digit timestamp and rerun build-and-test.

@Cjay-Cyber-2

Copy link
Copy Markdown
Contributor

CI Checks Failed — Action Needed

The current CI run has failing checks, so this PR remains unmergeable. Please inspect the failed jobs, fix the underlying issues, and push the correction to this PR branch.

Failing checks:

  • build-and-test: FAILURE

I will not merge while these checks are failing.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Implement graceful shutdown for HTTP, Prisma, Redis, and workers

2 participants