Conversation
…and Prisma in order Add a ShutdownCoordinator that owns SIGTERM/SIGINT handling and runs a fixed, bounded sequence: 1. HTTP: stop accepting connections, close idle keep-alive sockets, mark in-flight responses `Connection: close` and wait for them. 2. Workers: pause every BullMQ worker (no new jobs) and wait for active jobs, then close it. Steps 1-2 share SHUTDOWN_GRACE_PERIOD_MS (default 20000). On expiry, remaining connections are destroyed and busy workers are force-closed, so their jobs are retried via stalled-job recovery. 3. app.close(): run every Nest lifecycle hook. From beforeApplicationShutdown the coordinator closes registered resources phase by phase: queues -> redis -> database, each bounded. The process exits 0 after a clean drain and 1 if the grace period expired or any step failed. Repeated signals are ignored while shutdown runs once. Logs carry resource names and timings only, never job data or URLs. Nest's app.enableShutdownHooks() is not used: it disconnects Prisma in onModuleDestroy before the HTTP server stops, has no grace period and exits by re-raising the signal. Providers keep ownership of their connections and register how to close them: PrismaService (database), the shared LocksModule Redis client and the auth Redis client (redis). Queues and workers created by @nestjs/bullmq are discovered at bootstrap. RedisLock no longer disconnects the shared client it does not own, and PrismaService's beforeExit hook is replaced by the coordinator. Closes ASTROIDX556#60
|
@AdaBliss Great news! 🎉 Based on an automated assessment of this PR, the linked Wave issue(s) no longer count against your application limits. You can now already apply to more issues while waiting for a review of this PR. Keep up the great work! 🚀 |
Merge Conflict — Action NeededThis pull request has merge conflicts with the base branch ( What to do: update your branch by merging or rebasing against |
Merge Conflict — Action NeededThis pull request has merge conflicts with the base branch ( What to do: update your branch by merging or rebasing against |
|
CI fails in migration verification for the same duplicate Prisma migration timestamp: |
CI Checks Failed — Action NeededThe current CI run has failing checks, so this PR remains unmergeable. Please inspect the failed jobs, fix the underlying issues, and push the correction to this PR branch. Failing checks:
I will not merge while these checks are failing. |
closes #60
Summary
Adds coordinated graceful shutdown so deployments no longer abandon in-flight HTTP requests or BullMQ jobs, and all queue, Redis and Prisma connections close in a deterministic order.
Previously,
SIGTERMkilled the process immediately: Nest shutdown hooks were not enabled, and the only hook was abeforeExitlistener that never fires on a signal.Shutdown order
ShutdownCoordinator(src/common/shutdown) is installed fromsrc/main.tsafterapp.listen(). OnSIGTERMorSIGINTit runs:Connection: closeon in-flight responses, wait for themworker.pause()on every BullMQ worker (no new jobs; waits for active ones), thenworker.close()app.close()runs every Nest lifecycle hook, so each provider releases what it ownsQUITSteps 4-6 run from the coordinator's
beforeApplicationShutdown, so they also run when a module is closed outside a signal (tests, scripts).Default grace period:
SHUTDOWN_GRACE_PERIOD_MS=20000. This leaves about 10s for steps 3-6 inside Kubernetes' default 30sterminationGracePeriodSeconds. It is validated insrc/config/env.validation.tsand exposed as theshutdownconfig namespace (src/config/shutdown.config.ts).Behavior
0.worker.close(true). Their job locks lapse and BullMQ's stalled-job recovery re-queues them. A clear error is logged, cleanup still completes, and the process exits1.1.Why not
app.enableShutdownHooks()Nest's implementation:
onModuleDestroy, which disconnects Prisma, before it closes the HTTP server;The coordinator installs its own handlers and still runs every Nest lifecycle hook through
app.close().Ownership (dependency injection and config conventions)
Providers keep ownership of their connections and register how to close them. The coordinator only decides when.
PrismaServiceregisters in thedatabasephase. Under the coordinator, itsonModuleDestroyno longer disconnects early; without a coordinator it behaves as before.LocksModulefactory that creates the sharedREDIS_CLIENTregisters it.RedisLockno longer disconnects a client it does not own; previously that happened inonModuleDestroy, before queues were closed.Redisprovider registers too; previously it was never closed.@Processor) and queues (BullModule.registerQueue) created by@nestjs/bullmqare discovered at bootstrap. No domain module listens to process signals.Testing
shutdown-coordinator.service.spec.ts(17 cases) uses mocked HTTP server, workers, queues, Redis and Prisma, and invokes the lifecycle directly (no OS signals). It covers:Connection: closeduring a drain.closeRedisClientand forPrismaService's coordinated and uncoordinated disconnect.npm run typecheck,npm run lint,npm test(1105 passing),npm run build.Manual validation
Ran the built server (
dist/main.js) against real PostgreSQL and Redis 5, then deliveredSIGTERMtwice. On Windows,process.emitfires the same listeners a real signal would. Log output:The process exited with code 0. A handle inspection 200ms after the exit call found 0 open non-stdio handles, so there were no open-handle warnings. Startup and
npm run start:devare unchanged.Follow-ups, out of scope
SlidingWindowThrottlerGuard,AgentRateLimiterGuardandBalanceCacheServiceeach construct their own ioredis client outside DI and never close it. Moving them onto the shared client (or registering them) is a separate refactor. The finalprocess.exitstill releases them.environmentSchema. Once both are merged, addshutdownEnvSchemato it and a row todocs/configuration.md. The branches merge without textual conflicts (verified).