Repository navigation
Move to ECS Fargate: containerised app, Terraform, OIDC deploy pipeline - #37
Merged
Merged
Conversation
…rage - Verify the password with bcrypt in the local passport strategy; reject OAuth-created accounts that have no local password - Require a session on every /api/* data endpoint (QuickBooks, Xero, Excel) and on the QuickBooks sync routes. With no session the userid filter was undefined, which Prisma drops, so these returned every user's rows - Store sessions in Postgres via connect-pg-simple (new "session" table migration) instead of the in-process MemoryStore - Fail fast at startup when DATABASE_URL, SESSION_SECRET (>=32 chars, not a placeholder) or ENCRYPTION_KEY (exactly 32 bytes) are missing or weak; remove the hardcoded session-secret fallbacks - Fix .env.example ENCRYPTION_KEY placeholder, which was 34 bytes - Add route-level security tests and env validation tests
- Multi-stage Dockerfile (Node 24, arm64-ready, non-root): npm ci and prisma generate run at build time instead of on the server at deploy time; ships the RDS CA bundle so the DB certificate is verified - GET /health for load balancer checks, registered before logging and sessions so it never touches the database - Graceful shutdown on SIGTERM: stop accepting connections, drain in-flight requests, close the DB pool - Accept split DB_HOST/DB_NAME/DB_USER/DB_PASSWORD (injected from Secrets Manager on ECS) as an alternative to DATABASE_URL, for the app and the prisma CLI; verify DB TLS when DB_SSL_CA_PATH is set - Lint server.js too; drop its unused imports
Infrastructure (infra/): - Dedicated VPC over two AZs: ALB and tasks in public subnets (tasks accept traffic only from the ALB), RDS in private subnets with no internet route - RDS Postgres 17 restored from a snapshot of the old instance: private, gp3, forced TLS, RDS-managed rotating master password, deletion protection, final snapshot - One ALB with host routing (prod / staging), HTTP->HTTPS redirect, TLS 1.3 policy, single ACM cert for apex/www/staging (old staging cert expires 2026-09-25 and cannot renew) - ECS service module per environment: arm64 tasks, circuit-breaker rollback, ECS Exec, separate migrate task definition that alone can read the DB master secret - Secrets Manager: generated app secrets per env; integration credentials kept out of Terraform state - GitHub OIDC roles per GitHub environment, scoped to that environment's service, task families and roles - Apex/www DNS switch via prod_dns_target so cutover and rollback are one variable - S3 state backend with native locking (infra/bootstrap) Pipeline: - scripts/ecs-deploy.sh: migrate as a one-off task, roll the service, fail on circuit-breaker rollback - scripts/db-bootstrap.js: idempotent per-env database and least-privilege app role (data access only, no cross-environment connect) - CI on PRs (lint, tests, terraform validate, arm64 image build) and a deploy workflow: build once -> staging -> approval -> production - Remove the Elastic Beanstalk workflow so merging cannot redeploy EB - .gitattributes keeps LF endings for Dockerfile and Terraform
ECR scanning flagged 4 critical / 15 high CVEs in the bookworm base (perl, openssl, util-linux, zlib); bookworm is oldstable with LTS-only security support.
- Move each environment's ALB listener rule into the service module and make the ECS service depend on it: ECS rejects a target group that no listener uses yet, and the two were being created in parallel - State storage_encrypted = true on the restored RDS instance; left unset, Terraform planned to replace (destroy) the database - prevent_destroy on the database so any destructive plan fails outright - rds.force_ssl apply_method matches what RDS reports (no perpetual diff) - Put the RDS CA bundle in /etc/ssl/certs: ADD --chmod=644 also applied 644 to the /app/certs directory it created, so the non-root user could not open it and every task crashed at startup. Build now fails if the runtime user cannot read the bundle
Registration assigns roleid 4 and dashboards route on roleid 1-4, but the roles rows were never in a migration, so the new staging database rejected every registration with a foreign-key error. Names match production exactly; ON CONFLICT makes it a no-op there.
Prisma checksums migration.sql byte-for-byte; any autocrlf conversion would make applied migrations look modified.
The waiter can return immediately after update-service (the new deployment is not visible yet, so the old one looks stable) and also succeeds after a circuit-breaker rollback. Poll the deployment of the registered revision until COMPLETED, and fail on FAILED, INACTIVE (rolled back) or timeout.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Supersedes #36: this branch contains its security fixes.
Why
Elastic Beanstalk deploys had been failing since June. The single t3.micro ran
npm installon the server during every deploy and hung, a config change took the site down for about 30 minutes on 2026-09-23, and production was still running April's code (with a login bypass and unauthenticated data APIs).What's in here
App
/api/*route, Postgres-backed sessions, startup checks that refuse weak or missing secrets/health, graceful SIGTERM shutdown, verified DB TLS (RDS CA bundle), splitDB_*credentials from Secrets Managerrolesreference data (previously only inserted by hand, so a fresh DB rejected every registration)Infrastructure (
infra/, Terraform)prevent_destroyPipeline
ci.ymlon PRs: lint, tests,terraform validate, arm64 image builddeploy.ymlonmain: test → build once → staging (migrate + roll) → approval → production with the same imageAlready live and verified (applied by hand from this branch)
--connect-toprod_dns_target)What merging does
Runs the new Deploy workflow: staging deploys automatically, then production waits for approval. Public DNS does not change.
Test plan
terraform validate, shellcheck, actionlint