Skip to content

Move to ECS Fargate: containerised app, Terraform, OIDC deploy pipeline - #37

Merged
Kiveshan merged 10 commits into
mainfrom
infra/ecs-migration
Sep 23, 2026
Merged

Kiveshan merged 10 commits into
mainfrom
infra/ecs-migration

Conversation

@Kiveshan

Copy link
Copy Markdown
Owner

Supersedes #36: this branch contains its security fixes.

Why

Elastic Beanstalk deploys had been failing since June. The single t3.micro ran npm install on the server during every deploy and hung, a config change took the site down for about 30 minutes on 2026-09-23, and production was still running April's code (with a login bypass and unauthenticated data APIs).

What's in here

App

  • The Fix login password bypass, unauthenticated data APIs, and session storage #36 security fixes: password check on login, auth on every /api/* route, Postgres-backed sessions, startup checks that refuse weak or missing secrets
  • Multi-stage Dockerfile (Node 24, Debian 13, arm64, non-root). Dependencies install and the Prisma client generates at build time, never on the server
  • /health, graceful SIGTERM shutdown, verified DB TLS (RDS CA bundle), split DB_* credentials from Secrets Manager
  • Migration seeding the roles reference data (previously only inserted by hand, so a fresh DB rejected every registration)

Infrastructure (infra/, Terraform)

  • New VPC: RDS in private subnets with no public endpoint; tasks accept traffic only from the ALB
  • RDS restored from a snapshot of the old instance: encrypted, RDS-managed rotating master password, prevent_destroy
  • One ALB with host routing for prod and staging, HTTP→HTTPS, TLS 1.3 policy, new ACM cert (the old staging cert expires 2026-09-25)
  • ECS services with circuit-breaker rollback and ECS Exec; a separate migrate task that alone can read the DB master secret
  • Least-privilege DB roles per environment (data access only, no cross-environment connect)
  • GitHub OIDC roles per environment; no long-lived AWS keys

Pipeline

  • ci.yml on PRs: lint, tests, terraform validate, arm64 image build
  • deploy.yml on main: test → build once → staging (migrate + roll) → approval → production with the same image
  • The deploy fails if ECS rolls back, rather than looking green
  • The Elastic Beanstalk workflow is removed

Already live and verified (applied by hand from this branch)

  • Staging on ECS at https://staging.bizexecdata.co.za: register, login (password now checked), 401 on anonymous API calls, HSTS/CSP, unknown hosts get a 404
  • Production on ECS behind the new ALB: 2 healthy tasks in 2 AZs, migrations applied; tested via --connect-to
  • Prisma migration history on the restored DB matches the repo's checksums
  • Public DNS still points at Elastic Beanstalk. Cutover is a separate one-variable change (prod_dns_target)

What merging does

Runs the new Deploy workflow: staging deploys automatically, then production waits for approval. Public DNS does not change.

Test plan

  • 104 unit/route tests, lint, terraform validate, shellcheck, actionlint
  • Local container run: migrations, login, session survives restart, graceful shutdown
  • Staging and production-on-ECS smoke tests (above)
  • CI green on this PR
  • After merge: Deploy workflow completes staging, then production after approval

…rage

- Verify the password with bcrypt in the local passport strategy; reject
  OAuth-created accounts that have no local password
- Require a session on every /api/* data endpoint (QuickBooks, Xero, Excel)
  and on the QuickBooks sync routes. With no session the userid filter was
  undefined, which Prisma drops, so these returned every user's rows
- Store sessions in Postgres via connect-pg-simple (new "session" table
  migration) instead of the in-process MemoryStore
- Fail fast at startup when DATABASE_URL, SESSION_SECRET (>=32 chars, not a
  placeholder) or ENCRYPTION_KEY (exactly 32 bytes) are missing or weak;
  remove the hardcoded session-secret fallbacks
- Fix .env.example ENCRYPTION_KEY placeholder, which was 34 bytes
- Add route-level security tests and env validation tests
- Multi-stage Dockerfile (Node 24, arm64-ready, non-root): npm ci and
  prisma generate run at build time instead of on the server at deploy
  time; ships the RDS CA bundle so the DB certificate is verified
- GET /health for load balancer checks, registered before logging and
  sessions so it never touches the database
- Graceful shutdown on SIGTERM: stop accepting connections, drain
  in-flight requests, close the DB pool
- Accept split DB_HOST/DB_NAME/DB_USER/DB_PASSWORD (injected from Secrets
  Manager on ECS) as an alternative to DATABASE_URL, for the app and the
  prisma CLI; verify DB TLS when DB_SSL_CA_PATH is set
- Lint server.js too; drop its unused imports
Infrastructure (infra/):
- Dedicated VPC over two AZs: ALB and tasks in public subnets (tasks accept
  traffic only from the ALB), RDS in private subnets with no internet route
- RDS Postgres 17 restored from a snapshot of the old instance: private,
  gp3, forced TLS, RDS-managed rotating master password, deletion
  protection, final snapshot
- One ALB with host routing (prod / staging), HTTP->HTTPS redirect, TLS 1.3
  policy, single ACM cert for apex/www/staging (old staging cert expires
  2026-09-25 and cannot renew)
- ECS service module per environment: arm64 tasks, circuit-breaker
  rollback, ECS Exec, separate migrate task definition that alone can read
  the DB master secret
- Secrets Manager: generated app secrets per env; integration credentials
  kept out of Terraform state
- GitHub OIDC roles per GitHub environment, scoped to that environment's
  service, task families and roles
- Apex/www DNS switch via prod_dns_target so cutover and rollback are one
  variable
- S3 state backend with native locking (infra/bootstrap)

Pipeline:
- scripts/ecs-deploy.sh: migrate as a one-off task, roll the service, fail
  on circuit-breaker rollback
- scripts/db-bootstrap.js: idempotent per-env database and least-privilege
  app role (data access only, no cross-environment connect)
- CI on PRs (lint, tests, terraform validate, arm64 image build) and a
  deploy workflow: build once -> staging -> approval -> production
- Remove the Elastic Beanstalk workflow so merging cannot redeploy EB
- .gitattributes keeps LF endings for Dockerfile and Terraform
ECR scanning flagged 4 critical / 15 high CVEs in the bookworm base
(perl, openssl, util-linux, zlib); bookworm is oldstable with LTS-only
security support.
- Move each environment's ALB listener rule into the service module and
  make the ECS service depend on it: ECS rejects a target group that no
  listener uses yet, and the two were being created in parallel
- State storage_encrypted = true on the restored RDS instance; left unset,
  Terraform planned to replace (destroy) the database
- prevent_destroy on the database so any destructive plan fails outright
- rds.force_ssl apply_method matches what RDS reports (no perpetual diff)
- Put the RDS CA bundle in /etc/ssl/certs: ADD --chmod=644 also applied
  644 to the /app/certs directory it created, so the non-root user could
  not open it and every task crashed at startup. Build now fails if the
  runtime user cannot read the bundle
Registration assigns roleid 4 and dashboards route on roleid 1-4, but the
roles rows were never in a migration, so the new staging database rejected
every registration with a foreign-key error. Names match production exactly;
ON CONFLICT makes it a no-op there.
Prisma checksums migration.sql byte-for-byte; any autocrlf conversion
would make applied migrations look modified.
The waiter can return immediately after update-service (the new deployment
is not visible yet, so the old one looks stable) and also succeeds after a
circuit-breaker rollback. Poll the deployment of the registered revision
until COMPLETED, and fail on FAILED, INACTIVE (rolled back) or timeout.
@Kiveshan
Kiveshan merged commit 004e555 into main Sep 23, 2026
3 checks passed
@Kiveshan
Kiveshan deleted the infra/ecs-migration branch September 25, 2026 06:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant