Skip to content

About

Terraform module for a distributed, multi-region load testing platform on AWS — Cognito/API Gateway/Lambda control plane, Step Functions + ECS Fargate/k6 execution, EventBridge scheduling and failure handling, IoT Core live metrics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Distributed Load Testing

Terraform module that provisions a serverless, distributed load testing engine on AWS. A single execution fans a k6 script out across many parallel Fargate tasks, then collects each task's results into S3.

This is a lightweight, Terraform-native take on the pattern behind AWS's Distributed Load Testing on AWS solution. It now covers every piece of that reference architecture that's pure infrastructure: an authenticated API (Cognito + API Gateway + Lambda) for managing reusable test scenarios, starting/monitoring runs, and scheduling recurring or future runs via EventBridge Scheduler; multi-region test execution across a fixed set of regions; EventBridge-driven failure handling; and live metrics streaming to IoT Core. What's deliberately not built is the web console's actual frontend application — CloudFront + S3 hosting infra exists (console.tf) but serves a placeholder, since building a React/Amplify app is out of scope for a Terraform module — plus the optional Bedrock AgentCore MCP server, which the upstream solution itself treats as optional (see the reference architecture below for full context).

Reference architecture (upstream AWS solution)

This module is modeled on the Distributed Load Testing on AWS solution. Its full reference architecture, as published, is considerably larger than what this repo deploys:

flowchart TB
    subgraph Console["Console hosting (choose one)"]
        CF["CloudFront + S3\n(Amplify console)"]
        ALB["ALB + ECS Fargate\n(Amplify console)"]
        HL["Headless\n(self-hosted console ZIP)"]
    end

    Operator(["Console user / CLI"])
    Cognito["Cognito user pool\n(console, REST API, CLI, MCP auth)"]
    APIGW["API Gateway"]
    Lambda["Lambda microservices\n(test CRUD, scheduling)"]
    EB["EventBridge\n(scheduler + failure rules)"]
    FailFn["Failure handler Lambda"]
    SFN["Step Functions\n(test orchestration)"]
    S3["S3\n(scenarios, per-region results)"]
    DDB["DynamoDB\n(test/results metadata)"]
    ECS["ECS Fargate tasks\n(Taurus: JMeter / k6 / Locust)\nper selected Region"]
    CW["CloudWatch Logs"]
    LiveFn["Live-data Lambda"]
    IoT["AWS IoT Core\n(live metrics topic)"]
    MCPClient["MCP client\n(AI dev tool)"]
    AgentCore["Bedrock AgentCore Gateway"]
    MCPFn["DLT MCP Server Lambda"]

    Operator --> Console
    Console --> Cognito
    Operator -- "CLI" --> Cognito
    Cognito --> APIGW
    APIGW --> Lambda
    Lambda --> S3
    Lambda --> DDB
    Lambda --> EB
    EB -- "scheduled test" --> Lambda
    Lambda --> SFN
    EB -- "ECS/SFN failure events" --> FailFn
    SFN --> ECS
    ECS --> S3
    ECS --> CW
    S3 --> Lambda
    Lambda -- "aggregate results" --> DDB
    CW -.->|"live data option"| LiveFn
    LiveFn --> IoT
    IoT -.->|subscribe| Console
    MCPClient --> AgentCore
    AgentCore -- "validates Cognito token" --> MCPFn
    MCPFn --> DDB
    MCPFn --> S3
    MCPFn --> CW
Loading

Key pieces of that reference architecture: Cognito-gated console (CloudFront+S3, ALB+ECS Fargate, or headless), API Gateway/Lambda microservices for test CRUD and EventBridge-based scheduling, Step Functions orchestrating multi-region ECS Fargate tasks running Taurus (JMeter/k6/Locust), result aggregation in DynamoDB, an optional IoT Core live-metrics path, and an optional Bedrock AgentCore Gateway MCP server for AI-assisted analysis.

This repo now builds every piece of that diagram except: the console's frontend application (hosting infra exists, no app code), the ALB+ECS/headless console alternatives (CloudFront+S3 only), and the optional MCP server. Everything else — scheduling and ECS/Step-Functions failure-event routing (both bundled in the diagram's EB node), multi-region Fargate fan-out, and IoT Core live streaming — is implemented, described in This module's architecture below.

This module's architecture

Everything below is implemented, single primary region for the API/control-plane, fanning out to a fixed set of execution regions:

flowchart TB
    User[/"Operator\n(CLI / SDK)"/]
    Cognito["Cognito user pool\n(admin-created users only)"]
    APIGW["API Gateway REST API\n(Cognito authorizer)"]
    LamScenarios["scenarios Lambda\n/scenarios*"]
    LamTests["tests Lambda\n/tests*"]
    Scheduler["EventBridge Scheduler"]
    DDBScenarios[("DynamoDB\ntest_scenarios")]
    S3[("S3 bucket\nscripts/*, results/*")]
    Dispatcher["Dispatcher state machine\n(primary region)"]
    DDB[("DynamoDB\ntest_runs")]
    LamDispatch["region_dispatch Lambda"]
    EBRule["EventBridge rule\nexecution FAILED/TIMED_OUT/ABORTED"]
    LamFailure["failure_handler Lambda"]
    Console["CloudFront + S3\n(console, infra only)"]

    subgraph Region["Per region — eu-west-1 / eu-central-1 / ca-central-1"]
        RegionalSFN["Regional state machine"]
        ECS["ECS Fargate tasks\n(k6, N parallel)"]
        RegionalLogs["CloudWatch log group"]
        LamLive["live_data Lambda"]
        IoT["IoT Core\ntopic loadtest/*"]
    end

    User -- "sign in" --> Cognito
    User -- "Authorization: IdToken" --> APIGW
    APIGW -- "validates token" --> Cognito
    APIGW -- "/scenarios*" --> LamScenarios
    APIGW -- "/tests*" --> LamTests
    LamScenarios -- "CRUD" --> DDBScenarios
    LamScenarios -- "StartExecution" --> Dispatcher
    LamScenarios -- "Create/List/Delete Schedule" --> Scheduler
    Scheduler -- "invoke" --> LamScenarios
    LamTests -- "read status" --> DDB
    LamTests -- "list/read results" --> S3

    Dispatcher -- "RecordStart / RecordComplete" --> DDB
    Dispatcher -- "Map: one region-dispatch call per region" --> LamDispatch
    LamDispatch -- "start/check\n(cross-region boto3)" --> RegionalSFN
    RegionalSFN -- "runTask.sync" --> ECS
    ECS -- "fetch-script / upload-results" --> S3
    ECS -- "task logs" --> RegionalLogs
    RegionalLogs -- "subscription filter" --> LamLive
    LamLive -- "publish" --> IoT

    Dispatcher -. "execution fails" .-> EBRule
    EBRule --> LamFailure
    LamFailure -- "mark FAILED" --> DDB

    User -. "console (placeholder only)" .-> Console
Loading

Execution flow (via the API)

  1. Bootstrap an operator account once (see Authenticated API below), then sign in to get a bearer token.
  2. Upload a k6 test script to the results bucket under scripts/<key> (still done directly against S3 — no upload endpoint).
  3. Create a scenario: POST /scenarios with name, scriptKey, defaultTaskCount, optional regions (default: the primary region only). Returns a scenarioId.
  4. Start a test: POST /scenarios/{id}/start. The scenarios Lambda reads the scenario, resolves each requested region's state machine ARN (from REGIONAL_STATE_MACHINES), generates a testId, and calls states:StartExecution on the dispatcher with:
    {
      "testId": "<scenarioId>-<epoch>",
      "scenarioId": "<scenarioId>",
      "scriptKey": "checkout.js",
      "taskCount": 5,
      "regionSpecs": [
        { "region": "eu-west-1", "stateMachineArn": "...", "taskIndexes": [0, 1, 2, 3, 4] }
      ]
    }
    taskCount applies per region, not split across them — requesting 2 regions at count 5 runs 10 Fargate tasks total, 5 per region.
  5. The dispatcher writes a RUNNING row to test_runs (RecordStart), then fans a Map state out across regionSpecs (up to max_concurrent_regions regions in parallel). Each region's iteration invokes the region_dispatch Lambda to start_execution on that region's own state machine, then polls it every 15s via the same Lambda (describe_execution) until it's no longer RUNNING — this Lambda-mediated start/poll loop is the only correct way to reach another region, since Step Functions' native service integrations (ecs:runTask.sync, etc.) always resolve against the state machine's own home region and can't be redirected elsewhere.
  6. Each region's own state machine does exactly what the single-region version always did: fans a Map out across taskCount Fargate tasks (up to max_concurrent_tasks concurrently, ecs:runTask.sync, each task tagged with TestId/TaskIndex for the live-metrics correlation below). Each Fargate task is three containers sharing an ephemeral volume: fetch-script downloads scripts/<scriptKey> to /shared/script.js, k6 runs it and writes /shared/summary.json, upload-results uploads that to results/<testId>/<taskIndex>.json. (The k6 image ships without the AWS CLI, hence the sidecars.) All regions read/write the same global S3 bucket — a deliberate simplification over the upstream solution's per-region result buckets, accepting cross-region S3 transfer for a single, simpler store.
  7. Once every region reports non-RUNNING, the dispatcher marks the test COMPLETE (RecordComplete), recording per-region outcomes.
  8. Check status: GET /tests/{testId} returns the DynamoDB status row plus the list of result object keys under results/<testId>/ (not the inlined k6 summaries, to stay under API Gateway's payload limit on high-task-count tests).

The dispatcher can still be invoked directly via aws stepfunctions start-execution (see Usage) — the API is a convenience/auth layer in front of it, not a replacement.

Failure handling

An aws_cloudwatch_event_rule (eventbridge.tf) matches the dispatcher's execution reaching FAILED/TIMED_OUT/ABORTED and routes it to the failure_handler Lambda, which recovers testId from the execution's original input and conditionally marks the test_runs row failed (ConditionExpression guards against double-delivery and against overwriting a row not still RUNNING) — closing a real gap where an uncaught failure anywhere in the dispatcher or a regional state machine would otherwise leave a test stuck at RUNNING forever. Only one rule is needed: RunLoadTasks has no Catch/Retry, so any ECS task failure already surfaces as a dispatcher execution failure — a separate ECS-level rule would double-handle the same event.

Live metrics streaming

Each region's CloudWatch log group for its runner tasks has a subscription filter forwarding every log line to that region's own live_data Lambda, which resolves ecsTaskId (parsed from the log stream name) back to (testId, taskIndex) via ecs:DescribeTasks (using the TestId/TaskIndex tags set on the task at RunTask time — not container-log parsing, which would be fragile across three separately-logging containers), then publishes to IoT Core topic loadtest/<testId>/<taskIndex>/logs. This is publish-side only — there's no console yet to subscribe, and the subscriber-side IoT provisioning (a Cognito Identity Pool, IAM role, and IoT policy for Connect/Subscribe/Receive) depends on decisions the eventual console build will make, so it's intentionally not built here.

Console hosting

console.tf provisions CloudFront + a private S3 bucket (via Origin Access Control) wired to serve a console — but only a placeholder index.html ("Console not yet deployed"), not a real frontend application. The three values a future console build would need are already available as outputs: api_base_url, cognito_user_pool_id, cognito_app_client_id. No runtime-config injection mechanism exists yet (no build pipeline to regenerate it), and no custom domain/ACM certificate — just the default *.cloudfront.net domain.

Authenticated API

Cognito user pool (cognito.tf) is admin-create-users only — there's no public sign-up, matching the upstream solution's default-admin pattern. Terraform does not create any human user account; bootstrap the first operator manually after apply:

aws cognito-idp admin-create-user \
  --user-pool-id "$(terraform output -raw cognito_user_pool_id)" \
  --username <admin-email> \
  --user-attributes Name=email,Value=<admin-email> Name=email_verified,Value=true \
  --desired-delivery-mediums EMAIL

This emails a temporary password. Sign in once to set a real one, then get a bearer token for the API:

aws cognito-idp admin-initiate-auth \
  --user-pool-id "$(terraform output -raw cognito_user_pool_id)" \
  --client-id "$(terraform output -raw cognito_app_client_id)" \
  --auth-flow ADMIN_USER_PASSWORD_AUTH \
  --auth-parameters USERNAME=<admin-email>,PASSWORD=<password>

Use the resulting IdToken as the Authorization header value on every request below (API Gateway's Cognito authorizer validates it directly — no client secret is used).

Method Path Description
POST /scenarios Create a scenario: {"name", "description", "scriptKey", "defaultTaskCount", "regions"?} (regions: list of region strings, e.g. ["eu-west-1","ca-central-1"]; defaults to the primary region if omitted)
GET /scenarios List scenarios
GET /scenarios/{id} Get one scenario
PUT /scenarios/{id} Update a scenario
DELETE /scenarios/{id} Delete a scenario
POST /scenarios/{id}/start Start an execution; optional {"taskCount", "regions"} override the scenario defaults (taskCount applies per region, not split across them). Returns {"testId"}
POST /scenarios/{id}/schedules Create a schedule: {"scheduleExpression", "taskCount"?, "regions"?, "timezone"? (default "UTC"), "state"? (default "ENABLED")}. Returns {"scheduleId", ...}
GET /scenarios/{id}/schedules List schedules for this scenario
GET /scenarios/{id}/schedules/{scheduleId} Get one schedule's detail
DELETE /scenarios/{id}/schedules/{scheduleId} Delete a schedule
GET /tests List test runs (from test_runs)
GET /tests/{testId} Get status + result object keys for one run
curl -H "Authorization: <IdToken>" -X POST "$(terraform output -raw api_base_url)/scenarios" \
  -d '{"name":"checkout load","scriptKey":"checkout.js","defaultTaskCount":5}'

curl -H "Authorization: <IdToken>" -X POST "$(terraform output -raw api_base_url)/scenarios/<scenarioId>/start"

# Multi-region: taskCount applies per region (2 regions x 5 = 10 tasks total)
curl -H "Authorization: <IdToken>" -X POST "$(terraform output -raw api_base_url)/scenarios/<scenarioId>/start" \
  -d '{"regions":["eu-west-1","ca-central-1"]}'

curl -H "Authorization: <IdToken>" "$(terraform output -raw api_base_url)/tests/<testId>"

The test_scenarios table (dynamodb.tf) is CRUD-managed only by the scenarios Lambda; the existing test_runs table is untouched — it's still written only by the state machine.

Scheduling

Recurring or future-dated test runs use EventBridge Scheduler, created dynamically by the scenarios Lambda via scheduler:CreateSchedule — there's no Terraform-managed schedule resource for individual schedules, only the schedule group they live in (scheduler.tf). EventBridge Scheduler is the source of truth (queried live via ListSchedules/GetSchedule); nothing is mirrored into DynamoDB, so there's no risk of drift if a schedule is ever changed outside the API. A schedule's scheduleId is the literal EventBridge Scheduler schedule name (<scenarioId>--<uuid>), which is also what makes per-scenario listing work via a NamePrefix filter.

scheduleExpression uses EventBridge Scheduler's own syntax directly — at(2026-09-01T09:00:00) for a one-time run, rate(1 day) or cron(0 9 * * ? *) for recurring ones:

curl -H "Authorization: <IdToken>" -X POST "$(terraform output -raw api_base_url)/scenarios/<scenarioId>/schedules" \
  -d '{"scheduleExpression":"rate(1 day)","taskCount":5}'

curl -H "Authorization: <IdToken>" "$(terraform output -raw api_base_url)/scenarios/<scenarioId>/schedules"

curl -H "Authorization: <IdToken>" -X DELETE "$(terraform output -raw api_base_url)/scenarios/<scenarioId>/schedules/<scheduleId>"

When a schedule fires, EventBridge Scheduler invokes the scenarios Lambda directly (not through API Gateway) with a bare {"scenarioId", "taskCount"?} payload — the same lambda_handler detects this (no httpMethod key in the event) and calls the same internal start_test() function the HTTP POST /scenarios/{id}/start endpoint uses, so a scheduled run behaves identically to a manually-started one. Unlike the HTTP path, errors in this branch are not caught and turned into an error response — they propagate so EventBridge Scheduler's own retry policy engages instead of silently swallowing a failed scheduled run.

Provisioning order

Terraform infers creation order from resource attribute references, so no manual sequencing is normally required. Applying this module creates resources in roughly this order:

  1. S3 bucket, DynamoDB tables, and the IAM assume-role documents (s3.tf, dynamodb.tf, iam.tf) — independent, created in parallel.
  2. Global IAM roles (iam.tf): ecs_task_execution/ecs_task (depend on the S3 bucket ARN), state_machine_regional (depends on a hand-built list of regional ECS cluster ARNs — deliberately not read from the modules' outputs, to avoid a dependency cycle; see the comment above local.regional_cluster_arns in iam.tf), state_machine_dispatcher (depends on the test_runs table ARN and the region-dispatch Lambda ARN), lambda_region_dispatch (depends on hand-built regional state-machine/execution ARNs), lambda_live_data (depends on the same hand-built regional cluster ARN list), lambda_scenarios/lambda_tests (as before), scheduler_invoke, lambda_failure_handler.
  3. region_dispatch and failure_handler Lambda functions (lambda.tf) — depend on their own role's policy/attachment and log group.
  4. The dispatcher state machine (step_functions.tf) — depends on state_machine_dispatcher's role/policy, the test_runs table, and the region_dispatch Lambda ARN.
  5. The EventBridge failure rule (eventbridge.tf) — depends on the dispatcher's ARN and the failure_handler Lambda.
  6. Per region (eu-west-1, eu-central-1, ca-central-1, via regional_runners.tf's three module blocks, each on its own aliased provider from aws.tf): VPC/subnet lookup, ECS cluster/log group/security group/task definition, the region-local state machine, and that region's live_data Lambda + CloudWatch Logs subscription filter — see modules/regional_runner. Each module block carries an explicit depends_on on state_machine_regional's and lambda_live_data's policies, since it only receives their role ARNs as input variables (see the depends_on note below).
  7. Cognito user pool + app client (cognito.tf), test_scenarios table, and console hosting (S3 + CloudFront + bucket policy, console.tf) — all independent of the above, created in parallel.
  8. API Gateway (Cognito authorizer, /scenarios*//tests* resources/methods/integrations, Lambda permissions, deployment + stage) — depends on the user pool ARN and the scenarios/tests Lambda invoke ARNs.
  9. aws_scheduler_schedule_group (scheduler.tf) — independent.

One gap Terraform's implicit graph does not close on its own: several resources reference IAM roles, not the inline policies attached to those roles, so nothing forces the policies to exist first. That's harmless for terraform apply itself (ECS/Step Functions/API Gateway/EventBridge Scheduler don't validate permissions at registration time), but if a load test or API call is made immediately after apply, IAM's eventual consistency can cause an AccessDenied/403 on the very first attempt. To close that window, these resources carry explicit depends_on:

  • aws_ecs_task_definition.regional_runner (inside the module) → its region's execution/task role policies (unchanged from the single-region version, just relocated)
  • aws_sfn_state_machine.load_test (dispatcher) → aws_iam_role_policy.state_machine_dispatcher
  • aws_lambda_function.scenarios / .tests / .failure_handler / .region_dispatch → their own role's inline policy + managed attachment
  • Each module.runner_* block → aws_iam_role_policy.state_machine_regional, aws_iam_role_policy.lambda_live_data, aws_iam_role_policy_attachment.lambda_live_data_basic
  • aws_api_gateway_deployment.load_testing → both aws_lambda_permission resources (in addition to the resource/method/integration IDs already in its triggers hash)
  • aws_cloudwatch_log_subscription_filter.live_data (inside the module) → its region's aws_lambda_permission (a subscription filter attached before the Lambda's resource policy exists can fail creation outright, not just 403 on first use)

This orders policy creation before the resources that use it, which mitigates but cannot fully eliminate IAM propagation delay. If you hit an AccessDenied/403 immediately after a fresh apply, retry the request.

Two exceptions, deliberately without this treatment: aws_iam_role.scheduler_invoke (individual schedules are created at Lambda runtime via boto3, long after any propagation window, and there's no Terraform-managed schedule resource to hang a depends_on off) and the cross-region dispatch path (EventBridge Scheduler and, separately, EventBridge Scheduler-invoked regional executions both have built-in retry behavior that self-heals a first-invocation race, unlike the synchronous ECS/Step-Functions/API-Gateway paths above).

This was a real cutover, not purely additive

Introducing multi-region replaced the single-region ECS cluster, task definition, and state machine with per-region equivalents inside modules/regional_runner. ecs.tf and vpc.tf were deleted; nothing in DynamoDB or S3 was touched. If you'd already applied the single-region version of this module, this change destroys and recreates the ECS/state-machine layer — plan carefully around any in-flight test executions.

Resources created

File Resources
s3.tf Encrypted, private results bucket with a 90-day lifecycle expiry on results/
dynamodb.tf test_runs table (pay-per-request, PITR, encrypted) keyed on testId; test_scenarios table keyed on scenarioId
modules/regional_runner Per-region (instantiated 3x from regional_runners.tf): VPC/subnet lookup, ECS cluster + Fargate task definition (fetch-script / k6 / upload-results) + log group + security group (egress only), a region-local state machine, and a live_data Lambda + CloudWatch Logs subscription filter publishing to IoT Core
regional_runners.tf The three module instantiations (one per hand-coded region) + locals.regional_state_machines/regional_ecs_clusters maps consumed by the scenarios Lambda and outputs
step_functions.tf The dispatcher state machine (primary region): records run state, fans out across regions via the region_dispatch Lambda, waits on each region's completion, then records the final result
eventbridge.tf Rule matching dispatcher execution FAILED/TIMED_OUT/ABORTED, routed to failure_handler
cognito.tf Admin-only Cognito user pool + app client (no client secret) fronting the API
console.tf CloudFront + private S3 (via Origin Access Control) hosting infra for a future console, currently serving a placeholder page; bucket policy granting CloudFront read access lives here rather than iam.tf since it's a resource-based policy on the bucket, not a role
api_gateway.tf REST API, Cognito authorizer, /scenarios* (including /scenarios/{id}/schedules*) and /tests* resources/methods/integrations, Lambda permissions, deployment + stage
lambda.tf archive_file-packaged scenarios, tests, failure_handler, and region_dispatch Lambda functions (Python 3.13) + their log groups. (live_data is packaged per-region inside the module, not here.)
lambda/scenarios/handler.py, lambda/tests/handler.py, lambda/failure_handler/handler.py, lambda/region_dispatch/handler.py, lambda/live_data/handler.py Lambda source
scheduler.tf aws_scheduler_schedule_group that holds the dynamically-created per-scenario schedules
iam.tf ecs_task_execution/ecs_task roles; state_machine_regional role (ECS actions scoped to the hand-built list of regional cluster ARNs); state_machine_dispatcher role (DynamoDB + invoke region_dispatch); lambda_region_dispatch role (states:StartExecution/DescribeExecution across regional ARNs — the one legitimate cross-region reach, since it's plain boto3 code); lambda_scenarios role (scenarios/schedule CRUD, states:StartExecution on the dispatcher); lambda_tests role; lambda_failure_handler role (dynamodb:UpdateItem on test_runs); lambda_live_data role (region-wildcarded iot:Publish, ecs:DescribeTasks scoped to the regional cluster list); scheduler_invoke role
aws.tf Provider/version constraints (aws, archive) + 3 aliased aws providers, one per execution region
variables.tf Module inputs (see below)
outputs.tf Module outputs (see below)

Inputs

Variable Description Default
aws_region AWS region for the API/control-plane (Cognito, API Gateway, dispatcher, DynamoDB) eu-west-1
name_prefix Prefix applied to all resource names dlt
vpc_id VPC override for the primary region's default-VPC lookup pattern (each execution region resolves its own default VPC independently — see regional_network_overrides for per-region overrides) null
subnet_ids Subnets override, same scope as vpc_id null
assign_public_ip Assign runners a public IP (needed if subnets have no NAT/internet route) — applies to every region true
k6_image Container image used to run load-test scripts public.ecr.aws/docker/grafana/k6:latest
task_cpu Fargate task CPU units per runner 1024
task_memory Fargate task memory (MiB) per runner 2048
max_concurrent_tasks Upper bound on parallel runners per test, per region 20
max_concurrent_regions Upper bound on how many regions a single test fans out to in parallel 5
regional_network_overrides Per-region VPC/subnet overrides, keyed by region (e.g. {"ca-central-1" = {vpc_id = "vpc-...", subnet_ids = [...]}}). Omit a region to use its own default VPC {}
log_retention_days CloudWatch Logs retention for task/Lambda output 30
tags Tags applied to all resources {}
cognito_password_min_length Minimum password length enforced by the admin Cognito user pool 12
api_stage_name API Gateway deployment stage name v1
lambda_timeout Timeout (seconds) for the API/dispatch/live-data Lambdas 10
lambda_memory_size Memory (MiB) for the API/dispatch/live-data Lambdas 256

Note: the actual region list (eu-west-1, eu-central-1, ca-central-1) is not a variable — it's a fixed, hand-coded set of provider aliases in aws.tf plus module blocks in regional_runners.tf, and a matching local.regional_names list in iam.tf. Terraform can't generate provider configurations dynamically (no for_each/count on provider blocks), so adding, removing, or changing a region means editing those three places, not changing a variable's value.

Outputs

Output Description
state_machine_arn Start executions here to launch a load test (the dispatcher)
results_bucket Upload scripts under scripts/, read results under results/<testId>/ — shared across all regions
test_runs_table DynamoDB table tracking test run status
regional_state_machine_arns Per-region runner state machine ARNs, keyed by region
regional_ecs_clusters Per-region ECS cluster names, keyed by region
api_base_url Base URL for the scenarios/tests REST API
cognito_user_pool_id Cognito user pool backing the API; create operator accounts with aws cognito-idp admin-create-user
cognito_app_client_id Cognito app client ID used to authenticate against the API
test_scenarios_table DynamoDB table holding reusable test scenario definitions
console_url CloudFront URL serving the console bucket (infra only — no frontend app deployed)

Usage

terraform init
terraform plan
terraform apply

Upload a script (once, to the shared bucket) and start a test directly against the dispatcher:

aws s3 cp checkout.js "s3://$(terraform output -raw results_bucket)/scripts/checkout.js"

aws stepfunctions start-execution \
  --state-machine-arn "$(terraform output -raw state_machine_arn)" \
  --input '{
    "testId": "2026-08-19-checkout-load",
    "scriptKey": "checkout.js",
    "taskCount": 5,
    "regionSpecs": [
      {"region": "eu-west-1", "stateMachineArn": "<from regional_state_machine_arns>", "taskIndexes": [0,1,2,3,4]}
    ]
  }'

In practice, use the API instead (Authenticated API) — it resolves regionSpecs for you from a scenario's configured regions.

Check status and fetch results:

aws dynamodb get-item \
  --table-name "$(terraform output -raw test_runs_table)" \
  --key '{"testId":{"S":"2026-08-19-checkout-load"}}'

aws s3 sync "s3://$(terraform output -raw results_bucket)/results/2026-08-19-checkout-load/" ./results/

Notes

  • No web console application is provisioned — CloudFront + S3 hosting infra exists (console.tf) but serves a placeholder; the API is the real interface today, alongside direct CLI/SDK access to the underlying resources.
  • force_destroy = true on both S3 buckets (results and console) means terraform destroy deletes their contents — copy anything you need out first. The Cognito user pool is deliberately not force-destroyable (deletion_protection = "ACTIVE"): it holds every operator's identity, so an accidental replace would lock out all admins at once, unlike disposable script/result/console-asset data.
  • Fargate runners have egress-only security group rules and no inbound listeners, in every region.
  • Cognito is admin-create-users only; there is no public sign-up and Terraform does not manage any human user account (see Authenticated API).
  • The results bucket is single-region and shared by every execution region — Fargate tasks in eu-central-1/ca-central-1 read/write it over the network. This is a deliberate simplification versus the upstream solution's per-region result buckets; revisit if cross-region S3 transfer cost/latency becomes a real concern.
  • The region list itself (which regions exist, not which a given test targets) is fixed at Terraform-authoring time — see the note at the end of Inputs.
  • Review IAM scoping in iam.tf and any org-specific compliance requirements before running load tests against production endpoints.

About

Terraform module for a distributed, multi-region load testing platform on AWS — Cognito/API Gateway/Lambda control plane, Step Functions + ECS Fargate/k6 execution, EventBridge scheduling and failure handling, IoT Core live metrics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages