Skip to content

chore(tests): add LMI e2e suite and run-scoped shared capacity provider - #5465

Open
svozza wants to merge 18 commits into
mainfrom
feat/lmi-e2e-tests
Open

chore(tests): add LMI e2e suite and run-scoped shared capacity provider#5465
svozza wants to merge 18 commits into
mainfrom
feat/lmi-e2e-tests

Conversation

@svozza

@svozza svozza commented Jul 13, 2026

Copy link
Copy Markdown
Contributor

Summary

Adds an end-to-end test that validates Logger's InvokeStore-backed log-attribute isolation on Lambda Managed Instances (LMI), where multiple invocations run concurrently in the same execution environment. Also introduces the testing infrastructure to support it, including a run-scoped shared capacity provider so LMI suites don't each provision their own VPC + capacity provider.

Changes

  • Add a Logger LMI e2e suite (packages/logger/tests/e2e/lmi.test.ts) that saturates a deliberately small fleet, forces concurrent invocations to multiplex into shared execution environments, and asserts each invocation's appended keys stay isolated
  • Capture the handler's own log lines by teeing process.stdout.write (the production write path) and returning them in the response payload, making log collection deterministic — no CloudWatch polling or ingestion latency
  • Filter captured lines by the InvokeStore-scoped invocationKey, and additionally assert on function_request_id as a regression check now that the Lambda context is scoped per invocation under LMI (fix(logger): scope lambda context per invocation under LMI concurrency #5430)
  • Add TestLmiCapacityProvider (dual-stack IPv6-only VPC, egress-only IGW, CloudWatch Logs interface endpoint — no NAT) and extend TestNodejsFunction to attach a function to a capacity provider by construct or by ARN
  • Add packages/testing/src/lmi/cli.ts with deploy/destroy commands that manage a run-scoped LmiShared-<runId>-<arch> stack and print the capacity provider ARN, so a workflow run can provision one shared capacity provider per architecture up front and pass it to each suite via LMI_CAPACITY_PROVIDER_ARN (suites fall back to an ephemeral per-suite provider when unset)
  • Run the LMI suite unconditionally on every e2e dispatch (removes the RUN_LMI_TESTS gate and workflow input)

Issue number: closes #5518


By submitting this pull request, I confirm that you can use, modify, copy, and redistribute this contribution, under the terms of your choice.

Disclaimer: We value your time and bandwidth. As such, any pull requests created on non-triaged issues might not be successful.

@powertools-for-aws-oss-automation powertools-for-aws-oss-automation Bot added the size/XL PRs between 500-999 LOC, often PRs that grown with feedback label Jul 13, 2026
Comment thread packages/testing/src/lmi/cli.ts Fixed
@powertools-for-aws-oss-automation powertools-for-aws-oss-automation Bot added size/XL PRs between 500-999 LOC, often PRs that grown with feedback and removed size/XL PRs between 500-999 LOC, often PRs that grown with feedback labels Jul 13, 2026
@sonarqubecloud

Copy link
Copy Markdown

@svozza
svozza requested a review from dreamorosi July 13, 2026 21:12
Comment thread packages/testing/src/lmi/cli.ts Outdated
svozza and others added 18 commits July 27, 2026 22:20
… Managed Instances

Adds an opt-in (RUN_LMI_TESTS) e2e suite that provisions an ephemeral
Lambda Managed Instances capacity provider (dual-stack IPv6 VPC, no NAT,
CloudWatch Logs interface endpoint) and proves the InvokeStore-backed
log attribute isolation across invocations multiplexed into the same
execution environment via a module-scoped promise barrier.

Testing utils changes:
- TestLmiCapacityProvider construct + ./resources/capacity-provider export
- ExtraTestProps.lmi option on TestNodejsFunction (memorySize>=2048 floor,
  $LATEST.PUBLISHED qualified output)
- includeTailLogs opt-out on invoke helpers (Tail unsupported on LMI)

Related to #5092, findings posted on the issue.
… CloudWatch polling

The handler tees process.stdout.write (the Logger's production write
path) and returns its own log lines in the response payload, making log
collection deterministic — no FilterLogEvents polling, no ingestion
latency, no CloudWatch client in the test.

Captured lines are filtered by the InvokeStore-scoped invocationKey
rather than function_request_id: addContext stores the Lambda context
in instance state, so under LMI multiplexing the request id stamped on
log lines can belong to a different invocation. Tracked separately as
a bug (see lmi-request-id-bug-handoff.md); once fixed, the filter
should flip back to function_request_id as a regression check.
…ogs by request id

With the lambda context now scoped per invocation under LMI (#5430),
the handler selects its own log lines by function_request_id — an
independent per-invocation attribute — and the test asserts the
InvokeStore-scoped invocationKey on those lines, making the suite a
regression check for both the context scoping and appendKeys isolation.

The suite adds ~4 minutes to the logger e2e cell, inside the agreed
budget for running on every e2e dispatch, so the RUN_LMI_TESTS gate
and workflow input are removed.
The e2e suites are I/O-bound (waiting on CloudFormation), but vitest
sizes its worker pool from CPU cores, so on 2-core CI runners only 3 of
the 6 logger e2e files ran concurrently and the rest queued — each then
paying its own stack deploy after waiting. With 8 workers all suites
deploy their stacks up front and the cell duration approaches the
slowest single suite (~6 min measured) instead of a serialized ~8 min.

CloudFormation read-API pressure from the extra concurrent stack
monitors is bounded by the existing DescribeStackEvents polling patch
in the testing package (10s interval per stack).
With all 36 matrix cells sharing one account, the extra concurrent
stack operations from 8-worker logger cells pushed the account-wide
CloudFormation API rate over the edge: three unrelated cells failed
with 'Throttling: Rate exceeded' on stack deploys. Back to default
worker sizing; cell-duration work moves to the run-scoped shared
capacity provider follow-up, which removes per-cell VPC/CP stacks
entirely instead of racing them.
Instead of every LMI suite provisioning its own capacity provider + VPC,
a workflow run can deploy one shared capacity provider per architecture up
front and pass its ARN to each suite via LMI_CAPACITY_PROVIDER_ARN.

- add packages/testing/src/lmi/cli.ts with deploy/destroy commands that
  manage a run-scoped LmiShared-<runId>-<arch> stack and print the
  capacity provider ARN on stdout
- TestNodejsFunction accepts a capacity provider ARN string and attaches
  the function via L1 CfnFunction.capacityProviderConfig (an imported
  capacity provider has no addFunction)
- ExtraTestProps.lmi.capacityProvider widened to CapacityProvider | string
- TestLmiCapacityProvider ctor loosened to Pick<TestStack, 'stack'> so the
  CLI can build the stack standalone
- logger lmi.test.ts reads LMI_CAPACITY_PROVIDER_ARN, falling back to an
  ephemeral per-suite capacity provider when unset
The --run-id argument flows into a CloudFormation stack name and, via
join(tmpdir(), ...), into the assembly output path that is later read
back. A crafted value (e.g. containing ../) could escape tmpdir(), so
restrict it to the alphanumerics-and-hyphens set CloudFormation already
requires for stack names and reject anything else at the boundary.
The path-injection scanner's taint analysis doesn't recognise the
run-id regex guard as sanitisation, and its guidance is to validate the
constructed path before touching the file system. Resolve the assembly
output directory and assert it stays within os.tmpdir() before it is
written to or read back, keeping the run-id validation as the input-side
guard for defence in depth.
The constructor's cognitive complexity hit 17 (limit 15) once it carried
the nested ARN-vs-construct capacity-provider branching. Move that logic
into two private methods so the constructor keeps a flat structure and
the two attachment strategies read independently. No behaviour change.

@dreamorosi dreamorosi left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These are findings from a deep review pass, grouped as 2 blockers plus several smaller fixes. Happy to discuss any of them.

* resolve the ARN from the stack outputs instead (see the e2e workflow).
*/
const main = async (): Promise<void> => {
await Promise.all(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocker: Running two TestStack.deploy() calls concurrently in one process also runs two fromAssemblyBuilder synths concurrently. toolkit-lib explicitly does not support this unless clobberEnv: false: each synth temporarily replaces the global process.env with an immutable proxy via temporarilyWriteEnv, and interleaved disposal can leave the process with the wrong/stale env (including leaked CDK_OUTDIR/CDK_CONTEXT_JSON). This works today only because each App is constructed before the synth window, so it is correct by accident. Could we either pass clobberEnv: false in TestStack.#synthAssembly() (safe here because the builder ignores the injected env), or deploy the architectures sequentially?

* teardown from being attempted.
*/
const main = async (): Promise<void> => {
const results = await Promise.allSettled(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The same concurrent-synth concern described in deploySharedCapacityProvider.ts applies here.

# "re-run all jobs": this teardown destroys the stacks at the end of each
# attempt, so a re-run LMI cell fails fast at ARN resolution ("stack does
# not exist") until the setup job re-runs and redeploys them.
teardown-lmi-capacity-providers:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocker: If an LMI cell is cancelled or times out before its afterAll runs, the per-suite function stack survives with its function still attached to the shared capacity provider. This teardown then tries to delete the provider stack while associations/ENIs still exist—the classic DELETE_FAILED path—leaking an EC2-backed fleet plus VPC with no sweeper to catch it. Could we add at least a retry/second pass in the destroy script, and/or scheduled cleanup for LmiShared-* stacks and stacks tagged Service: Powertools-for-AWS-e2e-tests?

# suites and are subject to account vCPU quotas, so the suites share one per
# architecture instead of provisioning their own (see
# packages/testing/src/lmi/sharedCapacityProviderStack.ts).
setup-lmi-capacity-providers:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Related to the teardown concern: none of the three new jobs sets timeout-minutes, so the 360-minute default lets a hung deploy/cell keep two EC2 fleets alive for up to six hours before teardown even starts. The suite hooks themselves cap at 20 minutes each, so a roughly 60-minute job timeout seems easy to justify.

arn=$(aws cloudformation describe-stacks \
--stack-name "LmiShared-${GITHUB_RUN_ID}-${ARCH//_/-}" \
--query "Stacks[0].Outputs[?OutputKey=='CapacityProviderArn'].OutputValue" \
--output text)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should-fix: --output text exits 0 with an empty string when the JMESPath filter matches nothing, and prints None when outputs are null. An empty value is silently treated as “no shared provider” by the test, so all four cells fall back to provisioning their own VPC and capacity provider, defeating the shared-provider design and pressuring the VPC quota without warning. Please fail unless the value is non-empty and not None, for example: [[ -n "$arn" && "$arn" != "None" ]] || exit 1.

minExecutionEnvironments,
maxExecutionEnvironments,
} = scaling;
capacityProvider.addFunction(this, {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should-fix: The ARN path explicitly sets publishToLatestPublished = true at line 131, while this construct path relies on CDK's @default - True even though the underlying CFN property documents no default. Passing publishToLatestPublished: true here would make both attach paths emit identical templates and remove dependence on an undocumented service default. This is also the path used by every local run and the CI fallback.

"devDependencies": {
"@aws-lambda-powertools/testing-utils": "file:../testing"
"@aws-lambda-powertools/testing-utils": "file:../testing",
"@aws-sdk/client-cloudwatch-logs": "^3.1079.0",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should-fix: This dependency is not imported anywhere in the package and looks left over from the CloudWatch-polling approach, so it can be dropped. Conversely, @aws-sdk/client-lambda is imported by tests/e2e/lmi.test.ts but is not declared here (it currently resolves via workspace hoisting), so that dependency should be declared instead.


describe('Logger E2E - Lambda Managed Instances', () => {
// The LMI scheduler scales out to fresh execution environments until the
// capacity provider's fleet is saturated (8 environments with a 12 vCPU

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should-fix (comment accuracy): The real CI run (29740823390) reported 30/30 responses; 12 multiplexed across 24 execution environments in all four cells, so the “fleet saturates at 8 environments” model here—and in the similar file-header comment—does not match observed behavior, likely because both runtime cells now share each per-architecture provider. The assertion only needs one overlap, so the test held, but could we describe the actual mechanics so future readers do not tune counts against the wrong model?

const invokeFunctionOnce = async ({
functionName,
payload = {},
includeTailLogs = true,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should-fix: This new option has no callers; the LMI suite hand-rolls its own LambdaClient/InvokeCommand instead. Could we either wire the suite through invokeFunction or drop the option? Also, with includeTailLogs: false, callers currently receive silently empty TestInvocationLogs rather than an error.

@dreamorosi dreamorosi left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Apologies for the long wait - I just finally found some time to review this now.

Some review comments that I think are valid, let me know if you disagree and/or I misunderstood anything.

@dreamorosi

Copy link
Copy Markdown
Contributor

The infrastructure-leak concern raised in review—cancelled cells leaving functions attached to the shared capacity provider and causing DELETE_FAILED during teardown—will be backstopped by the scheduled sweeper tracked in #5519. Within this PR, it would be good to keep the scope to job timeout-minutes and a retry/second pass in the destroy script. Before merge, a fresh full e2e workflow run against the current head SHA would also be good, since the last green run (29740823390) predates the newest commits and the merge from main.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/XL PRs between 500-999 LOC, often PRs that grown with feedback

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Maintenance: Add Logger E2E tests for Lambda Managed Instances

3 participants