Skip to content

Add SSM Session Manager access to Stalwart nodes - #253

Merged
aatchison merged 2 commits into
mainfrom
aatchison/ssm-agent-pulumi-docs
Sep 30, 2026
Merged

aatchison merged 2 commits into
mainfrom
aatchison/ssm-agent-pulumi-docs

Conversation

@aatchison

@aatchison aatchison commented Sep 3, 2026 •

Copy link
Copy Markdown
Contributor

Verdict

Stage and prod management nodes are now reachable over AWS SSM Session Manager, no bastion or SSH key involved. Code-only change: deployment to existing running nodes is a separate manual step, and this PR does not run pulumi up.

Why

Port 8080 (the management dashboard) on the Stalwart management nodes (stage node 50, prod nodes 60/61) only accepts traffic from the private "management" load balancer's security group. The bastion isn't on that list, and this PR leaves that restriction alone on purpose.

Two methods now work without touching it, both by forwarding a node's own loopback port back to the operator instead of routing new traffic in:

  • SSH tunnel (existing, unchanged): two-hop ProxyCommand through the bastion, then -L 8080:localhost:8080 on the target node.
  • SSM port forwarding (new): aws ssm start-session --document-name AWS-StartPortForwardingSession --target <node>, run directly against the management node. This runs on the target itself and tunnels its own port back over the SSM agent's outbound WebSocket, so the node's inbound rules never get evaluated.

Verified live on stage (i-06b73cd0cf035ae0f): curl https://localhost:8080/ after starting the session returned the Stalwart Management login page.

AWS-StartPortForwardingSessionToRemoteHost (the bastion-relay variant) does not work here; it fails with Connection to destination port failed because that document proxies real, routed traffic through the jump host, which is still subject to the destination's security group. The README calls this out so nobody burns time trying it again mid-debug.

What changed

  • pulumi/stalwart/iam.py: attaches the AWS-managed AmazonSSMManagedInstanceCore policy to the Stalwart node IAM role.
  • pulumi/stalwart/__init__.py: threads the new RolePolicyAttachment through the constructor's return values, depends_on, self.finish() resources, and the class docstring.
  • pulumi/stalwart_instance_user_data.sh.j2: installs amazon-ssm-agent explicitly and enables/starts it.
  • README.md: new "Accessing the Management Dashboard via AWS SSO / SSM" section (prerequisites, the session-manager-plugin install including a no-root fallback, and the exact commands), plus a cross-reference from the SSH tunnel section noting they share the same trick.

Notes

The IAM role is shared cluster-wide, not per-node: stalwart_iam.iam() builds one role/instance profile reused by every node regardless of node_roles or services. There's no per-node role to scope this to just the management nodes, so it lands on the shared one. That grants SSM access cluster-wide rather than just to the 2-3 management nodes, but it's only AmazonSSMManagedInstanceCore (Session Manager plus agent registration), so the wider scope isn't a real risk.

These nodes run Amazon Linux 2023 (al2023-ami-minimal-kernel-6.1-x86_64), which usually ships the SSM agent preinstalled and enabled. We still install and enable it explicitly in user-data rather than relying on that, since it's not guaranteed across every AMI build and the install step is a no-op when the agent's already there.

User-data only runs on first launch, so this only affects future or replaced nodes. The existing running nodes (50, 60, 61) need the agent confirmed/enabled by hand; the IAM attachment applies live to the shared role with no restart needed.

No security group changes. pulumi up/preview was not run against real infra as part of this PR.

Attach AmazonSSMManagedInstanceCore to the shared Stalwart node IAM
role and explicitly install/enable amazon-ssm-agent in the node
user-data, so operators can reach the management dashboard (port 8080)
via AWS-StartPortForwardingSession with no bastion and no SG changes.

Document both the existing SSH tunnel and the new SSM method in
README.md, including why AWS-StartPortForwardingSessionToRemoteHost
does not work here (the management port's SG only allows the private
load balancer as a source).

Co-Authored-By: Claude Code <noreply@anthropic.com>
@aatchison
aatchison marked this pull request as draft September 3, 2026 00:03
Strip the inline/docstring comments added in the previous commit,
keeping the code logic unchanged. Also replace the OS-specific
session-manager-plugin install instructions in README.md with a link
to AWS's own install docs, since the exact steps vary by OS.

Co-Authored-By: Claude Code <noreply@anthropic.com>
@aatchison aatchison changed the title Add SSM Session Manager access to Stalwart management nodes Add SSM Session Manager access to Stalwart nodes Sep 3, 2026
@aatchison
aatchison marked this pull request as ready for review September 17, 2026 16:06
@aatchison
aatchison requested a review from ryanjjung September 17, 2026 16:06
@Sancus
Sancus self-requested a review September 17, 2026 17:59

@Sancus Sancus left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this is fine and much better than the bastion so lets do it ASAP.

@ryanjjung ryanjjung left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This looks okay. Seems we still need the bastion route available for us to SSH into the machines to run bootstrapping/restarts/etc. Is that an easy add here, or should we ticket that and the removal of the bastion configs?

@aatchison
aatchison merged commit 5ef0d0f into main Sep 30, 2026
1 check passed
@aatchison
aatchison deleted the aatchison/ssm-agent-pulumi-docs branch September 30, 2026 15:30
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants