Skip to content

Fix outbound SMTP EHLO - IP mismatch #245

Description

@Sancus

The below is AI generated(Opus 5) but I edited it significantly myself for both length and language.

Summary

  1. Prod announces mail.thundermail.com in the SMTP EHLO on outbound delivery, but mail actually leaves from one of the four NAT gateway addresses currently attached out of the eight-address 5.60.96.8/29 pool, each of which has its own PTR in the smtp000–smtp007.thundermail.com scheme.
  2. mail.thundermail.com resolves to the inbound load balancer, not to any sending address, so the announced name does not correspond to the connecting IP.

Impact

  1. The EHLO name does not resolve to the sending IP. mail.thundermail.com resolves to 18.195.225.241 and 3.78.21.144 (the public load balancer). A receiver forward-resolving the EHLO name to verify it against the connecting address gets no match. This is the substantive issue, and it fires SpamAssassin's RCVD_HELO_IP_MISMATCH (~1.7–2.4 points).
  2. The EHLO string differs from the PTR string. A softer heuristic, but it is what mail testers report and it is visible to any receiver that compares the two.

TLDR;

Option A

We should probably collapse the reverse names into one, smtp.thundermail.com and set Stalwart to always use that for EHLO.

Option B

We could also make a 1:1 IP:node mapping to solve this, which would retain the numbered hostnames, but has other downsides.


Background and Detailed Explanation

Current state

Where the EHLO name comes from. server.hostname is set to mail.thundermail.com in the Postgres settings store (not in config.toml). No override exists at any level: queue.connection contains only default.description and default.timeout.connect, and queue.source-ip and queue.strategy are both empty. Stalwart's resolution order is per-source-IP override, then connection strategy ehlo-hostname, then server.hostname.

How mail egresses. All four sending nodes sit in private subnets with no public IP:

Node Private IP AZ Subnet
mailstrom-prod-stalwart-10 10.0.34.169 eu-central-1a subnet-054346d611223f189
mailstrom-prod-stalwart-11 10.0.126.243 eu-central-1b subnet-0b7564960babec972
mailstrom-prod-stalwart-12 10.0.57.148 eu-central-1a subnet-054346d611223f189
mailstrom-prod-stalwart-13 10.0.97.38 eu-central-1b subnet-0b7564960babec972

A single route table (mailstrom-prod-vpc-private-rt) covers both private subnets and sends all default traffic to one NAT gateway, nat-09d2add6f5ecd87e0.

The thundermail-public-pool-0 BYOIP block holds eight addresses, 5.60.96.8/29, each with its own PTR and a matching forward A record. Four are currently attached to the NAT gateway and the other four are unused.

IP Allocation PTR Forward A Attached
5.60.96.8 eipalloc-0fd0c117569440a4f smtp000-003.thundermail.com 5.60.96.8 NAT gateway
5.60.96.15 eipalloc-081748c1059a9affd smtp004-007.thundermail.com 5.60.96.15 unassociated

The SPF record already authorises the whole /29, so it anticipates all eight being used.

Why the obvious fix does not apply

Stalwart supports queue.source-ip.<ip>.ehlo-hostname, but it only consults those entries for addresses listed in queue.connection.<id>.source-ips, which it binds as a local source address. Our instances only ever hold 10.0.x.x on their NICs; the public address is applied by the NAT gateway. Stalwart therefore cannot key an EHLO name off the address the connection will actually use.

Any fix has to make the IP-to-name relationship deterministic.


Option A — One shared reverse identity across the pool

Collapse the reverse names into one, so the announced name is correct regardless of which address NAT selects.

  1. Create smtp.thundermail.com with A records covering the whole pool: 5.60.96.8 through .15. Do this first; AWS validates forward DNS before accepting a PTR change. Including the four currently unattached addresses means a spare can be attached later without silently reintroducing this bug.

  2. Repoint all eight PTRs, e.g. aws ec2 modify-address-attribute --allocation-id eipalloc-014cf4b0ce1d34629 --domain-name smtp.thundermail.com, repeated for the other seven allocations listed above.

  3. Set the outbound EHLO name, leaving server.hostname untouched so the inbound banner, MX, and client autoconfig are unaffected:

    [{"type":"insert","prefix":null,"assert_empty":false,
      "values":[["queue.connection.default.ehlo-hostname","smtp.thundermail.com"]]}]

    POST /api/settings on a management node, then GET /api/reload.

Keeps: the NAT gateway and every address in the pool. SPF unchanged. Each IP still has exactly one PTR. No config.toml change — the Stalwart side is a single settings-store key.

Gives up: per-IP reverse identity. Remote logs, blocklist lookups, and reputation dashboards will show one name for all eight addresses, so the IPs can no longer be distinguished or warmed independently. The smtp000–smtp007 naming scheme is retired.


Option B — An elastic IP per sending instance, egress via the internet gateway

Give each sending node its own public address so the egress IP is deterministic per node, then have each node announce the PTR name of its own address.

  1. Move the sending nodes to public subnets, or attach a second ENI in one, and route their egress through the internet gateway instead of the NAT gateway.
  2. Assign one pool address per node. The four currently unattached addresses (5.60.96.12–.15, smtp004–smtp007) already have correct PTR and forward records, so no new DNS work is needed for a four-node deployment, and the pool leaves headroom to scale to eight.
  3. Configure each node's EHLO name to match the PTR of its own address (see below).

Keeps: distinct per-IP PTRs and exact string equality between EHLO and rDNS. Clean 1:1 node-to-address mapping. NAT data-processing charges disappear for mail traffic.

Also preserves the existing smtp000–smtp007 naming scheme, which the pool is already set up for.

Gives up / risks: each mail node gains a public interface, so security group discipline matters more. Verify before committing that AWS's port 25 throttle removal applies to these addresses when attached directly to instances rather than to the NAT gateway. This is a meaningful change to the Pulumi network topology.

Option B requires a config.toml change

Because each node announces a different name and Stalwart's settings store is cluster-global, the EHLO value has to become node-local. config.local-keys accepts exact-match entries rather than only prefix.* globs, so it can be scoped to precisely one key:

Recommendation

Option A.

It resolves both problems above, requires no infrastructure change, and the only thing forfeited is per-IP distinguishability in logs.

The one downsides is that the way things were set up implies that per-IP identity was an intentional design choice. Option A discards it. If we want to keep it, then we'll have to do something like Option B.

In any case, we should test this on stage first.

Verification

After the change, from a receiving host:

  • dig -x <sending IP> returns the announced name.
  • Forward-resolving the announced name includes the sending IP.
  • A test message's Received header shows matching helo= and rdns= values.
  • Re-run the mail tester and confirm the HELO/rDNS finding is cleared.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Fields

    No fields configured for issues without a type.

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions