Email 11 - Relay servers
Email Solution · Previous: High availability and key synchronization · Next: Mailbox server
Everything described so far happens inside the management network. Two small servers per site are the only part of the email system that talks to the internet: the relay servers. They take what the internal pair hands them, remove the headers that describe the inside, and deliver it. They hold no mailboxes, no directory integration and no state besides the queue, which is why this Article is short and why most of it is about names, routes and what the servers refuse to do.
One job
The design gives the relay one function: accept mail from the internal email servers and send it to the right server on the internet. The install notes add a second one in a single line of the transport table: a message for the site's own domain goes back inside, to the mailbox server.
| Host | Role | Internal VLAN 1168, eth0 | Relay VLAN 1172, eth1 | Public address after NAT |
|---|---|---|---|---|
DC2-A-VCMSR001 | primary relay, datacenter A | 10.12.19.34 | 10.12.19.65 | 203.0.113.44 |
DC2-B-VCMSR001 | fallback relay, datacenter B | 10.12.19.35 | 10.12.19.66 | 203.0.113.48 |
Site 1 has the same pair in the design: DC1-A-VCMSR001 and DC1-B-VCMSR001, relay addresses 10.11.19.65 and 10.11.19.66, public addresses 198.51.100.44 and 198.51.100.48. I have the configuration trees of site 2 only, so everything below is site 2.
flowchart TB x["Internal pair DC2-A-VCMSX001 and DC2-B-VCMSX001"] subgraph ra["DC2-A-VCMSR001, primary"] ra0["eth0 10.12.19.34"] rap["Postfix, smtp_header_checks"] ra1["eth1 10.12.19.65, default gateway"] ra0 --> rap --> ra1 end subgraph rb["DC2-B-VCMSR001, fallback"] rb0["eth0 10.12.19.35"] rbp["Postfix, smtp_header_checks"] rb1["eth1 10.12.19.66, default gateway"] rb0 --> rbp --> rb1 end x == "relayhost, SMTP 25" ==> ra0 x -. "smtp_fallback_relay, SMTP 25" .-> rb0 ra1 --> fw["Network firewall, NAT"] rb1 --> fw fw -- "from 203.0.113.44 or 203.0.113.48, SMTP 25, STARTTLS if offered" --> mx["Mail server on the internet"] rap -. "ad-dc2.example.net only, back through eth0" .-> x3["DC2-A-VCMSX002 10.12.19.43"]
Two interfaces and the routes between them
A relay stands in two VLANs. eth0 is in the internal VLAN 1168 (10.12.19.32/28), together with the internal servers. eth1 is in the relay VLAN 1172 (10.12.19.64/28), whose only purpose is the path to the internet, and it carries the default gateway 10.12.19.78. The address on eth1 has a DNS name of its own, dc2-a-vcmsn001.adm.example.net, so the server is …vcmsr001 from the inside and …vcmsn001 towards the firewall.
A server with its default route on the outside interface would send every answer to DNS, NTP, syslog, monitoring and my SSH session out through the relay VLAN. Three static routes on eth0 prevent that: the three private address blocks of RFC 1918 go to 10.12.19.46, the gateway of the internal VLAN.
output 4 lines
# Route to internal networks 192.168.0.0/16 via 10.12.19.46 10.0.0.0/8 via 10.12.19.46 172.16.0.0/12 via 10.12.19.46
The files are Config documents: ifcfg-eth0, ifcfg-eth1, route-eth0 and route6-eth0. For IPv6 there are no summary routes; route6-eth0 lists the management segments one by one.
The two relays were not built the same way, and diff shows it. On the first relay the interface files are tidy and ifcfg-eth1 begins with ## Ansible managed. On the second relay ifcfg-eth0 looks like the installer's file edited by hand. Its ifcfg-eth1 carries the IPv6 address 2001:db8:a2:b6f::f:2/64, which is the address of eth0 of the first relay and lies in the prefix of the wrong VLAN; the design wanted 2001:db8:a2:b6b::f:2. The first design document says "SMTP server supports IPv4 and IPv6, but this solution is based on IPv4", and Postfix on the relays has inet_protocols = all, so a relay would try IPv6 for a destination that publishes it. I have no record of whether IPv6 delivery ever worked from the second relay, or was ever noticed not to.
NAT and the public names
The firewall translates the two eth1 addresses one to one. The only rule the design lists for the relay VLAN is outbound: TCP from the two NAT-ed addresses to any address, port 25. This is where the Source material stops agreeing with itself.
| Source | Public name of DC2-A-VCMSR001 | Public name of DC2-B-VCMSR001 |
|---|---|---|
| Design, DNS tables (A, MX and PTR records) and NAT table | mx1.ad-dc2.example.net | mx2.ad-dc2.example.net |
| Design, network diagram of site 2 | mx3.ad.example.net | mx4.ad.example.net |
| Install notes, server information | mx3.ad.example.net | mx4.ad.example.net |
main.cf, myhostname | mx3.ad.example.net | mx4.ad.example.net |
So the servers introduce themselves in SMTP as mx3 and mx4 of ad.example.net, the mail domain of site 1, where mx1 and mx2 are the relays of site 1. mydomain is ad.example.net as well, on servers that relay for ad-dc2.example.net. The DNS tables of the design promise different names in a different zone. I cannot tell from the archive which records existed in the public DNS. It matters, because a receiving server may compare the name in the greeting with the PTR record of the connecting address, and the design's PTR records say mx1.ad-dc2.example.net and mx2.ad-dc2.example.net. The NAT table has a slip of its own: it calls the internal name of the second relay dc2-b-vcmsn002, where the network table and the install notes say dc2-b-vcmsn001. More on the tables in Network, DNS and firewall.
Postfix as configured
The whole file is a Config document: Postfix main.cf, relay servers. Compared with the internal servers (Internal servers: Postfix) it is what remains when authentication, LDAP, the content filter and the server side of TLS are taken away.
# My networks = Internal Email Server only mynetworks = 10.12.19.33/32 10.12.19.41/32 10.12.19.42/32 # Allowed clients are only from my_networks = Internal Email Server only smtpd_recipient_restrictions = reject_non_fqdn_recipient, reject_unknown_recipient_domain, permit_mynetworks, reject_unauth_destination, reject
| Parameter | Value | What it means here |
|---|---|---|
myhostname | mx3.ad.example.net, mx4.ad.example.net | The name in the greeting and in HELO towards the internet |
mynetworks | the VIP 10.12.19.33 and the two nodes of the internal pair | The only clients that may relay. The mailbox server 10.12.19.43 is not in the list; it sends through the internal pair |
smtpd_recipient_restrictions | ends with reject | Whatever permit_mynetworks did not let in is refused |
| SASL | not configured | No authentication on the relays. The trusted-network approach was chosen in the first design: three fixed addresses, nothing with a password |
smtp_tls_security_level | may | Outbound: STARTTLS when the other side offers it, clear text when it does not, no certificate check |
smtp_tls_protocols, smtp_tls_mandatory_protocols | !SSLv2, !SSLv3 | Outbound protocol floor |
smtpd_tls_cert_file, smtpd_tls_security_level | not set | Inbound: the relay offers no STARTTLS, so the hop from the internal pair to the relay is clear text inside VLAN 1168 |
transport_maps | hash:/etc/postfix/transport | One entry, see below |
smtp_header_checks | pcre:/etc/postfix/header_checks | The header cleaning, see below |
message_size_limit | 20971520 | 20 MB, the same as on the internal servers |
notify_classes, bounce_notice_recipient | commented out | Sending non-delivery notices to one functional mailbox was prepared and not switched on |
The install notes generate a key and a certificate request for each relay and three Diffie-Hellman parameter files, under the heading "Certificates (currently not used on this server, but certificate is generated)". The design lists the certificate paths for both relays; main.cf references none of them. The smtpd_tls_protocols lines in the hardening block therefore restrict a TLS server that is not switched on.
The hardening block (disable_vrfy_command, smtpd_helo_required, the connection and rate limits, strict_rfc821_envelopes, reject_unauth_pipelining, smtpd_forbidden_commands) is the same block as on the internal servers and is explained there.
master.cf is the distribution's file: one smtpd on port 25, no submission, no SMTPS, no filter. diff finds no difference between the two relays.
The transport table has one line.
output 1 line
ad-dc2.example.net smtp:[10.12.19.43]
Without it a relay would treat the site's own domain like any other and look up its MX records, which point at the relays themselves. The comment in the file says what the line is for: "Mainly non-delivery email notification". When a server on the internet refuses a message, it is the relay that writes the bounce, and the bounce is addressed to an internal sender.
Is mail from the internet accepted?
The design publishes MX records for both mail domains, priority 10 for the first relay and 20 for the second, which reads like an inbound path with a preference. The configuration shows that there is none, on three levels. The firewall table has no rule from the internet to the relay VLAN, only the outbound one. The first design document made it a requirement: "connecting to Relay Email Server will not be allowed from internet". And Postfix itself would refuse: a client that is not one of the three addresses in mynetworks falls through reject_unauth_destination to the final reject, whatever the recipient.
The MX records are there for SPF. The TXT record of the design is v=spf1 mx -all for each domain: only the hosts named in the MX records may send mail for the domain, everything else fails hard. With the public addresses of the relays behind the MX names, the record authorises exactly the two NAT addresses without repeating them. The side effect is that a non-delivery report generated later by a remote server, or any reply to an address in the domain, is addressed to an MX that does not answer. The first design document accepted that in so many words: the main function of the relay is sending mail to external domains, not receiving it.
Removing internal information from headers
Outbound mail passes header_checks in the SMTP client, at the moment of delivery.
Six rules: every Received: line is dropped (the appliance or application that sent the message, the internal server, the content filter and the relay's own line), Mime-Version is rewritten to the bare 1.0, and User-Agent, X-Mailer, X-Enigmail and X-Originating-IP are dropped. The Config document has the rule-by-rule table.
The reason is in the comment, "Remove sensitive information from email header": host names, addresses and software versions of a management network should not travel in every alert sent to a partner. The price is the trace. A recipient sees a message that begins at the relay's public address, with no hop and no timestamp before it, so a delay can only be investigated in the logs of the internal servers and the relay. Because smtp_header_checks works in the SMTP client, it applies to everything a relay delivers, the transport line to the mailbox server included. And the rule does not catch everything: a Message-ID generated by an internal Postfix still ends with that server's host name.
Redundancy on SMTP level
There is no cluster and no virtual address on the relay side. The design calls it "redundant on SMTP level": the internal servers know both relays and prefer one.
relayhost = [10.12.19.34] smtp_fallback_relay = [10.12.19.35]
Both nodes of the internal pair carry these two lines (Postfix main.cf, internal servers). Postfix hands a message to the fallback when the primary cannot be reached. The relays do not know about each other. A message that already sits in the queue of a relay that goes down waits there until the server is back; nothing replicates a queue. The first design document had asked for less, a single relay per site with "HA on virtualization layer will be sufficient"; the second relay came with the HA concept of 2019. Monitoring per the design is two checks per relay: the postfix process must run and the file systems must stay under their thresholds.
Installation
The install notes of DC2-A-VCMSR001 are short. The commands below ran as root; the second relay got the same ones without mc.
$ yum install postfix mc $ systemctl stop postfix $ mkdir -p /etc/postfix-original $ cp -R /etc/postfix/* /etc/postfix-original/ $ rm -rf /etc/postfix/* $ vi /etc/postfix/main.cf $ vi /etc/postfix/transport $ vi /etc/postfix/header_checks $ systemctl start postfix $ systemctl enable postfix
The notes keep the stock configuration aside, empty the directory and write the files new. They do not record postmap /etc/postfix/transport, which a hash: table needs before Postfix can read it, nor how master.cf came back after the rm; the archived master.cf is the stock one. Then the host firewall, also as root.
$ firewall-cmd --permanent --add-service=ssh $ firewall-cmd --permanent --add-service=smtp $ firewall-cmd --reload
The service is added to the default zone, so port 25 is open on both interfaces as far as the host is concerned. What keeps the internet out is the network firewall and the last line of smtpd_recipient_restrictions. The notes end with the key and certificate request described above and with subscription-manager config --rhsm.manage_repos=0, which stops the subscription manager from managing the repository file. The install notes of the relays have no SELinux section: the first design document sets enforcing mode with the targeted policy, and unlike the internal servers the relays got no local policy module.
The first concept
The relay has a design document of its own from August 2017, a draft with the version 00.01, written for site 1 when the system was one internal server and one relay.
| Topic | First concept, 2017 | As built, 2019 |
|---|---|---|
| Relays per site | one, DC1-A-VCMSR001 | two |
| Availability | "High Availability on application Level is not required" | primary and fallback relay |
| Allowed client | mynetworks = 10.11.19.33/32, the internal server | the VIP and both nodes of the pair |
| Public name | mx1.ad.example.net, one A, MX, PTR record and the SPF record v=spf1 mx -all | two of each per site |
message_size_limit | 10485760 | 20971520 |
queue_minfree | 20971520 | 31457280 |
| Sizing | 2 vCPU, 4 GiB RAM, 32 GiB disk | see Virtual machines and OS build |
Its firewall section has two rules, written as iptables rules: TCP from the internal server, source ports 1024 to 65535, to port 25 of the relay, and TCP from the relay to port 25 of any address. For the network firewall between the segments it says only that the rules "are the same". The header rules are already there, character for character. The chapter "Known issues and/or limitations" has one sentence: "No known issues or limitations have been found."
What I would do differently
- One name per relay, everywhere.
myhostname, the A record, the PTR record and the name in the design should be the same string. Here the design says one thing and the servers another, and nobody can check today which was true. - Use the certificate that was generated. The key and the request exist on both relays. Switching on
smtpd_tls_security_level = maywith them would have encrypted the hop from the internal pair, which carries every message the content filter was told not to encrypt. - Build both relays from the same source. One interface file says "Ansible managed", its twin on the other server is hand-edited and has a wrong address. Two servers that must be identical should not be typed twice.
- Leave the
Received:lines alone. RFC 5321, section 4.4, says a mail program must not change or delete aReceived:line that was added before; the first rule ofheader_checksdoes exactly that. Internal host names can be kept out of the trace in less blunt ways. - Sign the mail. The design has SPF and nothing else. Since February 2024 Gmail and Yahoo require SPF or DKIM, valid forward and reverse DNS and TLS from every sender, and SPF, DKIM and DMARC together from bulk senders. Postfix has no DKIM signer of its own, so a relay built today gets an external one.