LINUXOR.SK ... open source notes ...

Email 03 - Network, DNS and firewall

category: solutionz · date: 2019-12-31 · updated: 2026-10-02 · author: LALA

Email Solution · Previous: Logical design · Next: Virtual machines and OS build

A mail system is judged from outside by three things it does not control alone: the address its mail comes from, what DNS says about that address, and what the firewalls let through. This part has the two VLANs, every address, the public DNS records, the firewall flows and the routing of the relays, which stand in both VLANs. It is also the part where the design contradicts itself most, so the tables of the design were checked against the configuration of the five site 2 servers, and the result of that check is at the end.

Two VLANs

VLANPurposeSite 1 (DC1)Site 2 (DC2)
1168, internalClients to the internal email servers; internal servers to the relays10.11.19.32/28, 2001:db8:a1:b6f::/6410.12.19.32/28, 2001:db8:a2:b6f::/64
1172, relayRelays to the mail servers on the internet10.11.19.64/28, 2001:db8:a1:b6b::/6410.12.19.64/28, 2001:db8:a2:b6b::/64

I read both VLANs as stretched over datacenters A and B of a site: the design does not say so, but a virtual address could not move between the two nodes of the pair otherwise. In site 2 the gateway of the internal network is 10.12.19.46 and 2001:db8:a2:b6f::1, that of the relay network 10.12.19.78 and 2001:db8:a2:b6b::1. Everything that leaves VLAN 1172 passes a firewall that translates the relays' addresses to public ones.

The split repeats on the network level what the concept does with servers. The three internal servers have one interface, in VLAN 1168, and their default route points into the management infrastructure. Only the relays have a second interface, and only that interface has a way to the internet.

Site 2

mermaid
flowchart TB
  inet["Internet"]
  fw["Firewall, NAT"]
  n1["10.12.19.65 is translated to 203.0.113.44"]
  n2["10.12.19.66 is translated to 203.0.113.48"]
  v1172["VLAN 1172, relay, 10.12.19.64/28"]
  v1168["VLAN 1168, internal, 10.12.19.32/28"]
  vip(["VIP on eth0, 10.12.19.33 and 2001:db8:a2:b6f::f:1"])
  subgraph dca["Datacenter A, hypervisors DC2-A-CKVM001 to 004"]
    ra["DC2-A-VCMSR001"]
    x2["DC2-A-VCMSX002"]
    xa["DC2-A-VCMSX001"]
  end
  subgraph dcb["Datacenter B, hypervisors DC2-B-CKVM001 to 004"]
    rb["DC2-B-VCMSR001"]
    xb["DC2-B-VCMSX001"]
  end
  inet --- fw
  fw --- v1172
  n1 -.- fw
  n2 -.- fw
  v1172 -- "eth1 10.12.19.65, default gateway" --- ra
  v1172 -- "eth1 10.12.19.66, default gateway" --- rb
  v1168 -- "eth0 10.12.19.34" --- ra
  v1168 -- "eth0 10.12.19.35" --- rb
  v1168 -- "eth0 10.12.19.43" --- x2
  v1168 -- "eth0 10.12.19.41" --- xa
  v1168 -- "eth0 10.12.19.42" --- xb
  vip -.- xa
  vip -.- xb

The addresses below are those of the design, and each IPv4 address was found in the ifcfg-eth* file of the server it belongs to.

ServerVLANInterfaceIPv4IPv6DNS name in adm.example.net
DC2-A-VCMSX0011168eth010.12.19.412001:db8:a2:b6f::f:4dc2-a-vcmsx001
DC2-B-VCMSX0011168eth010.12.19.422001:db8:a2:b6f::f:5dc2-b-vcmsx001
VIP of the pair1168eth010.12.19.33/282001:db8:a2:b6f::f:1/64dc2-s-xcmsx001
DC2-A-VCMSX0021168eth010.12.19.432001:db8:a2:b6f::f:6dc2-a-vcmsx002
DC2-A-VCMSR0011168eth010.12.19.342001:db8:a2:b6f::f:2dc2-a-vcmsr001
DC2-A-VCMSR0011172eth110.12.19.652001:db8:a2:b6b::f:1dc2-a-vcmsn001
DC2-B-VCMSR0011168eth010.12.19.352001:db8:a2:b6f::f:3dc2-b-vcmsr001
DC2-B-VCMSR0011172eth110.12.19.662001:db8:a2:b6b::f:2dc2-b-vcmsn001

A relay has two names: vcmsr for the interface the internal servers talk to and vcmsn for the one that is translated.

Site 1

The network figure of site 1 is the same picture with other labels: the same two VLAN numbers, the same five servers with DC1- names on hypervisors DC1-A-CKVM001 to 004 and DC1-B-CKVM001 to 004, the same last octets. A second diagram would add nothing; these are the values.

ServerVLANInterfaceIPv4IPv6
DC1-A-VCMSX0011168eth010.11.19.412001:db8:a1:b6f::f:4
DC1-B-VCMSX0011168eth010.11.19.422001:db8:a1:b6f::f:5
VIP of the pair, dc1-s-xcmsx0011168eth010.11.19.33/282001:db8:a1:b6f::f:1/64
DC1-A-VCMSX0021168eth010.11.19.432001:db8:a1:b6f::f:6
DC1-A-VCMSR0011168eth010.11.19.342001:db8:a1:b6f::f:2
DC1-A-VCMSR0011172eth110.11.19.652001:db8:a1:b6b::f:1
DC1-B-VCMSR0011168eth010.11.19.352001:db8:a1:b6f::f:3
DC1-B-VCMSR0011172eth110.11.19.662001:db8:a1:b6b::f:2

These are planned values. In October 2019 the new concept was not yet deployed in site 1, and there is no site 1 configuration in the Source material to check them against.

NAT and public DNS

The firewall translates the eth1 address of each relay to one public IPv4 address. The design describes NAT for IPv4 only and says nothing about how, or whether, the relays reach the internet over IPv6.

SiteRelay interfaceInternal addressPublic address
1dc1-a-vcmsn00110.11.19.65198.51.100.44
1dc1-b-vcmsn00110.11.19.66198.51.100.48
2dc2-a-vcmsn00110.12.19.65203.0.113.44
2dc2-b-vcmsn00110.12.19.66203.0.113.48

For a receiving mail server to accept what the relays send, the public DNS of example.net needs four kinds of records per site. This is the set the design specifies for site 1; site 2 has the same set with the domain ad-dc2.example.net and the two 203.0.113 addresses.

TypeNameValueWhy
Amx1.ad.example.net198.51.100.44The name of the first relay as the internet sees it
Amx2.ad.example.net198.51.100.48The second relay
MXad.example.net10 mx1.ad.example.netThe mail exchangers of the domain
MXad.example.net20 mx2.ad.example.netThe second, lower preference
TXTad.example.netv=spf1 mx -allSPF: only the hosts named in the MX records may send for this domain, everything else fails
PTR44.100.51.198.in-addr.arpa.mx1.ad.example.netReverse lookup of the sending address gives the name back
PTR48.100.51.198.in-addr.arpa.mx2.ad.example.netThe same for the second relay

The MX records are there for SPF, not for receiving. The SPF policy is written as mx, so the domain must publish MX records that resolve to the relays' public addresses, but the 2017 relay design states that connections from the internet to the relay "will not be allowed", and the firewall tables of 2019 have no flow from the internet inwards. A mail server on the internet that tries to deliver to one of these addresses cannot connect. The Source material holds the specification of the records, not a zone file or a query against the public DNS, so I cannot show what was actually published.

Firewall flows

The design has four tables of six columns, one per server role and site, that repeat the same flows with other addresses. Condensed to one table per role, with the site 2 values, they are these. The source port is 1024-65535 throughout.

SourceDestinationProtocol and portPurpose
Email clients, written as "Any ???" in the designVIP 10.12.19.33 and both nodes, 10.12.19.41, 10.12.19.42TCP 25, TCP 465SMTP with STARTTLS, SMTPS
Admin VPN 10.11.20.0/2510.12.19.41, 10.12.19.42, 10.12.19.43TCP 22SSH
Admin VPN 10.11.20.0/2510.12.19.43TCP 443Webmail
Internal serversDomain controllers 10.12.16.209, 10.12.16.210"AD (LDAP)", TCP and UDPLDAPS on 636 for Postfix and Dovecot, and the host's own AD integration through sssd
Internal serversInfoblox 10.11.16.145, 10.11.18.145TCP and UDP 53; 123DNS, NTP
Internal serversCentral syslog 10.12.17.113TCP and UDP 514Logs
Internal serversProxy 10.12.16.113TCP 3128Squid
Internal serversGitLab 10.12.17.33TCP 22etckeeper pushes /etc
Internal serversCobbler 10.12.16.97TCP 443Package repository
Internal serversGraphite 10.12.18.17, 10.12.18.18, 10.12.18.19TCP 80, TCP 443Metrics
Internal serversSensu servers 10.12.16.129, 10.12.16.130, 10.12.16.131ICMPPing
Internal serversRabbitMQ 10.12.17.1, 10.12.17.2TCP 5672Sensu transport

The relays have the same nine outgoing flows to the infrastructure services, from 10.12.19.34 and 10.12.19.35, the same SSH flow from the admin VPN, and one flow that nobody else has.

SourceDestinationProtocol and portPurpose
Admin VPN 10.11.20.0/2510.12.19.34, 10.12.19.35TCP 22SSH
Relays, eth0The nine infrastructure destinations of the table aboveas aboveas above
Relays, eth1: 10.12.19.65, 10.12.19.66, translatedAnyTCP 25SMTP to the mail servers on the internet

Site 1 has the same flows with its own addresses: the servers in 10.11.19.x, four domain controllers instead of two (10.11.88.145, 10.11.88.146, 10.11.16.209, 10.11.16.210), and the infrastructure services at the same last octets in 10.11.16.x to 10.11.18.x. Infoblox and the admin VPN have the same addresses in the tables of both sites: the Infoblox pair exists in site 1 only, and the VPN network given for site 2 is the site 1 one.

Four remarks on these tables. The client row was never finished: "Any ???" stands in the released version 1.0 of both sites, so the design does not say which networks may reach the mail service. The outgoing rows of the internal table name only the two nodes of the pair as source, not the third server, which needs the directory, DNS and the rest just as much. NTP is listed as TCP; it is UDP 123. And the tables are IPv4 only, although every server has an IPv6 address and the VIP has one too.

The traffic between the mail servers themselves stays inside VLAN 1168 and does not cross the network firewall, so it is in none of the tables. From the configuration it is this.

FromToPortWhat
The pairRelays, 10.12.19.34 then 10.12.19.35TCP 25relayhost and smtp_fallback_relay
The pairThird internal server, 10.12.19.43TCP 25Transport for ad-dc2.example.net
Third internal serverThe pair, 10.12.19.41 then 10.12.19.42TCP 25Its relayhost and fallback
RelaysThird internal serverTCP 25Bounces for the own domain
Node A and node B of the pairEach otherVRRP, unicastKeepalived
Node A and node B of the pairEach otherTCP 22rsync over SSH of keys and certificates

Each server also runs firewalld. The install notes open the following, as root, on the pair.

bash
$ firewall-cmd --permanent --add-service=ssh
$ firewall-cmd --permanent --add-service=smtp
$ firewall-cmd --permanent --add-port=465/tcp
$ firewall-cmd --reload
ServerOpened in firewalld
The pairssh, smtp, 465/tcp
Third internal serverssh, smtp, 465/tcp, https, and in a second block of the notes also imap
Relaysssh, smtp

Two things do not fit. The design says IMAP is reachable from localhost only, while the notes of the third server open the imap service; the notes contain the firewall block twice, once without and once with it, and do not say which is the final state. And the notes of the pair have no rule for VRRP, IP protocol 112, which a default firewalld zone does not accept from the other node; how the Keepalived advertisements were let through is not recorded.

Routing on the relays

A host with two interfaces has to decide where the default route goes. On the relays it goes out of eth1, towards the internet, because the destinations there cannot be enumerated; the destinations inside can. From ifcfg-eth0 and ifcfg-eth1 of DC2-A-VCMSR001:

ini
# ifcfg-eth0
IPADDR="10.12.19.34"
NETMASK="255.255.255.240"
GATEWAY="10.12.19.46"
DEFROUTE="no"
IPV6_DEFAULTGW="2001:db8:a2:b6f::1"
IPV6_DEFROUTE="no"
# ifcfg-eth1
IPADDR="10.12.19.65"
NETMASK="255.255.255.240"
GATEWAY="10.12.19.78"
DEFROUTE="yes"
IPV6_DEFAULTGW="2001:db8:a2:b6b::1"
IPV6_DEFROUTE="yes"

Both files carry NM_CONTROLLED="no": the interfaces are brought up by the classic network scripts, which read two more files for static routes. route-eth0 has three lines, all through the internal gateway 10.12.19.46: 192.168.0.0/16, 10.0.0.0/8 and 172.16.0.0/12. They are the three private ranges of RFC 1918; together they say "everything private goes inside". The management infrastructure, the directory, DNS, monitoring and the internal mail servers on other subnets are all reached this way, and only what is left over, public addresses, follows the default route to the NAT firewall. The file is a Config document: route-eth0, relay servers.

IPv6 has no private ranges to summarize, so route6-eth0 is a list of twelve /64 prefixes through 2001:db8:a2:b6f::1, one per management network the relay talks to. The prefixes of the Sensu servers (bb1), RabbitMQ (baf) and Graphite (b85) can be recognized in it. The file is identical on both relays and has three flaws as archived: one prefix is listed twice, one line (2001:db8:a2:b99::/64) has no via, and the list is not complete. The prefix of the central syslog server, 2001:db8:a2:b9b::/64, is not in it, and neither is that of the first IPv6 resolver in resolv.conf, 2001:db8:a1:c0b::/64; the list has 2001:db8:a2:c0b::/64 instead. On the first relay whatever is missing leaves through eth1, which has the IPv6 default route; for the second relay, whose eth1 is misconfigured for IPv6, the resulting route is not recorded. A list like this has to be maintained by hand for every new network, which the three IPv4 lines do not. The file is a Config document: route6-eth0, relay servers.

Postfix on the relay needs no setting for any of this. It listens on all interfaces (inet_interfaces = all) and lets the kernel choose the outgoing one; what keeps the internet-facing interface from being a way in is the firewall in front of VLAN 1172 and mynetworks, which holds the VIP and the two nodes of the pair and nothing else. The rest of the relay is in Relay servers and in the Config document Postfix main.cf, relay servers.

What the check against the configuration found

Where in the designIt saysThe configuration and notes say
Network table, VLAN 1172DC1 10.12.19.64/28, DC2 10.11.19.64/28The IPv4 networks are swapped between the sites: eth1 of both site 2 relays is in 10.12.19.64/28. The IPv6 prefixes in the same rows are right
Network summary of site 1DC1-B-VCMSR001 eth0 is 10.12.19.35A site 2 address in a site 1 table; the firewall table of site 1 has 10.11.19.35
NAT tableInternal names dc1-b-vcmsn002, dc2-b-vcmsn002The network summaries and the install notes have dc2-b-vcmsn001: the interface name of the first relay in datacenter B, not a second device
DNS records and NAT table, site 2Public names mx1.ad-dc2.example.net, mx2.ad-dc2.example.netThe network figure of site 2, the install notes and Postfix say mx3.ad.example.net and mx4.ad.example.net for the same two public addresses, and mydomain on both relays is ad.example.net
Client settings, site 2VIP IPv4 10.12.16.3310.12.19.33 everywhere else and in keepalived.conf
Firewall table of site 1DC2-B-VCMSX001 / 10.11.19.42A site 2 name with a site 1 address; DC1-B-VCMSX001 is meant

The fourth row is the one with consequences. The relays of site 2 introduce themselves as mx3.ad.example.net and mx4.ad.example.net, hosts of the site 1 domain, and that is myhostname in both main.cf files, so it is what they say in HELO. The design's record tables instead plan mx1 and mx2 under ad-dc2.example.net. Receiving servers compare the HELO name, the PTR record of the connecting address and the SPF record of the sender's domain, and the two naming schemes cannot both be consistent with one set of records. Which records were published for site 2 I cannot tell from the Source material. The mail domain of site 2 itself is not in doubt: the internal servers use ad-dc2.example.net, with one exception on node B that is described in Internal servers: Postfix.

The configuration trees have mistakes of their own that the design does not have.

HostFileFound
DC2-B-VCMSX001, DC2-A-VCMSX002ifcfg-eth0NETMASK="255.255.255.0" on a /28 network; node A and both relays have 255.255.255.240
DC2-B-VCMSR001ifcfg-eth1IPV6ADDR=2001:db8:a2:b6f::f:2/64: an address from the internal prefix, and the one that belongs to eth0 of the other relay, where the design plans 2001:db8:a2:b6b::f:2; IPV6_DEFROUTE is no
DC2-B-VCMSR001Install notes, headerThe IPv6 address of the internet side is given as 2001:db8:a2:b6b::f:1, copied from the first relay

With the wrong netmask a server treats all of 10.12.19.0/24 as directly connected, which includes the relay network 10.12.19.64/28; the addresses it actually talks to are either inside its own /28 or outside the /24, so the mistake stays without effect as long as nothing internal needs to reach the relays' outer interfaces. The second relay, as archived, has no usable IPv6 address on its outer interface. Its ifcfg-eth1 sets IPV6_DEFAULTGW to the gateway of the relay network and IPV6_DEFROUTE="no"; which IPv6 default route that produced under the classic network scripts is not recorded. Since the design treats the way to the internet as IPv4 with NAT, it is quite possible that nobody noticed; the Source material has no test of outgoing IPv6.

← solutionz