LINUXOR.SK ... open source notes ...

Email 10 - High availability and key synchronization

category: solutionz · date: 2019-12-31 · updated: 2026-10-02 · author: LALA

Email Solution · Previous: Automatic email encryption · Next: Relay servers

Requirement NR2 calls the email system mission critical and asks for high availability "on application layer". The first concept of 2017–2018 had one internal server and one relay per site; the final design of 2019 doubled both. This Article is about what the doubling consists of: one virtual address held by Keepalived, a fallback line in Postfix, and two small services that keep the recipients' certificates and keys the same on both internal servers. It is also about what is not redundant.

Three server roles, three answers

RoleServers at site 2RedundancyWhat a client has to do
Internal SMTPDC2-A-VCMSX001, DC2-B-VCMSX001A virtual address (VIP) that moves between two complete, running serversNothing; it talks to dc2-s-xcmsx001.adm.example.net
RelayDC2-A-VCMSR001, DC2-B-VCMSR001On SMTP level: the internal servers know a primary and a fallback relayNothing; the client here is Postfix
Mailboxes and webmailDC2-A-VCMSX002NoneWait

The design says of the internal pair that it "forms a Active-Passive cluster, but from service perspective it is a Active-Active deployment", because the only cluster resource is the address: Postfix and Dovecot run on both nodes all the time, with inet_interfaces = all, and each node accepts mail on its own address as well. Nothing is started or stopped at a failover.

The relays need no cluster. main.cf of the internal servers has relayhost = [10.12.19.34] and smtp_fallback_relay = [10.12.19.35]; if the first relay cannot be reached or answers with a temporary error, Postfix hands the message to the second. That is in Relay servers.

The third internal server is called "stand alone" in the design, and it is the single point of the system by design: the mailboxes are a Maildir on its local data disk, and nothing replicates them. While it is down, the pair keeps accepting mail for the site's own domain and holds it in the queue; the webmail is gone. See Mailbox server.

mermaid
flowchart TB
  cl["Servers, applications, appliances"] -- "SMTP 25, SMTPS 465" --> vip["VIP dc2-s-xcmsx001, 10.12.19.33"]
  subgraph a1["DC2-A-VCMSX001, MASTER, priority 100"]
    ka["Keepalived"]
    pa["Postfix"]
    da["Dovecot"]
    sa["smime and pgp directories"]
    ka -. "killall -0 master, every second" .-> pa
    ka -. "killall -0 dovecot, every second" .-> da
  end
  subgraph b1["DC2-B-VCMSX001, BACKUP, priority 50"]
    kb["Keepalived"]
    pb["Postfix"]
    db["Dovecot"]
    sb["smime and pgp directories"]
    kb -. "killall -0 master, every second" .-> pb
    kb -. "killall -0 dovecot, every second" .-> db
  end
  vip --- pa
  vip -. "moves here on failure" .- pb
  ka <-- "VRRP, unicast, every second" --> kb
  sa -- "rsync over ssh, with delete" --> sb
  sb -- "rsync over ssh, with delete" --> sa
  pa -- "relayhost" --> r1["DC2-A-VCMSR001"]
  pa -. "smtp_fallback_relay" .-> r2["DC2-B-VCMSR001"]
  pa -- "own domain" --> mb["DC2-A-VCMSX002, mailboxes, single"]

Keepalived

The configuration is a Config document: keepalived.conf.

SettingValue at site 2Why
virtual_ipaddress10.12.19.33/28 and 2001:db8:a2:b6f::f:1/64 on eth0The address behind the name the clients use, in both protocols
state, priorityMASTER and 100 on node A, BACKUP and 50 on node BNode A holds the address whenever it can
unicast_src_ip, unicast_peer10.12.19.41 and 10.12.19.42, crossedVRRP between two known peers, nothing asked of the network
authenticationauth_type PASS with a shared passwordEach member must authenticate, says the design
vrrp_script check_postfix_masterkillall -0 master, interval 1, fall 2, rise 2The address must leave a node without Postfix
vrrp_script check_dovecotkillall -0 dovecot, interval 1, fall 2, rise 2Without Dovecot nobody can authenticate, so Postfix alone is useless
advert_int1One advertisement per second

killall -0 sends no signal; it only reports whether a process of that name exists. master is the Postfix master process. For site 1 the design gives the same roles and the address 10.11.19.33/28 and 2001:db8:a1:b6f::f:1/64 under the name dc1-s-xcmsx001.adm.example.net.

The logical diagram of the design labels the line between the two Keepalived boxes "MULTICAST/VRRP". The text of the same design says "cluster members communicate with other members thru unicast", and the configuration is unicast. The diagram is wrong.

The install notes are short. As root, on both nodes:

bash
$ yum install keepalived
$ vi /etc/sysctl.d/postfix.conf
$ init 6
$ mkdir -p /etc/keepalived-original
$ cp -R /etc/keepalived/* /etc/keepalived-original/
$ vi /etc/keepalived/keepalived.conf
$ systemctl start keepalived
$ systemctl enable keepalived

The reboot in the middle activates sysctl.d/postfix.conf, which sets net.ipv4.ip_nonlocal_bind and net.ipv6.ip_nonlocal_bind to 1. Postfix on both nodes sends from the virtual address (smtp_bind_address = 10.12.19.33 and smtp_bind_address6), and the relays and the mailbox server list that address in mynetworks. A node that does not hold the address could not even bind to it without this setting.

Keepalived runs confined by SELinux, and keepalived_t may neither look at other daemons nor signal them. The module that allows the two checks is keepalived-local-1.0.te; it is a harvest of the audit log and carries much more than the two signull rules that matter.

The notes record no test: no forced failover, no measured time. They also show no firewalld rule for VRRP (IP protocol 112); only ssh, smtp and 465/tcp are opened, and the firewall tables of the design have no row for traffic between the two nodes. With firewalld's default behaviour the backup would not receive the advertisements of the master, and both would hold the address. One of the links I kept at the end of the notes is a question titled "Keepalived VIP is active on both servers". Whether a rule was added outside the notes, I cannot tell from the Source material.

A failover, step by step

mermaid
sequenceDiagram
  participant C as Client
  participant A as Node A, priority 100
  participant B as Node B, priority 50
  A->>B: VRRP advertisement every second
  C->>A: SMTP to the VIP
  Note over A: Postfix master or Dovecot stops
  A->>A: check fails twice, instance in fault state
  A->>A: VIP removed from eth0
  Note over B: advertisements from A stop
  B->>B: becomes master, adds VIP to eth0
  C->>B: SMTP to the VIP
  Note over A: service is started again
  A->>A: check passes twice
  A->>B: VRRP advertisement, priority 100
  B->>B: back to backup, VIP removed
  C->>A: SMTP to the VIP

Neither check has a weight, so a failed check does not lower the priority; it takes the instance out of the election. There is no nopreempt, so node A takes the address back as soon as its checks pass again. Every move cuts the SMTP sessions that were open on the address; a client that retries gets the other node.

What moves and what does not

Synchronizing certificates and keys

The encryption filter of Automatic email encryption looks for <recipient>.cer in smime/ and <recipient>.asc in pgp/ on the node where the message happens to arrive. Both nodes must therefore have the same files, and the design wants that an administrator uploads a new one to either node only. Two services per node do that, one per directory: sync-smime.sh and sync-pgp.sh, started by sync-smime.service and sync-pgp.service.

bash
while true; do
    inotifywait -r -e modify,attrib,close_write,move,create,delete $PGP_KEYS_DIR
    chown bash-postfix-encrypt-filter:bash-postfix-encrypt-filter $PGP_KEYS_DIR/*
    chmod 400 $PGP_KEYS_DIR/*
    rsync -avz -e "ssh -o StrictHostKeyChecking=no -i /home/bash-postfix-encrypt-filter/.ssh/id_rsa" $PGP_KEYS_DIR/ $USER@$HOST:$PGP_KEYS_DIR/ --delete
done

inotifywait blocks until something in the directory is written, created, moved, deleted or changes its attributes, and then exits. The loop gives the files to the filter account with mode 400, which is why a key can be copied in as root without further care, and pushes the directory to the other node: -a keeps owner, mode and times, -z compresses, --delete removes on the other side whatever is not on this side. Then it waits again. The loop runs as root, but logs in on the peer as the filter account with that account's key.

The key pairs were made as the filter account, on each node, and the public halves exchanged by hand.

bash
$ yum install rsync inotify-tools
$ su - bash-postfix-encrypt-filter
$ ssh-keygen -t rsa
$ cat ~/.ssh/id_rsa.pub
$ vi ~/.ssh/authorized_keys
$ chmod 600 ~/.ssh/authorized_keys

The first line ran as root, the rest as the account; authorized_keys on each node received the public key of the other. Then, as root on both nodes:

bash
$ cd /usr/local/bin/
$ chown bash-postfix-encrypt-filter:bash-postfix-encrypt-filter ./bash-postfix-encrypt-filter-sync*
$ chmod 550 ./bash-postfix-encrypt-filter-sync*
$ systemctl start bash-postfix-encrypt-filter-sync-pgp.service
$ systemctl start bash-postfix-encrypt-filter-sync-smime.service
$ systemctl enable bash-postfix-encrypt-filter-sync-pgp.service
$ systemctl enable bash-postfix-encrypt-filter-sync-smime.service

The design states that systemd restarts a synchronization service it finds not running. The unit files have no Restart= line, so systemd does not; a loop that has died stays dead until the next boot or a manual start. Nothing monitors it, as far as the design's monitoring tables go.

Only the files are synchronized. The GnuPG keyring that the filter really encrypts with lives in the home directory of the account on each node and is filled from the .asc file at first use there.

What two pushes with delete do to each other

"Bidirectional" here means two one-way mirrors pointed at each other. Each push makes the peer equal to the pusher; there is no merge and no memory of what was deleted. Reasoning from the script, without any recorded incident:

SituationResult
A file is added on one node, the other is idleIt is pushed. The arrival is an event on the peer, which pushes back; nothing differs any more, so that push is empty
A file is deleted on one nodeThe push deletes it on the peer
Files are added on both nodes at nearly the same timeEach push carries its own new file and deletes the one it does not know. Depending on timing, one or both files can disappear
The same file is changed on both nodesThe last push wins; rsync is not told to keep the newer one
A key is deleted on node A while the service or the whole node B is downNothing is pushed later on its own, because an event that nobody waited for is lost. The next event on node B pushes the deleted key back to A
A node is rebuilt with empty directoriesThe first event in its directory pushes that state, and the peer loses everything that is not on the new node
A change arrives while the loop is busy with rsyncIt is not seen. The directories differ until the next event on that node

The procedure that avoids all of this is the one the design implies: change one node at a time, and look at the other afterwards.

What I would do differently

Checked against Keepalived 2.4, systemd and inotify-tools 4

Keepalived at the time was the version RHEL 7 shipped; the Source material does not record which. Current upstream is 2.4.3, and CentOS Stream 9 and 10, the upstream of RHEL 9 and 10, carry 2.2.8.

As builtToday
vrrp_script with script, interval, fall, rise, no weightKeywords unchanged. weight defaults to 0, and with weight 0 a failing script puts the instance into the fault state
Process check with killall -0Keepalived has vrrp_track_process, which follows processes through the kernel and needs no script; the manual has a section on why not to use pgrep, pidof or killall
No enable_script_security, no script_userScripts run as the user keepalived_script if it exists, otherwise as the user Keepalived runs as. enable_script_security refuses root scripts whose path a non-root user can write
auth_type PASSStill parsed. The manual calls it non-compliant, since authentication was removed from VRRP by RFC 3768 in 2004, and to be avoided "except when using unicast, where it can be helpful". Only the first eight characters of the password are used
No stronger authentication availableKeepalived 2.4.0 (June 2026) added auth_hmac, an HMAC-SHA256 trailer, meant especially for unicast. It is not in the 2.2.8 that CentOS Stream 9 and 10, the upstream of RHEL 9 and 10, carry
IPv4 and IPv6 address in one virtual_ipaddress blockThe manual says all addresses in virtual_ipaddress must be of the same family; a mixture belongs in virtual_ipaddress_excluded. How a 2.x release reacts to the block as built was not tested
unicast_src_ip, unicast_peerUnchanged; check_unicast_src and TTL checks per peer were added
No nopreemptUnchanged; for nopreempt the initial state must not be MASTER
No vrrp_versionDefault is 2, but IPv6 instances use version 3
net.ipv4.ip_nonlocal_bind, net.ipv6.ip_nonlocal_bindBoth still exist
Units with PIDFile= and no Type=All options are still valid. systemd recommends PIDFile= only for Type=forking and says PID files should be avoided in new units
After = network.targetnetwork-online.target exists for units that need a configured network
inotify-tools from the distributionUpstream is active (4.26), but the package is not in RHEL 10; it comes from EPEL
rsync3.4 is in RHEL 10. I did not re-check the options or what current OpenSSH thinks of the RSA key

The one point the manual forbids in this configuration is the mixed address block: put the IPv6 address into virtual_ipaddress_excluded. Whether a 2.x release refuses the block as built was not tested.

← solutionz