Email 14 - Accounts, clients and operations
Email Solution · Previous: Webmail
The last part is about living with the system: what has to be done in Active Directory before a device can send its first alert, what is typed into that device, how the servers are watched and logged, in which order a site was built, and what was left unfinished. It closes with the list of contradictions I found when I read my own documents again.
How an account gets mail
Nothing about a mail user is stored on the mail servers. An account exists in Active Directory, and membership in groups decides what it may do. The lookups behind this are in Active Directory integration; here is the procedure as an operator sees it.
First the account. A device, a server or an application that sends mail gets its own account, named after the service with the suffix _mail: gitlab_mail for GitLab, sensu_mail for the monitoring. A person who only reads mail in the webmail has a normal user account.
| Kind of account | Created under | Used for |
|---|---|---|
| Device or software component | OU=msx_SVC,OU=Users_SVC,DC=ad,DC=example,DC=net | SMTP login of the device |
| User | OU=Users_STD,DC=ad,DC=example,DC=net | Webmail, to read mail; the design does not authorize these accounts to send |
The login name is the user principal name, account@ad.example.net in site 1 and account@ad-dc2.example.net in site 2, and it is also the mail address. There is nothing else to set: the address is the account.
Then the groups. Each permission has one group per organisation, so that it stays visible who let whom in, and one role group that contains them. Postfix and Dovecot ask only for the role group and follow the nesting.
| To allow | Add the account to one of | Role group that the servers check |
|---|---|---|
| Activating the address: sending to addresses of the own mail domains | SMTP_ORG, SMTP_PARTNER1, SMTP_PARTNER2 | SMTP_ACCESS |
| Sending to external domains | ESMTP_ORG, ESMTP_PARTNER1, ESMTP_PARTNER2 | ESMTP_ACCESS |
| Access to the webmail | IMAP_ORG, IMAP_PARTNER1, IMAP_PARTNER2 | IMAP_ACCESS |
The per-organisation groups live in OU=RBAC_GROUPS,OU=RBAC, the role groups in OU=RBAC_ROLES,OU=RBAC. These paths, like the two in the first table, are the notation of the design: the archived lookup files of site 2 search under DC=ad-dc2,DC=example,DC=net and name the role groups in OU=APPS,OU=RBAC_ROLES,OU=RBAC; see Active Directory integration. A device account that is in one SMTP_ group and nothing else is meant to alert the operators' mailboxes and not to send a single message to the internet. A notification that must reach a partner's mailbox outside needs the ESMTP_ group too, and then the content filter decides whether it leaves encrypted; see Automatic email encryption. Distribution lists are groups as well, under OU=DISTRIBUTION_GROUPS.
A change of membership needs nothing on the mail servers. There is no map to rebuild and no service to reload, because every check is an LDAP query at the time of the connection.
Client configuration
Clients are given the virtual address of the internal pair. The design does not forbid a node's own address: it says SMTP can also be used on the addresses that are not the virtual one, and its firewall rule for clients lists both nodes beside the virtual address. The design gives two ways in and calls the first one preferred.
| Parameter | SMTPS, preferred | SMTP with TLS |
|---|---|---|
| Port | TCP 465 | TCP 25 |
| Connection security | SSL/TLS from the first byte | STARTTLS |
| Authentication methods | PLAIN, LOGIN | PLAIN, LOGIN |
| User name | the account's address | the account's address |
| Password | the account's password in Active Directory | the account's password in Active Directory |
| Sender address | the account's address | the account's address |
The sender address is meant to be the account's own address. The design intended Postfix to compare the two and reject a message whose MAIL FROM belongs to somebody else. As archived, the order of the sender restrictions defeats that check: the group lookup permits every member of SMTP_ACCESS before the comparison is reached, so an authenticated account can use the address of any other account in that group. See Internal servers: Postfix.
The values that differ per site are these.
| Parameter | Site 1 (DC1) | Site 2 (DC2) |
|---|---|---|
| SMTP server, name | dc1-s-xcmsx001.adm.example.net | dc2-s-xcmsx001.adm.example.net |
| SMTP server, IPv4 | 10.11.19.33 | 10.12.19.33 |
| SMTP server, IPv6 | 2001:db8:a1:b6f::f:1 | 2001:db8:a2:b6f::f:1 |
| User name and sender address | account@ad.example.net | account@ad-dc2.example.net |
Two remarks on this table. The client tables of the design give 10.12.16.33 as the IPv4 address for site 2; the Keepalived configuration, the install notes and the design's own chapter on the virtual address all say 10.12.19.33, and I take the client table for a typing error. And the site 1 values carry a note in the design that they apply only after the new concept is implemented there: in October 2019 site 1 still ran the first concept, where clients were given the name of the single internal server. Its address was 10.11.19.33, the one the design then assigns to the virtual address, so a client configured by address would not have noticed the change.
The design does not say why SMTPS is the preferred one. The two ports are not equal on the server side, and not in the way the word "preferred" suggests: the listener on 465 carries its own, shorter recipient check. That is the first item under known issues below, and the server side of both ports is in Internal servers: Postfix.
Clients with neither authentication nor TLS
Requirement FR4 says that some clients, mostly specialised hardware appliances, support neither SMTP authentication nor TLS, and that the server "will be ready for such clients in the form of setting exceptions". I checked what the configuration really does with such a client. The relevant lines of main.cf on the internal pair:
mynetworks = 10.12.19.34/32, 10.12.19.35/32, 10.12.19.43/32 smtpd_client_restrictions= permit_sasl_authenticated, permit_mynetworks, reject smtpd_tls_security_level = encrypt smtpd_tls_auth_only = yes
The mechanism for an exception is mynetworks: an address listed there passes the client, sender and recipient restrictions without logging in, as long as the mail is for the own domain. For external recipients a sender address of the own domain still has to be in ESMTP_ACCESS, because both check_sender_access lines of the recipient restrictions stand before permit_mynetworks. As archived, the list holds three addresses and all three are mail servers, the two relays and the mailbox server. No appliance is in it. A client that does not authenticate and is not in the list is rejected.
And the list alone would not be enough. smtpd_tls_security_level = encrypt makes TLS mandatory on port 25 for everybody, listed or not, so a client that cannot do STARTTLS is refused at MAIL FROM. A real exception for a client without TLS would have needed a second listener in master.cf with the security level lowered to may and access limited to named addresses. The Source material holds no such listener and no record of an appliance that needed one. So the honest statement is: FR4 was provided for in the design, the exception for authentication is one line away, the exception for TLS was not built. The whole file is a Config document: Postfix main.cf, internal servers.
Monitoring
The email servers sit in the same infrastructure services as every other server of the management infrastructure. The design has one diagram for that, redrawn here for site 2.
flowchart LR adm["Administrators and operators"] -- "SSH, TCP 22" --> sshd subgraph vm["Email server, virtual machine"] sshd["sshd"] sssd["sssd"] res["DNS resolver"] ntpd["ntpd"] aud["auditd"] rsys["rsyslogd"] sen["sensu-client"] etck["etckeeper"] end aud --> rsys sssd -- "AD, LDAP" --> ad["Active Directory, DC2-A-VCAD001 and 002"] res -- "DNS, 53" --> ibx["Infoblox, site 1"] ntpd -- "NTP, 123" --> ibx rsys -- "syslog, TCP and UDP 514" --> sys["Central syslog, DC2-S-XCSYS001"] sen -- "Sensu messages, TCP 5672" --> rmq["RabbitMQ, DC2-A and DC2-B-VCRMQ001"] sen -- "metrics, TCP 80" --> gra["Graphite, DC2-S-XCGRA001"] vm -- "ICMP" --> sns["Sensu servers"] etck -- "git over SSH, TCP 22" --> git["GitLab"]
Monitoring is Sensu. A client on each server runs the checks and publishes the results to RabbitMQ; the machine pings the Sensu servers, which is the direction the design gives for ICMP; metrics go to Graphite. Every server gets the base template.
| Level | Check | Condition |
|---|---|---|
| System | Processes sshd, sssd, ntpd, qemu-ga, rsyslogd, auditd, crond | Must be running |
| System | All standard mount points | Usage below the threshold |
On top of the template each role has its own short list.
| Role | Servers in site 2 | Additional checks |
|---|---|---|
| Internal pair | DC2-A-VCMSX001, DC2-B-VCMSX001 | Processes postfix and dovecot must be running |
| Mailbox and webmail | DC2-A-VCMSX002 | Processes postfix, dovecot, httpd, mariadb must be running; usage of /data below the threshold |
| Relays | DC2-A-VCMSR001, DC2-B-VCMSR001 | Process postfix must be running |
The site 1 servers have the same lists. The design gives no thresholds and no check definitions, only these tables.
Every check here asks whether a process exists. None asks whether mail gets through: no check of the queue length, no test message from end to end, no look at the expiry date of a certificate, and nothing watches which node holds the virtual address or whether the two key synchronization services are alive. The monitoring system itself sends its alerts through this mail system, with the account sensu_mail, which is on the list of senders whose mail is never encrypted. When the mail system is the thing that is broken, the alert about it travels through the broken thing.
Logging
| What | Where it is written | Rotation |
|---|---|---|
| Postfix, all servers | syslog, facility mail; the SMTPS listener logs as postfix/smtps | by the system |
| Dovecot | syslog, facility mail | by the system |
| Roundcube | syslog, facility mail, identity roundcube | by the system |
| Content filter | its own file, /var/spool/postfix/bash-postfix-encrypt-filter/log/bash-postfix-encrypt-filter.log | daily, seven kept, copytruncate |
| Operating system audit | auditd | by the system |
rsyslogd forwards everything that reaches syslog to the central syslog server of the site, dc2-s-xcsys001.adm.example.net in site 2 and dc1-s-xcsys001.adm.example.net in site 1. The forwarding rule is part of the operating system template and is not in the archived configuration trees.
The filter is the exception. It is a shell script and writes its own log, one line per step with the mail ID in front, so that the path of one message through the filter can be followed with grep: entered the filter, sender or recipient on an exception list, certificate or key found, encrypted, handed back. That file stays on the node. It is not in syslog and therefore not on the central server, and since either node of the pair may have handled a message, a search means looking on both. The rotation is a Config document: logrotate rule for the filter. copytruncate is there because the script appends to the file by name on every run and nothing could be told to reopen it. For logrotate to be allowed into the Postfix spool at all, SELinux needed a file context and a local module; see Virtual machines and OS build.
Besides the log, the filter keeps a copy of every message, the original and the encrypted one, in its archive directory for one day. I read this as a troubleshooting aid, the Source material does not say what it was for; it also means that clear text of mail that left encrypted lies on the internal servers for a day. The script is in bash-postfix-encrypt-filter.sh.
Configuration tracking
etckeeper is on every server as part of the base build: /etc is a git repository, changes are committed, and the diagram of the design shows the repository pushed to the GitLab of the management infrastructure. The email system used it and did not configure it; the Source material has no etckeeper settings. What it did not cover is everything outside /etc: the filter script and the two synchronization scripts in /usr/local/bin, the exception lists and the recipients' keys under /var/spool/postfix. The design names a GitLab project of its own for the source code of the filter. Where the exception lists were kept besides the servers, the Source material does not say.
The order of installation
Chapter 7 of the design, "Installation notes", is four references to text files and one empty heading, that of the second relay, for site 2, and five times the sentence that the notes "will be added after implementation of new email concept" for site 1. The archive holds five files of notes, one per server of site 2. They are grouped by topic and not strictly by time (the certificate, which Postfix on the internal servers cannot start without, is the second to last section in each; the notes of the relays say their certificate is generated and currently not used), so the order below is the order of the notes with that caveat.
| Step | Internal pair | Mailbox and webmail | Relays |
|---|---|---|---|
| 1 | Install Postfix and Dovecot, stop them, set the stock configuration aside | The same | Install Postfix, stop it, set the stock configuration aside |
| 2 | vmail account, Dovecot configuration | /data/vmail, its SELinux label, vmail account, Dovecot | Postfix: main.cf, transport to the mailbox server, header_checks |
| 3 | Postfix: virtual address, relay hosts, transport | Postfix: relay through the pair, delivery to Dovecot | Start Postfix |
| 4 | Content filter: packages, account, script, directories, exception lists, master.cf, log rotation | firewalld: SSH, SMTP, 465, HTTPS, IMAP | firewalld: SSH, SMTP |
| 5 | Key synchronization: rsync, inotify-tools, SSH keys, two systemd units | SELinux: two modules, one boolean | Certificate request |
| 6 | firewalld: SSH, SMTP, 465 | Apache and PHP, MariaDB, Roundcube installer | Repositories |
| 7 | Keepalived: non-local bind, reboot, configuration | Certificate request | |
| 8 | SELinux: file context, three modules | Repositories | |
| 9 | Certificate request with the name of the virtual address | ||
| 10 | Repositories |
The notes do not record the order between the servers. The dependencies give one: the relays and the mailbox server first, because the internal pair needs somewhere to send; node A of the pair before node B; the SSH keys of the synchronization last, since each node needs the other's public key. The notes for node A and node B are the same file with a handful of differences: names, addresses, Keepalived state and priority, peer.
Known issues and limitations
Both design documents of the first concept end with a chapter of this name, and in both it is one sentence: "No known issues or limitations have been found." The final design has no such chapter. Reading everything again for this write-up, I found the following. Each item is in the Source material; none of them is in its list of issues.
- Port 465 skips the check for external sending. The
smtpslistener inmaster.cfreplaces the recipient restrictions ofmain.cfwithpermit_sasl_authenticated,reject. As I read the configuration, an account that is inSMTP_ACCESSand not inESMTP_ACCESSis held to internal recipients on port 25 and is not on port 465, the port the design recommends. The details are in Internal servers: Postfix. - Site 1 was not done. The final design of October 2019 describes the HA concept for both sites and states that it "is not yet deployed" in site 1. All configuration shown in this Solution is from site 2, the pre-production site. Production still ran one internal server and one relay.
mydomainon node B. Node A of the internal pair hasmydomain = ad-dc2.example.net, node B hasad.example.net, the domain of the other site. See Internal servers: Postfix.- The relays do not say the names DNS has for them. The design's records for site 2 are
mx1.ad-dc2.example.netandmx2.ad-dc2.example.net; the relays havemyhostnameset tomx3.ad.example.netandmx4.ad.example.net. See Relay servers. - Two spellings of the virtual name. The component table and the certificate table of the design call the clustered service
DC2-S-VCMSX001; the address chapter, the client tables and the certificate request in the install notes saydc2-s-xcmsx001. - IMAP is open. The design says IMAP is reachable from localhost only. On the mailbox server Dovecot listens on every IPv4 address (
listen = *) and the install notes add theimapservice to firewalld. - Unanswered questions in the firewall tables. The source column of the client rules reads "EMAIL clients / Any ???", question marks included.
- FR4 is half built, as described above.
- A recipient without a key gets no alert. When the filter has neither a certificate nor a PGP key for an external recipient who is not on an exception list, it replaces the body with a notice asking for a key. The original message is not delivered. For a system whose mail is alerts, that is a limitation an operator has to know.
- One mailbox server. Mail for the site's own domain ends on a single machine in datacenter A. If it is down, the pair queues.
- Site 2 resolves names and gets its time from site 1. The Infoblox appliances exist only there.
- LDAP over IPv4 only. The servers speak IPv6 to clients, but the design notes that Active Directory is "integrated only on IPv4 addresses due to problems in LDAP libraries".
- Health means "the process exists". Both Keepalived and Sensu test for a running process. A Postfix that is up and cannot reach Active Directory keeps the virtual address and stays green. See High availability and key synchronization.
What I would do differently
The list above is what was wrong then. This one is what has changed since, checked against current upstream documentation in October 2026.
- The base of the build is past its end of life. Maintenance of RHEL 7 ended on 2024-06-30. MariaDB 5.5, GnuPG 2.0, OpenSSL 1.0.2 and PHP 5 are no longer supported upstream, and Roundcube 1.1 is five minor lines behind. Postfix went from 2.10.1 to 3.11, Dovecot from 2.2 to 2.4, and Dovecot 2.4 does not accept a 2.2 configuration. A rebuild is a new build with this design as the specification, not an upgrade.
- Relay control in its own list. Postfix prefers relay permission rules in
smtpd_relay_restrictions, which is evaluated before the recipient list, and its default permits every authenticated client. I kept everything insmtpd_recipient_restrictionsand never set the relay list. That is not what lost the check for external sending on port 465: the-o smtpd_recipient_restrictionsoverride inmaster.cfdid that alone. Today the stockmaster.cfcalls that servicesubmissions, and I would write its options by hand and not take the template. - A health check that is not
killall. Keepalived has a nativevrrp_track_processand explains in its manual why not to usepgrep,pidoforkillall. It is still a test for a process. The password authentication between the nodes was removed from the VRRP standard in 2004; the manual calls it non-compliant and tolerable with unicast peers, which is what the pair used. - Time from chrony.
ntpdis not available since RHEL 8; the NTP protocol is implemented there bychronydonly. - Mail to the internet needs more than SPF now. Since February 2024 Gmail and Yahoo require SPF or DKIM and valid forward and reverse DNS from every sender, Gmail also TLS, and both require SPF, DKIM and DMARC from bulk senders. The design has an SPF record and neither DKIM nor DMARC, and relays that introduce themselves with names that differ from the DNS records of the design. I would not count on that being enough now.
- Monitoring that sends a message. This one needs no research. A check that submits a test message on 465 with a device account and finds it in a mailbox a minute later would have covered Active Directory, both listeners, the filter, the mailbox server and the certificates at once, and none of the process checks did.