Vault 07 - TLS, DNS and certificates
Vault Solution · Previous: Linux hardening · Next: Load balancer
Neither load balancer on the way to Vault terminates TLS. A client that connects to the public name, or to the virtual address, or straight to a node, ends its TLS session on a Vault node and checks that node's certificate. This article explains what that does to the names in a certificate, how certificates were issued, and what renewal looked like.
Names
| Record | PROD | NONPROD | COMMON | Visible |
|---|---|---|---|---|
| Bastion host | prod-vault-bastion1.example.net | nonprod-vault-bastion1.example.net | common-vault-bastion1.example.net | Internally |
| Load-balancer nodes | prod-vault-lb1, prod-vault-lb2 | nonprod-vault-lb1, nonprod-vault-lb2 | common-vault-lb1, common-vault-lb2 | Internally |
| Virtual address | prod-vault.example.net | nonprod-vault.example.net | common-vault.example.net | Internally |
| Vault nodes | prod-vault-node1 to node5 | nonprod-vault-node1 to node5 | common-vault-node1 to node3 | Internally |
| Public endpoint | prod-vault.pub.example.net | nonprod-vault.pub.example.net | none | On the Internet |
All internal names are A records in one internal zone that the platform's DNS servers answer. The two public names are in a separate, publicly visible zone and point at the reverse proxy described in Network design and firewall flows. Records were ordered from the DNS team with a spreadsheet per environment; nothing here registers itself.
The names a cluster needs to run are also in /etc/hosts on every server, so that a DNS outage cannot stop a quorum: /etc/hosts.
Which name a client checks
flowchart LR o["Client on the Internet<br/>connects to prod-vault.pub.example.net"] --> rp["Reverse proxy<br/>TCP passthrough"] i["Client inside<br/>connects to prod-vault.example.net"] --> lb["HAProxy<br/>TCP passthrough"] rp --> lb lb --> n["Active Vault node<br/>presents its own certificate"] p["Peer node, administrator<br/>connects to prod-vault-node3.example.net"] --> n
Any of the five nodes can be the active one, so every node certificate has to be valid for every name a client might have used to get there.
| Entry | Value | Needed by |
|---|---|---|
| Common name | The node's own name, prod-vault-node1.example.net | Convention |
DNS.1 | The node's own name | Peers joining the cluster, administrators, the health check |
DNS.2 | The virtual address name, prod-vault.example.net | Internal API clients, and the nodes themselves through api_addr |
DNS.3 | The public name, prod-vault.pub.example.net | API clients on the Internet |
IP.1 | The node's address | Clients that connect by address |
IP.2 | The virtual address, 10.10.2.14 | The same |
IP.3 | The public address, 198.51.100.32 | The same |
COMMON has no public endpoint and its certificates carry DNS.1 and DNS.2 only.
The certificate authority
All certificates were issued by the organisation's internal certificate authority, from its web-server template.
| Property | Value |
|---|---|
| Template | Web service: server authentication and client authentication |
| Validity | Two years |
| Key | RSA 4096; the authority requires at least 3072 |
| Subject | Country, state, locality, organisation and unit as the authority prescribes; the common name is the node |
| Challenge password | Not set |
The authority issues only for names that cannot be registered publicly or that the organisation owns, and only to the people responsible for the server. Its root is not in any public trust store. Every server and every API client therefore needs the authority's certificate installed, which for the clients was part of the integration work.
Issuing a certificate
The same steps for every node. First the working variable and the subject alternative names, which OpenSSL takes from its configuration file.
$ export FQDN=prod-vault-node1.example.net $ vi /etc/pki/tls/openssl.cnf
[ req ] req_extensions = req_ext [ req_ext ] subjectAltName = @alt_names [ alt_names ] DNS.1 = prod-vault-node1.example.net DNS.2 = prod-vault.example.net DNS.3 = prod-vault.pub.example.net IP.1 = 10.10.1.34 IP.2 = 10.10.2.14 IP.3 = 198.51.100.32
Then the key and the request, both checked before they leave the server.
$ openssl genrsa -out /opt/vault/tls/$FQDN-key.pem 4096 $ openssl req -new -out /opt/vault/tls/$FQDN.csr -key /opt/vault/tls/$FQDN-key.pem $ openssl req -text -noout -verify -in /opt/vault/tls/$FQDN.csr $ openssl rsa -in /opt/vault/tls/$FQDN-key.pem -check
The private key never leaves the node. The request is carried to a workstation in the organisation's directory domain, because the authority is a Microsoft certificate service and accepts requests only from an authorised domain account.
C:\> certreq -submit -attrib "CertificateTemplate:<WEB_SERVICE_TEMPLATE>" prod-vault-node1.example.net.csr prod-vault-node1.example.net.cer C:\> certreq -retrieve <REQUEST_ID> prod-vault-node1.example.net.cer
The certificate comes back to the node as $FQDN-cert.pem, with the authority's certificate beside it as $FQDN-ca.pem.
$ openssl x509 -in /opt/vault/tls/$FQDN-cert.pem -text -noout $ chown root:root /opt/vault/tls/*-cert.pem /opt/vault/tls/*-ca.pem $ chown root:vault /opt/vault/tls/*-key.pem $ chmod 0644 /opt/vault/tls/*-cert.pem /opt/vault/tls/*-ca.pem $ chmod 0640 /opt/vault/tls/*-key.pem $ cp /opt/vault/tls/*-ca.pem /etc/pki/ca-trust/source/anchors/ $ update-ca-trust $ restorecon -RvF /opt/vault/tls
File in /opt/vault/tls | Owner | Mode | Used as |
|---|---|---|---|
<fqdn>-key.pem | root:vault | 0640 | tls_key_file |
<fqdn>-cert.pem | root:root | 0644 | tls_cert_file |
<fqdn>-ca.pem | root:root | 0644 | tls_client_ca_file, and a system trust anchor |
<fqdn>.csr | root:root | 0644 | The request, kept beside the key |
Putting the authority into the system trust store is what lets one Vault node verify another when it joins the cluster, and lets a PROD node verify COMMON when it unseals. Neither retry_join nor the seal block names a CA file; see Vault server configuration.
Validity and renewal
| Environment | First certificates | Current certificates | Expire |
|---|---|---|---|
| NONPROD | 2022 | Issued 12 February 2024 | 11 February 2026 |
| PROD | 2023 | Issued 12 February 2024 | 11 February 2026 |
| COMMON | 14 June 2023 | The same | 13 June 2025 |
The February 2024 certificates were not a renewal forced by expiry. They were issued when the public endpoint came into use, to add the public name and address as DNS.3 and IP.3; the earlier certificates of PROD still had more than a year to run. That is the practical lesson of TLS passthrough: a new way in to the cluster means new certificates on every node.
Replacing a certificate is the issuing procedure again with the new file put in place, followed by a restart of Vault on that node, one node at a time, the active node last.
$ systemctl restart vault $ vault status $ vault operator raft list-peers
A restarted node of PROD or NONPROD unseals itself through COMMON and rejoins as a follower; the cluster never loses its majority. Restarting the active node last means one leader election instead of several.
Expiry was watched from outside: the monitoring probe that checks the virtual address also reads the certificate it is shown and raises an alert before it expires. With five nodes and one active, that probe sees one certificate at a time, so the expiry dates were also kept in the design document, per environment.
Reading it today
- A restart is not needed. Vault reloads the listener's certificate and key on
SIGHUP.systemctl reload vaultreplaces a certificate without sealing the node, which removes the dependency on COMMON from the procedure. - Two years is long. Public authorities are down to much shorter lifetimes, and internal ones follow. Nineteen manual certificate requests every two years were tolerable; the same every few months would not be. Vault's own PKI engine, or an ACME-capable internal authority, is the way out, and both need the automation this build did not have.
- The dates in the table have passed. COMMON's certificates ran out in June 2025 and the others in February 2026. A reader who finds a cluster like this one should check that first.
tls_skip_verifywas never needed and never used. Trust came from the system store, and that still works in 2.1. Naming the CA explicitly, withleader_ca_cert_fileinretry_joinandtls_ca_certin the seal block, is clearer and does not depend on what else the host trusts.- TCP listener
- Transit seal