Vault 08 - Load balancer
Vault Solution · Previous: TLS, DNS and certificates · Next: Vault server configuration
A Vault cluster has one active node, and which one changes. Clients should not have to know. In front of each cluster stand two small servers that give it one name and one address and send every connection to whichever node is active at that moment.
Why not the platform's load balancer
The private cloud has a load balancer built into its network virtualisation. The platform team itself advised against it for this kind of use and offered a pattern instead: two virtual machines with HAProxy, made highly available with Keepalived. That is two well-known open-source daemons and no new product, which is what the maintainability requirement in Use cases and requirements asked for.
Two jobs, two daemons
flowchart TB c["API client"] -- "TCP 443" --> vip(["Virtual address 10.10.2.14"]) subgraph lb1["prod-vault-lb1, MASTER, priority 100"] k1["Keepalived"] h1["HAProxy"] k1 -. "killall -0 haproxy, every second" .-> h1 end subgraph lb2["prod-vault-lb2, BACKUP, priority 50"] k2["Keepalived"] h2["HAProxy"] k2 -. "killall -0 haproxy, every second" .-> h2 end vip --- h1 vip -. "moves here on failure" .- h2 k1 <-- "VRRP, unicast, every second" --> k2 h1 -- "GET /v1/sys/health over TLS, every 5 s" --> n1["node1: 429 standby"] h1 --> n2["node2: 429 standby"] h1 == "TCP 8200" ==> n3["node3: 200 active"] h1 --> n4["node4: 429 standby"] h1 --> n5["node5: 429 standby"]
| Daemon | Question it answers | How |
|---|---|---|
| Keepalived | Which of the two load-balancer nodes holds the virtual address? | VRRP between the two nodes; the node with the higher priority and a running HAProxy wins |
| HAProxy | Which Vault node gets the connection? | An HTTP health check against every node; only a node that answers 200 is in rotation |
HAProxy
The configuration is a Config document: haproxy.cfg. Three choices in it carry the design.
TCP mode. HAProxy does not look inside the connection. It accepts TCP on port 443 of the virtual address and opens TCP to port 8200 of a Vault node. TLS is between the client and Vault, so the load balancer holds no certificate and no key, and has nothing worth stealing. The price is that it cannot add a header saying who the client was; see the last section.
The health endpoint. Vault answers GET /v1/sys/health without authentication and puts the node's state into the HTTP status code.
| Code | State of the node | In rotation |
|---|---|---|
| 200 | Initialized, unsealed, active | Yes |
| 429 | Unsealed, standby | No |
| 472 | Disaster-recovery secondary, Enterprise | No |
| 473 | Performance standby, Enterprise | No |
| 501 | Not initialized | No |
| 503 | Sealed | No |
The backend expects exactly 200. At any moment four of five servers are "down" in HAProxy's view, and that is the healthy state.
backend prod-vault
option httpchk GET /v1/sys/health
http-check expect rstatus 200
server prod-vault-node1 prod-vault-node1.example.net:8200 check check-ssl verify none inter 5000The check is HTTP over TLS, the traffic is not inspected. check-ssl makes the health check a TLS client of its own, separate from the passed-through client connections.
After a leader election the new active node starts answering 200 and the old one stops. With a five-second interval and the default rise and fall counts, clients see a gap of some seconds in which no node is in rotation and connections are refused. API clients have to retry.
HAProxy runs in a chroot under /var/lib/haproxy and logs to rsyslog on the loopback address; see Audit logging and log shipping for where those lines go.
$ yum install haproxy $ mkdir -p /etc/haproxy-original $ cp -R /etc/haproxy/* /etc/haproxy-original/ $ vi /etc/haproxy/haproxy.cfg $ mkdir -p /var/lib/haproxy/dev $ restorecon -RvF /var/lib/haproxy/dev $ systemctl enable haproxy $ systemctl start haproxy
The state of the backend can be read from the stats socket at any time. The output looks like this.
$ echo "show stat" | nc -U /var/lib/haproxy/stats | cut -d "," -f 1,2,18 | column -s, -t
output 8 lines
# pxname svname status prod-vault-lb FRONTEND OPEN prod-vault prod-vault-node1 DOWN prod-vault prod-vault-node2 DOWN prod-vault prod-vault-node3 UP prod-vault prod-vault-node4 DOWN prod-vault prod-vault-node5 DOWN prod-vault BACKEND UP
Keepalived
The configuration is a Config document: keepalived.conf.
| Setting | Value | Why |
|---|---|---|
state, priority | MASTER and 100 on lb1, BACKUP and 50 on lb2 | lb1 holds the address whenever it can |
vrrp_script | killall -0 haproxy, every second, two failures to fail | The address must not stay on a node whose HAProxy has died |
unicast_src_ip, unicast_peer | The two nodes' own addresses | Unicast between two known peers needs nothing from the network underneath |
virtual_ipaddress | The cluster's virtual address | The one address clients use |
authentication | A shared password | Keeps a misconfigured third node out |
notify | A script that records each transition | Lets an operator ask a node what it is |
The node that does not hold the address still runs HAProxy, configured to listen on it. Binding an address the node does not have is forbidden by default, so the load-balancer nodes set one kernel parameter more than the others.
$ yum install keepalived $ vi /etc/sysctl.d/haproxy.conf $ sysctl -p /etc/sysctl.d/haproxy.conf
output 1 line
net.ipv4.ip_nonlocal_bind = 1
The notification script, keepalived_notify.sh, writes the last transition to a file, one line like the one below.
$ cat /var/run/keepalived_status
output 1 line
Tue Aug 8 10:12:44 CEST 2023 - INSTANCE haproxy_service has transitioned to the MASTER state with a priority of 100
The password for VRRP was generated on the spot and is the same on both nodes of a pair.
$ date | sha256sumSELinux
Both daemons run confined, and both needed something the stock policy does not allow.
| Need | Symptom | Fix |
|---|---|---|
| HAProxy connects to port 8200, which is not a web port in the policy | Backend connections are denied | setsebool -P haproxy_connect_any on |
| rsyslog creates the log socket inside HAProxy's chroot | No HAProxy log lines | A small policy module, rsyslog-haproxy.te |
| Keepalived runs the check and notify scripts | Denials for keepalived_t in the audit log | A module generated from those denials with audit2allow |
$ checkmodule -M -m /etc/selinux/rsyslog-haproxy.te -o /etc/selinux/rsyslog-haproxy.mod $ semodule_package -o /etc/selinux/rsyslog-haproxy.pp -m /etc/selinux/rsyslog-haproxy.mod $ semodule -i /etc/selinux/rsyslog-haproxy.pp $ grep keepalived /var/log/audit/audit.log | audit2allow -M keepalived-modify $ semodule -i keepalived-modify.pp $ semodule --list-modules=full | grep -E 'keepalived|rsyslog-haproxy'
The first module is written by hand and says exactly what it allows. The second is whatever the audit log contained when the command was run, which on the first node was one line, allow keepalived_t self:capability dac_override;. A generated module should always be read before it is loaded; how I work with local SELinux policy in general is in SELinux local policy maintenance workflow.
Where this configuration came from
The Keepalived file was not written for Vault. It is the one from an earlier mail system of mine, where the same pair of daemons kept a virtual address in front of Postfix and Dovecot, and one comment in the NONPROD and COMMON files still says so. The pattern is independent of what stands behind it: a process check, a priority, a notify script.
Reading it today
- Vault sees the load balancer, not the client. In TCP mode every request arrives from
10.10.2.10or10.10.2.11, and that is the remote address in every audit entry. Vault supports the PROXY protocol for exactly this:send-proxy-v2on HAProxy'sserverlines, andproxy_protocol_behaviorwithproxy_protocol_authorized_addrsin the Vault listener. Thex_forwarded_forsettings do not help in TCP mode, because no HTTP header is added. With a second passthrough proxy in front, as on the public path, that one has to send the PROXY header too. - Two new health codes. Since Vault 1.19 a standby that cannot reach the active node answers 474 and a node removed from the cluster answers 530. Neither is 200, so this backend treats both correctly without a change.
- The health check verifies nothing.
verify nonewas a shortcut; the CA was on the host. - The gap after a leader change can be shortened with
fastinterandrise 1on theserverlines. - Health endpoint
- TCP listener, PROXY protocol