Vault 04 - Network design and firewall flows
Vault Solution · Previous: High-level design · Next: Virtual machines and OS build
Each environment has three small networks, and each network has one job. This article lists the networks, every flow that was permitted between them and to the outside, and the two places where each flow was enforced.
Three networks per environment
flowchart TB subgraph lbnet["Load-balancer network 10.10.2.8/29"] vip(["prod-vault 10.10.2.14"]) lb1["lb1 10.10.2.10"] lb2["lb2 10.10.2.11"] end subgraph svcnet["Service network 10.10.1.32/28"] n1["node1 .34"] n2["node2 .35"] n3["node3 .36"] n4["node4 .37"] n5["node5 .38"] end subgraph admnet["Admin network 10.10.3.16/29"] b["bastion1 10.10.3.18"] end lbnet -- "TCP 8200" --> svcnet admnet -- "SSH" --> lbnet admnet -- "SSH" --> svcnet
| Network | PROD | NONPROD | COMMON | Holds |
|---|---|---|---|---|
| Service | 10.10.1.32/28 | 10.20.1.32/28 | 10.30.1.32/28 | The Vault nodes |
| Load balancer | 10.10.2.8/29 | 10.20.2.8/29 | 10.30.2.8/29 | Two load-balancer nodes and the virtual address |
| Admin | 10.10.3.16/29 | 10.20.3.16/29 | 10.30.3.16/29 | The bastion host |
All three are software-defined networks of the private cloud, stretched over both availability zones, so a machine keeps its address when the platform restarts it in the other zone.
Addresses
| Host | PROD | NONPROD | COMMON |
|---|---|---|---|
<env>-vault-node1 | 10.10.1.34 | 10.20.1.34 | 10.30.1.34 |
<env>-vault-node2 | 10.10.1.35 | 10.20.1.35 | 10.30.1.35 |
<env>-vault-node3 | 10.10.1.36 | 10.20.1.36 | 10.30.1.36 |
<env>-vault-node4 | 10.10.1.37 | 10.20.1.37 | none |
<env>-vault-node5 | 10.10.1.38 | 10.20.1.38 | none |
<env>-vault-lb1 | 10.10.2.10 | 10.20.2.10 | 10.30.2.10 |
<env>-vault-lb2 | 10.10.2.11 | 10.20.2.11 | 10.30.2.11 |
<env>-vault, the virtual address | 10.10.2.14 | 10.20.2.14 | 10.30.2.14 |
<env>-vault-bastion1 | 10.10.3.18 | 10.20.3.18 | 10.30.3.18 |
| Public endpoint | 198.51.100.32 | 203.0.113.22 | none |
Flows inside an environment
| From | To | Port | Purpose |
|---|---|---|---|
| API clients | Virtual address | TCP 443 | Vault API |
| Load-balancer nodes | Vault nodes | TCP 8200 | Vault API, passed through, and the health check |
| Vault nodes | Vault nodes | TCP 8200 | Joining the cluster |
| Vault nodes | Vault nodes | TCP 8201 | Raft replication and request forwarding |
| Load-balancer node | Load-balancer node | VRRP, IP protocol 112 | Keepalived |
| Bastion host | Load-balancer and Vault nodes | TCP 22 | Administration |
Every one of these segments is encrypted. Clients and the health check use TLS with the node certificates; between Vault nodes, port 8201 carries mutually authenticated TLS with certificates the cluster generates for itself.
Flows that leave an environment
| From | To | Port | Purpose |
|---|---|---|---|
| PROD and NONPROD Vault nodes | Virtual address of COMMON, 10.30.2.14 | TCP 443 | Transit auto-unseal |
| All Vault nodes | Audit collector of the security operations centre | TCP 50543 | Vault audit device |
| All servers | Central log collector, 10.40.2.11 | TCP 50515 | Operating-system security events |
| All servers | DNS and NTP servers of the private cloud | 53, 123 | Name resolution and time |
| All servers | Package repository, 10.40.2.10 | TCP 80 | Operating-system packages |
| Vault nodes | Forward proxy | TCP 8084 | The vendor's package repository on the Internet |
| Vault nodes | Object storage, 10.40.6.10 and 10.40.6.11 | TCP 443 | Snapshots and rotated audit logs |
Monitoring network 10.40.1.16/28 | All servers | TCP 10050 | Zabbix agent, polled |
| All servers | Monitoring network | TCP 10051 | Zabbix agent, active checks |
Monitoring probe, 10.43.1.8 | Virtual address | TCP 443 | Availability and certificate expiry |
Administrator access networks 10.41.1.16/28, 10.41.1.32/28 | Bastion host | TCP 22 | Administrators, after the organisation's access gateway |
Clients from the Internet
The orchestration platform ran in a public cloud, outside the organisation's network. The only approved way in from the Internet was the private cloud's reverse-proxy service, a pair of load balancers with a public address in front and a source-translated address behind.
flowchart LR subgraph inet["Internet"] o["orchestrator<br/>192.0.2.24, 192.0.2.34"] end subgraph rp["Reverse-proxy area"] r["Reverse proxy<br/>public 198.51.100.32<br/>TLS passthrough"] end subgraph env["PROD"] l["prod-vault 10.10.2.14<br/>HAProxy, TLS passthrough"] a["Active Vault node"] end o -- "TCP 443" --> r r -- "TCP 443 from 10.42.3.8" --> l l -- "TCP 8200" --> a
Neither proxy terminates TLS. The orchestrator's TLS session ends on the Vault node, which is why the node certificates carry the public name and address as well; see TLS, DNS and certificates. The reverse proxy admits only the orchestrator's two outbound addresses, and behind it the host firewall of the load-balancer nodes admits only the reverse proxy.
| Hop | Source filter |
|---|---|
| Reverse proxy, public side | The two outbound addresses of the orchestrator's environment |
| Load-balancer node, virtual address | The reverse proxy's source-translated address 10.42.3.8, and the two addresses 10.42.1.10 and 10.42.1.11 it probes from |
| Vault node, port 8200 | The load-balancer network and the service network |
COMMON has no public endpoint and no API clients. Its virtual address accepts the service networks of PROD and NONPROD and the monitoring probe, and nothing else.
Where the rules live
Each flow was permitted twice, by two different teams with two different tools.
| Layer | Enforced by | Granularity | Managed through |
|---|---|---|---|
| Between networks | The distributed firewall of the private cloud | Network or host objects, named services | Change requests in the firewall management portal |
| On each host | firewalld, zone public, rich rules | Source, destination and port per rule | firewall-cmd, by the Vault team |
The platform firewall permitted the platform's own standard services for every project by default. Everything else in the tables above was requested rule by rule, with objects named after what they are, for example NET_PROD_VAULT_SERVICE to IP_COMMON_VAULT_LB_VIP, service HTTPS.
On the hosts, the predefined ssh, cockpit and dhcpv6-client services were removed from the zone, so that nothing is open that is not written down as a rich rule with a source and a destination.
$ firewall-cmd --permanent --remove-service=dhcpv6-client $ firewall-cmd --permanent --remove-service=cockpit $ firewall-cmd --permanent --remove-service=ssh $ firewall-cmd --permanent --zone=public --add-rich-rule='rule family=ipv4 source address=10.10.2.8/29 destination address=10.10.1.34/32 port port=8200 protocol=tcp accept' $ firewall-cmd --reload $ firewall-cmd --zone=public --list-rich-rules
The complete rule sets are Config documents:
General firewalld usage, zones and rich-rule syntax are in my older note firewalld notes and are not repeated here.
What the notes got wrong
Writing rules by hand on eight servers per environment produced two small inconsistencies, both visible in the Config documents: a monitoring rule on the bastion host whose command names the wrong destination network, and a rule on the first load balancer that opens the Zabbix port under a comment that says HTTPS while its twin on the second opens 443. Neither opened anything dangerous. Both are the argument for generating host firewalls from one description of the flows, which is what the tables in this article would be.
Reading it today
Vault's ports are unchanged in 2.1: 8200 for the API and 8201 for the cluster. Two things are new on the path in. Vault 2.0 rejects request paths that are not clean, with //, /./ or /../ in them, which matters to any proxy or client that builds URLs carelessly. And the listener can redact the version, cluster name and addresses from unauthenticated endpoints such as the health check, which is worth switching on for an endpoint reachable from the Internet.