HashiCorp Vault - secrets management for a service platform
Between 2021 and 2024 I designed, built and ran the secrets management of a multi-tenant network service platform. Customers of that platform hand over credentials: cloud account keys, VPN pre-shared keys, private keys of certificates, routing passwords. Before Vault, those values travelled through the same data model as everything else. Afterwards a customer portal wrote them into Vault, an orchestration platform read them out when it configured a device, and nothing in between ever held them.
This Solution is the whole of that system: three self-managed Vault clusters on Integrated Storage, one of which exists only to unseal the other two, behind HAProxy and Keepalived, on hardened Oracle Linux 8, with audit logs and snapshots shipped to object storage.
The system in one picture
flowchart LR subgraph consumers["API clients"] portal["service-portal<br/>writes secrets"] orch["orchestrator<br/>reads secrets"] end subgraph prod["PROD"] plb["HAProxy + Keepalived<br/>prod-vault:443"] pv["prod-vault<br/>5 nodes, Raft"] plb --> pv end subgraph nonprod["NONPROD"] nlb["HAProxy + Keepalived<br/>nonprod-vault:443"] nv["nonprod-vault<br/>5 nodes, Raft"] nlb --> nv end subgraph common["COMMON"] clb["HAProxy + Keepalived<br/>common-vault:443"] cv["common-vault<br/>3 nodes, Raft"] clb --> cv end portal --> plb orch --> plb portal -. test and dev .-> nlb orch -. test and dev .-> nlb pv -- "Transit auto-unseal" --> clb nv -- "Transit auto-unseal" --> clb s3[("S3 object storage<br/>snapshots, audit logs")] pv --> s3 nv --> s3 cv --> s3
The fictional environment
Every article and every configuration file uses the same names and addresses.
| Environment | Purpose | Vault nodes | Service network | Load-balancer network | Admin network | Client endpoint |
|---|---|---|---|---|---|---|
| PROD | Production secrets | 5 | 10.10.1.32/28 | 10.10.2.8/29 | 10.10.3.16/29 | prod-vault.example.net, 10.10.2.14 |
| NONPROD | Test and development secrets | 5 | 10.20.1.32/28 | 10.20.2.8/29 | 10.20.3.16/29 | nonprod-vault.example.net, 10.20.2.14 |
| COMMON | Transit auto-unseal for the other two | 3 | 10.30.1.32/28 | 10.30.2.8/29 | 10.30.3.16/29 | common-vault.example.net, 10.30.2.14 |
Hosts are named <env>-vault-node1 to node5, <env>-vault-lb1 and lb2, and <env>-vault-bastion1, all under example.net. Shared services (DNS, NTP, package repository, monitoring, log collector, object storage) sit in 10.40.0.0/16, administrators and API clients arrive from 10.41.0.0/16 and above, and addresses that were public in reality are taken from the documentation ranges of RFC 5737.
Articles
Read in this order; it is the order in which the system was built.
| # | Article | What it covers |
|---|---|---|
| 1 | Use cases and requirements | Who stores what, who reads it, and what the system had to guarantee |
| 2 | Secret model | The KV layout, and the move from KV v1 paths to KV v2 objects |
| 3 | High-level design | Three clusters, Raft, two availability zones, and what fails when |
| 4 | Network design and firewall flows | Three networks per environment and every permitted flow |
| 5 | Virtual machines and OS build | Sizing, the cloud template, disks, and the base build of a server |
| 6 | Linux hardening | CIS benchmark, kernel, SSH, SELinux, and the documented deviations |
| 7 | TLS, DNS and certificates | Names, subject alternative names, issuance and renewal |
| 8 | Load balancer | HAProxy in TCP mode, Keepalived, and how the active node is found |
| 9 | Vault server configuration | Installing Vault and reading vault.hcl block by block |
| 10 | Initialization and Transit auto-unseal | Bringing up COMMON first, then the clusters it unseals |
| 11 | Authentication and policies | AppRole, userpass, and a writer that cannot read |
| 12 | Audit logging and log shipping | Two audit devices, rsyslog, logrotate, and logs to object storage |
| 13 | Raft snapshots, backup and restore | The snapshot agent, the rsync detour, and restore |
| 14 | Monitoring and day-2 operations | Zabbix, security alerts, renewals, and moving to Vault 2.x |
Configuration files
Each file is a document of its own: what it is, where it lives, the file with its comments, what differs per node or environment, and what has changed since.
Vault
| File | Explained in |
|---|---|
| vault.hcl | 9, 10 |
| snapshot.json | 13 |
| vault-snapshot.service | 13 |
Vault policies
| File | Explained in |
|---|---|
| admin-policy.hcl | 11 |
| orchestrator-policy-kv2.hcl | 11 |
| orchestrator-ui-policy-kv2.hcl | 11 |
| service-portal-policy-kv2.hcl | 11 |
| snapshots-policy.hcl | 11, 13 |
| autounseal-prod-vault.hcl | 10, 11 |
Load balancer
Linux build and hardening
| File | Explained in |
|---|---|
| Cloud template | 5 |
| /etc/hosts | 5, 7 |
| sysctl.conf | 6 |
| sshd_config | 6 |
| CIS hardening playbook | 6 |
| firewalld rules, Vault node | 4 |
| firewalld rules, load balancer | 4 |
| firewalld rules, bastion host | 4 |
Logging
| File | Explained in |
|---|---|
| rsyslog.conf, Vault node and bastion | 12 |
| rsyslog.conf, load balancer | 12 |
| logrotate.conf | 12 |
| logrotate.d/vault | 12 |
| logrotate.d/haproxy | 12 |
| auditd-logrotate | 12 |
| vault-logs-s3-sync.sh | 12 |
| crontab | 12 |
Software
| Component | Version as built | Role |
|---|---|---|
| HashiCorp Vault, Community edition | 1.11.1 at the first build, 1.14.x from 2023 | Secrets management |
| vault_raft_snapshot_agent | 0.3.1 | Periodic Raft snapshots |
| Oracle Linux | 8 | Operating system of every server |
| HAProxy | 1.8, from the distribution | TCP load balancer in front of each cluster |
| Keepalived | from the distribution | Virtual address between two load-balancer nodes |
| Zabbix agent | 6.0 | Monitoring |
| Ansible with the RHEL8-CIS role | 2.9 | Hardening |
What I would do differently
The articles say this where it belongs. The short list:
- The unseal token had a fixed three-year lifetime. It was created in July 2023; a cluster restarted after July 2026 with that token would stay sealed. See article 10.
- The unseal token sat inline in
vault.hcl, in a file with mode0644. See vault.hcl. - The audit log shows the load balancer as the client of every request, because HAProxy passes TCP through without the PROXY protocol. See article 8.
- The KV v2 policies grant more than they say, because one wildcard stands where
dataormetadatashould. See article 11. - Two availability zones cannot give a five-node cluster a quorum on their own. The design relied on the platform restarting virtual machines in the other zone. See article 3.