Vault 14 - Monitoring and day-2 operations
Vault Solution · Previous: Raft snapshots, backup and restore
The design document asked for three kinds of monitoring: operational, security and performance. This last article says which of them the build delivered, how the routine work on a running cluster was done, and what moving this system to Vault 2.x would involve.
What watched the system
flowchart LR subgraph env["One environment"] s["Every server<br/>Zabbix agent"] vip(["Virtual address :443"]) v["Vault<br/>audit devices"] os["rsyslog, auditd"] end zp["Zabbix proxies<br/>monitoring network"] -- "10050, passive checks" --> s s -- "10051, active checks" --> zp probe["External probe"] -- "HTTPS, certificate" --> vip v -- "audit entries" --> soc["Security operations centre"] os -- "security events" --> soc scan["Daily scan client"] --> comp["Vulnerability and compliance scanner"] s --- scan
| Kind | Asked for in the design | Delivered by | Gap |
|---|---|---|---|
| Operational | Seal status, leadership changes, audit failures, request latency, replication | Zabbix agent on every server; an external probe on the virtual address | No sign that Vault's own metrics were collected |
| Security | Root token use, policy changes, new auth methods, failed logins, unusual hours | Audit trail to the security operations centre, see Audit logging and log shipping | None |
| Performance | CPU, memory, disk, storage latency, Raft last contact | Zabbix agent, for the host figures | Vault and Raft figures not collected |
Zabbix
Monitoring is a shared service of the organisation: a central Zabbix reached through proxies in a monitoring network, an agent on each server, and a self-service portal where a system's hosts are registered.
The agent came as the service's own installer package and was installed on every server under a dedicated account.
$ lvextend -L +2G /dev/vg00/lv_opt $ resize2fs /dev/mapper/vg00-lv_opt $ useradd -c "Zabbix agent" -m -d /opt/zabbix_agent zabbix $ tar -xf zabbix_agent_installer.unix.6.0_1.08.tar $ source install.sh -proxyip 10.40.1.28 -ignore_install_in_root
The installer needs the address of a proxy, and the service asked projects to spread themselves over its two proxies. Before any of that, the two flows have to exist in the platform firewall and on the host.
$ firewall-cmd --permanent --zone=public --add-rich-rule='rule family=ipv4 source address=10.40.1.16/28 destination address=10.30.1.32/28 port port=10050 protocol=tcp accept' $ firewall-cmd --reload $ nmap -p 10051 10.40.1.17
The monitoring page of the design is not among my sources. This is what the install notes and the firewall rules show.
| Check | From | What it tells |
|---|---|---|
| The agent's standard host items | Agent | The server is healthy and its filesystems are not filling up |
| HTTPS on the virtual address | External probe | A client can reach an active Vault node through the load balancer |
| Certificate on the virtual address | External probe | Days until the certificate the active node presents expires |
The probe on the virtual address is the most valuable single check, because it tests the whole chain a client depends on: address, load balancer, health check, active node, TLS. It answers "is Vault serving" and nothing about why not.
What the notes show no check for
Initialization and Transit auto-unseal showed that this design has dependencies which fail silently, because a running Vault does not use them. My notes show no check for any of them.
| Silent dependency | Fails when | A check that would catch it |
|---|---|---|
| The unseal token | It expires, three years after creation | Look the token up by accessor on COMMON; alert on remaining lifetime |
| COMMON being unsealed | A COMMON node restarts and waits for key shares | Health endpoint of each COMMON node, expecting 200 or 429 |
| Certificates on standby nodes | They expire; the probe only ever sees the active node's | Read the certificate from each node's port 8200 directly |
| Snapshots arriving | The agent stops, or its SecretID is revoked | Age of the newest object under snapshots/ in the bucket |
| Audit log shipping | The S3 key is rotated, or the script fails | Age of the newest object under logs/ |
| A sealed standby | A node restarts and cannot unseal | Health endpoint of every node; 503 is an alert even though the cluster still serves |
Each is one item with one threshold. Together they are the difference between a cluster that is up and a cluster that will come up again.
Vault's own telemetry was not configured. The server file has no telemetry block, so request latency, Raft commit time, leadership changes and audit-device failures, all named in the design, were visible only in the logs after the fact.
Routine checks
What an administrator looks at on a node, in this order.
$ export VAULT_ADDR='http://127.0.0.1:8200' $ vault status $ vault login -method=userpass username=admin01 $ vault operator raft list-peers $ vault operator raft autopilot state $ vault audit list -detailed $ journalctl -u vault --since "1 hour ago"
| Output | Healthy |
|---|---|
vault status | Sealed false, HA Mode active on exactly one node and standby on the others, Raft committed and applied index equal |
list-peers | Every node present, one leader, all voters |
autopilot state | Healthy: true, failure tolerance 2 for five nodes and 1 for three |
audit list | Both devices |
On a load-balancer node:
$ cat /var/run/keepalived_status $ ip -brief address show $ echo "show stat" | nc -U /var/lib/haproxy/stats | cut -d "," -f 1,2,18 | column -s, -t
Restarting a cluster without an outage
Kernel updates, certificate changes and configuration changes all end in restarting Vault on every node. The rule is one node at a time, standbys first, the leader last, and no next node before the previous one is back.
| Step | On | Command | Go on when |
|---|---|---|---|
| 1 | Any node | vault operator raft list-peers | The leader is known |
| 2 | A standby | systemctl restart vault | |
| 3 | The same standby | vault status | Sealed false |
| 4 | Any node | vault operator raft autopilot state | Healthy: true, all five voters |
| 5 | Repeat 2 to 4 for each standby | ||
| 6 | The leader | vault operator step-down | Another node is leader |
| 7 | The former leader | systemctl restart vault, then 3 and 4 |
Step 3 is where the dependency on COMMON shows: a PROD node that does not come back unsealed within seconds means COMMON, the path to it or the token, and the restart stops there with four healthy nodes instead of continuing to two. Before restarting anything in PROD, check that COMMON answers.
$ curl -s -o /dev/null -w '%{http_code}\n' https://common-vault.example.net/v1/sys/health
For COMMON itself the same order applies, with one more step per node: two share holders unseal it.
Stepping down first makes the leader change happen when the administrator chooses, with every other node already restarted, instead of when the process stops.
Recurring work
| Work | How often | Described in |
|---|---|---|
| Operating-system patches and reboots | Regularly | Above |
| Node certificates | Every two years, per environment | TLS, DNS and certificates |
| Unseal token | Before it expires | Initialization and Transit auto-unseal |
| AppRole SecretIDs | When a consumer asks, or after an incident | Authentication and policies |
| Administrator accounts | When the team changes | The same |
| Hardening re-check | With every benchmark revision; findings from the daily scan as they come | Linux hardening |
| Restore test | Should have been quarterly | Raft snapshots, backup and restore |
Vault versions
| When | Version | Where |
|---|---|---|
| 2022 | 1.11.1 | The first NONPROD build notes |
| Mid 2023 | 1.14.0, 1.14.1 | The PROD build and its initialization |
| Late 2023 | 1.14.4 | Seen in the audit log |
The notes record the versions and no upgrade procedure. The procedure that fits this build is a package update with the rolling restart above: update the package on a standby, restart it, check, go on, leader last.
$ yum update vault $ systemctl restart vault $ vault status
The order matters for upgrades in particular: the vendor's guidance is standbys first and the active node last.
Moving this system to Vault 2.x
Vault 1.14 was the last release under the MPL licence and is long out of support. This is what a move to the current release, 2.1.1, would have to deal with, taken from the release documentation and checked against the configuration in this Solution.
| Since | Change | Effect on this system |
|---|---|---|
| 1.15 | Licence changed to the Business Source License | A decision for the organisation, not a technical step. OpenBao is the fork that stayed on the MPL. |
| 1.19 | Health endpoint gained codes 474 and 530 | None; HAProxy accepts only 200 |
| 1.20 | disable_mlock must be set explicitly with Integrated Storage | None; it is set |
| 1.21 | A key written twice in an HCL block is an error, in configuration and in policies | Check every file; the ones published here have none |
| 1.21 | List values in allowed_parameters are matched element by element | None; the KV v2 policies have no parameter rules |
| 2.0 | sys/rekey and sys/generate-root need a Vault token as well as key shares | Rewrite the recovery runbook; extend admin-policy |
| 2.0 | Request paths that are not clean are rejected | Test both API clients; a doubled slash in a built URL now fails |
| 2.0.1 | Wildcards in rendered policy templates are refused | None; no templated policies |
| any | Audit file devices with an executable mode are refused | None; the default mode was used |
The configuration file should be cleaned up first.
| Remove before upgrading | Why |
|---|---|
disable_cache = "true" | A large performance cost for nothing |
raw_storage_endpoint = "true" | A privileged endpoint nobody needed |
disable_sealwrap, disable_printable_check | Enterprise-only, and undocumented |
tls_skip_verify in storage "raft" | Never a parameter |
The inline seal token | Move it to the environment or a file reference |
The upgrade guide for 2.x prescribes no intermediate versions and says to read the important changes of every release in between. Whether a direct jump from 1.14 to 2.1 is supported I could not verify. For a system holding other people's secrets I would not find out in production: restore the latest PROD snapshot into a scratch cluster on 1.14, upgrade that through one release at a time, and run both clients' test suites against it. That exercise is the missing restore test and the upgrade rehearsal in one.
The snapshot agent does not make the trip. It has to be replaced before the upgrade, not after.
What I would build the same way
- Three clusters, one of them only for unsealing. It kept people out of the restart path of production and cost three small machines.
- Integrated Storage. In three years there was no second product to patch, monitor or explain.
- TLS to the node. No proxy on the path ever held a key or saw a secret.
- A writer that cannot read. The most exposed system could not collect what it stored.
- Two audit devices and an outside team reading one of them.
- Named deviations. Every hardening rule that was not applied has a line saying why.
And differently: the base build from code and not from notes, a periodic unseal token outside the configuration file, the PROXY protocol so that the audit log names the client, policies that spell out data and metadata, a check for every silent dependency, and a restore test on the calendar.