LINUXOR.SK ... open source notes ...

Vault 14 - Monitoring and day-2 operations

category: solutionz · date: 2024-12-31 · updated: 2026-10-02 · author: LALA

Vault Solution · Previous: Raft snapshots, backup and restore

The design document asked for three kinds of monitoring: operational, security and performance. This last article says which of them the build delivered, how the routine work on a running cluster was done, and what moving this system to Vault 2.x would involve.

What watched the system

mermaid
flowchart LR
  subgraph env["One environment"]
    s["Every server<br/>Zabbix agent"]
    vip(["Virtual address :443"])
    v["Vault<br/>audit devices"]
    os["rsyslog, auditd"]
  end
  zp["Zabbix proxies<br/>monitoring network"] -- "10050, passive checks" --> s
  s -- "10051, active checks" --> zp
  probe["External probe"] -- "HTTPS, certificate" --> vip
  v -- "audit entries" --> soc["Security operations centre"]
  os -- "security events" --> soc
  scan["Daily scan client"] --> comp["Vulnerability and compliance scanner"]
  s --- scan
KindAsked for in the designDelivered byGap
OperationalSeal status, leadership changes, audit failures, request latency, replicationZabbix agent on every server; an external probe on the virtual addressNo sign that Vault's own metrics were collected
SecurityRoot token use, policy changes, new auth methods, failed logins, unusual hoursAudit trail to the security operations centre, see Audit logging and log shippingNone
PerformanceCPU, memory, disk, storage latency, Raft last contactZabbix agent, for the host figuresVault and Raft figures not collected

Zabbix

Monitoring is a shared service of the organisation: a central Zabbix reached through proxies in a monitoring network, an agent on each server, and a self-service portal where a system's hosts are registered.

The agent came as the service's own installer package and was installed on every server under a dedicated account.

bash
$ lvextend -L +2G /dev/vg00/lv_opt
$ resize2fs /dev/mapper/vg00-lv_opt
$ useradd -c "Zabbix agent" -m -d /opt/zabbix_agent zabbix
$ tar -xf zabbix_agent_installer.unix.6.0_1.08.tar
$ source install.sh -proxyip 10.40.1.28 -ignore_install_in_root

The installer needs the address of a proxy, and the service asked projects to spread themselves over its two proxies. Before any of that, the two flows have to exist in the platform firewall and on the host.

bash
$ firewall-cmd --permanent --zone=public --add-rich-rule='rule family=ipv4 source address=10.40.1.16/28 destination address=10.30.1.32/28 port port=10050 protocol=tcp accept'
$ firewall-cmd --reload
$ nmap -p 10051 10.40.1.17

The monitoring page of the design is not among my sources. This is what the install notes and the firewall rules show.

CheckFromWhat it tells
The agent's standard host itemsAgentThe server is healthy and its filesystems are not filling up
HTTPS on the virtual addressExternal probeA client can reach an active Vault node through the load balancer
Certificate on the virtual addressExternal probeDays until the certificate the active node presents expires

The probe on the virtual address is the most valuable single check, because it tests the whole chain a client depends on: address, load balancer, health check, active node, TLS. It answers "is Vault serving" and nothing about why not.

What the notes show no check for

Initialization and Transit auto-unseal showed that this design has dependencies which fail silently, because a running Vault does not use them. My notes show no check for any of them.

Silent dependencyFails whenA check that would catch it
The unseal tokenIt expires, three years after creationLook the token up by accessor on COMMON; alert on remaining lifetime
COMMON being unsealedA COMMON node restarts and waits for key sharesHealth endpoint of each COMMON node, expecting 200 or 429
Certificates on standby nodesThey expire; the probe only ever sees the active node'sRead the certificate from each node's port 8200 directly
Snapshots arrivingThe agent stops, or its SecretID is revokedAge of the newest object under snapshots/ in the bucket
Audit log shippingThe S3 key is rotated, or the script failsAge of the newest object under logs/
A sealed standbyA node restarts and cannot unsealHealth endpoint of every node; 503 is an alert even though the cluster still serves

Each is one item with one threshold. Together they are the difference between a cluster that is up and a cluster that will come up again.

Vault's own telemetry was not configured. The server file has no telemetry block, so request latency, Raft commit time, leadership changes and audit-device failures, all named in the design, were visible only in the logs after the fact.

Routine checks

What an administrator looks at on a node, in this order.

bash
$ export VAULT_ADDR='http://127.0.0.1:8200'
$ vault status
$ vault login -method=userpass username=admin01
$ vault operator raft list-peers
$ vault operator raft autopilot state
$ vault audit list -detailed
$ journalctl -u vault --since "1 hour ago"
OutputHealthy
vault statusSealed false, HA Mode active on exactly one node and standby on the others, Raft committed and applied index equal
list-peersEvery node present, one leader, all voters
autopilot stateHealthy: true, failure tolerance 2 for five nodes and 1 for three
audit listBoth devices

On a load-balancer node:

bash
$ cat /var/run/keepalived_status
$ ip -brief address show
$ echo "show stat" | nc -U /var/lib/haproxy/stats | cut -d "," -f 1,2,18 | column -s, -t

Restarting a cluster without an outage

Kernel updates, certificate changes and configuration changes all end in restarting Vault on every node. The rule is one node at a time, standbys first, the leader last, and no next node before the previous one is back.

StepOnCommandGo on when
1Any nodevault operator raft list-peersThe leader is known
2A standbysystemctl restart vault
3The same standbyvault statusSealed false
4Any nodevault operator raft autopilot stateHealthy: true, all five voters
5Repeat 2 to 4 for each standby
6The leadervault operator step-downAnother node is leader
7The former leadersystemctl restart vault, then 3 and 4

Step 3 is where the dependency on COMMON shows: a PROD node that does not come back unsealed within seconds means COMMON, the path to it or the token, and the restart stops there with four healthy nodes instead of continuing to two. Before restarting anything in PROD, check that COMMON answers.

bash
$ curl -s -o /dev/null -w '%{http_code}\n' https://common-vault.example.net/v1/sys/health

For COMMON itself the same order applies, with one more step per node: two share holders unseal it.

Stepping down first makes the leader change happen when the administrator chooses, with every other node already restarted, instead of when the process stops.

Recurring work

WorkHow oftenDescribed in
Operating-system patches and rebootsRegularlyAbove
Node certificatesEvery two years, per environmentTLS, DNS and certificates
Unseal tokenBefore it expiresInitialization and Transit auto-unseal
AppRole SecretIDsWhen a consumer asks, or after an incidentAuthentication and policies
Administrator accountsWhen the team changesThe same
Hardening re-checkWith every benchmark revision; findings from the daily scan as they comeLinux hardening
Restore testShould have been quarterlyRaft snapshots, backup and restore

Vault versions

WhenVersionWhere
20221.11.1The first NONPROD build notes
Mid 20231.14.0, 1.14.1The PROD build and its initialization
Late 20231.14.4Seen in the audit log

The notes record the versions and no upgrade procedure. The procedure that fits this build is a package update with the rolling restart above: update the package on a standby, restart it, check, go on, leader last.

bash
$ yum update vault
$ systemctl restart vault
$ vault status

The order matters for upgrades in particular: the vendor's guidance is standbys first and the active node last.

Moving this system to Vault 2.x

Vault 1.14 was the last release under the MPL licence and is long out of support. This is what a move to the current release, 2.1.1, would have to deal with, taken from the release documentation and checked against the configuration in this Solution.

SinceChangeEffect on this system
1.15Licence changed to the Business Source LicenseA decision for the organisation, not a technical step. OpenBao is the fork that stayed on the MPL.
1.19Health endpoint gained codes 474 and 530None; HAProxy accepts only 200
1.20disable_mlock must be set explicitly with Integrated StorageNone; it is set
1.21A key written twice in an HCL block is an error, in configuration and in policiesCheck every file; the ones published here have none
1.21List values in allowed_parameters are matched element by elementNone; the KV v2 policies have no parameter rules
2.0sys/rekey and sys/generate-root need a Vault token as well as key sharesRewrite the recovery runbook; extend admin-policy
2.0Request paths that are not clean are rejectedTest both API clients; a doubled slash in a built URL now fails
2.0.1Wildcards in rendered policy templates are refusedNone; no templated policies
anyAudit file devices with an executable mode are refusedNone; the default mode was used

The configuration file should be cleaned up first.

Remove before upgradingWhy
disable_cache = "true"A large performance cost for nothing
raw_storage_endpoint = "true"A privileged endpoint nobody needed
disable_sealwrap, disable_printable_checkEnterprise-only, and undocumented
tls_skip_verify in storage "raft"Never a parameter
The inline seal tokenMove it to the environment or a file reference

The upgrade guide for 2.x prescribes no intermediate versions and says to read the important changes of every release in between. Whether a direct jump from 1.14 to 2.1 is supported I could not verify. For a system holding other people's secrets I would not find out in production: restore the latest PROD snapshot into a scratch cluster on 1.14, upgrade that through one release at a time, and run both clients' test suites against it. That exercise is the missing restore test and the upgrade rehearsal in one.

The snapshot agent does not make the trip. It has to be replaced before the upgrade, not after.

What I would build the same way

And differently: the base build from code and not from notes, a periodic unseal token outside the configuration file, the PROXY protocol so that the audit log names the client, policies that spell out data and metadata, a check for every silent dependency, and a restore test on the calendar.

← solutionz