LINUXOR.SK ... open source notes ...

Vault 13 - Raft snapshots, backup and restore

category: solutionz · date: 2024-12-31 · updated: 2026-10-02 · author: LALA

Vault Solution · Previous: Audit logging and log shipping · Next: Monitoring and day-2 operations

Five replicas are not a backup. Raft copies a mistake to every node as faithfully as it copies everything else. This article covers how the clusters were backed up with snapshots, the detour the first design took, and how a restore works.

What a snapshot is

A Raft snapshot is one file with the complete state of the cluster at a moment: every secret, every policy, every auth method and token, all still encrypted by Vault's own keys. It is taken from the active node through the API.

Two properties follow from "still encrypted". The file can be stored somewhere less trusted than Vault itself. And it is useless without the seal that protected the cluster it came from: a PROD snapshot can only be opened by a cluster that unseals with the same Transit key on COMMON, and a COMMON snapshot only with COMMON's unseal shares.

Automated snapshots are an Enterprise feature. With the Community edition something else has to call the API on a schedule.

The snapshot agent

That something was vault_raft_snapshot_agent, a small open-source Go program. It runs on every node of a cluster, asks its local Vault every interval whether that node is the leader, and only on the leader takes a snapshot and stores it.

mermaid
flowchart LR
  subgraph n1["node1, standby"]
    a1["agent"] -. "leader? no" .-> v1["Vault"]
  end
  subgraph n3["node3, active"]
    a3["agent"] -- "leader? yes" --> v3["Vault"]
    a3 -- "AppRole login, GET sys/storage/raft/snapshot" --> v3
  end
  subgraph n5["node5, standby"]
    a5["agent"] -. "leader? no" .-> v5["Vault"]
  end
  a3 -- "upload" --> s3[("Bucket: snapshots/")]

Running it everywhere means no configuration changes when leadership moves, and no snapshot is taken twice.

Building and installing

There was no package. The agent was compiled once on a build host and the binary copied to the nodes.

bash
$ yum install go
$ wget https://github.com/Lucretius/vault_raft_snapshot_agent/archive/refs/tags/v0.3.1.zip
$ unzip v0.3.1.zip
$ cd vault_raft_snapshot_agent-0.3.1
$ go build

On each Vault node the binary is put in place and labelled.

bash
$ cp vault_raft_snapshot_agent /usr/local/bin
$ chmod +x /usr/local/bin/vault_raft_snapshot_agent
$ chown vault:vault /usr/local/bin/vault_raft_snapshot_agent
$ restorecon -RvF /usr/local/bin/vault_raft_snapshot_agent

Its identity in Vault

The agent logs in with an AppRole whose policy has one rule. The policy is a Config document: snapshots-policy.hcl.

bash
$ vault policy write snapshots-policy /etc/vault.d/policy/snapshots-policy_v01.hcl
$ vault write auth/approle/role/snapshots policies=snapshots-policy
$ vault write auth/approle/role/snapshots token_no_default_policy=true
$ vault read auth/approle/role/snapshots/role-id
$ vault write -f auth/approle/role/snapshots/secret-id

Configuration and service

Two Config documents:

bash
$ chown vault:vault /etc/vault.d/snapshot.json
$ restorecon -RvF /etc/vault.d/snapshot.json
$ restorecon -RvF /etc/systemd/system/vault-snapshot.service
$ systemctl start vault-snapshot
$ systemctl enable vault-snapshot
EnvironmentIntervalKept
PROD, NONPRODEvery two hoursFirst 168 on local disk, which is two weeks; later unlimited, in object storage
COMMONOnce a day in the first deployment14 on local disk in the first deployment; the design document gives all clusters the same final configuration

The detour: local files and rsync

The first version of the backup stored snapshots on the node, in /data/raft/snapshots. That has an obvious flaw. Only the leader takes snapshots, so only the leader has them, and the leader is the node whose loss one is preparing for.

The fix at the time was to copy the directory to the other four nodes whenever it changed. A systemd path unit watched the directory and started a service, and the service ran a script.

ini
[Unit]
Description="Monitor the /data/raft/snapshots directory for changes"

[Path]
PathChanged=/data/raft/snapshots
Unit=vault-snapshot-monitor.service

[Install]
WantedBy=multi-user.target
bash
#!/bin/bash
VAULT_SNAPSHOTS_DIR=/data/raft/snapshots
USER=recovery
VAULT2=nonprod-vault-node2.example.net

/usr/bin/rsync -avz --omit-dir-times --no-o --no-g --no-perms \
  -e "ssh -o StrictHostKeyChecking=no -i /home/recovery/.ssh/id_rsa" \
  $VAULT_SNAPSHOTS_DIR/ $USER@$VAULT2:$VAULT_SNAPSHOTS_DIR/

The script had one such rsync line per peer, and each node's copy named the other four. It worked. Looking at what it needed is the reason it was removed.

It neededWhich meant
SSH between all Vault nodesA firewall rule on every node that the design otherwise did not have
A private SSH key on every Vault node, for an account in wheelWhoever got that key from one node could log in to all of them
StrictHostKeyChecking=noThe copy would go to whatever answered at that name
Group write access to the snapshot directory for that accountA second identity with access to Vault's data directory tree
The same disk as the Raft databaseA full snapshot directory would stop Vault

And after all that, every snapshot was still inside the same eight virtual machines in the same availability zone.

In December 2023 the agent was pointed at the private cloud's object storage and the whole mechanism was taken out again: the path unit, the service, the script, the key, the rsync package, the directory permissions and the five firewall rules. The removal is in the notes as carefully as the installation, command by command.

bash
$ systemctl disable vault-snapshot-monitor.path
$ rm --force /etc/systemd/system/vault-snapshot-monitor.path /etc/systemd/system/vault-snapshot-monitor.service
$ rm --force /home/recovery/.ssh/id_rsa /usr/local/bin/vault-snapshot-sync.sh
$ yum remove rsync
$ chown vault:vault /data/raft /data/raft/snapshots
$ chmod 700 /data/raft /data/raft/snapshots
$ firewall-cmd --permanent --zone=public --remove-rich-rule='rule family=ipv4 source address=10.20.1.32/28 destination address=10.20.1.34/32 port port=22 protocol=tcp accept'
$ firewall-cmd --reload

Final state

mermaid
flowchart LR
  l["Leader of each cluster<br/>snapshot agent, every 2 h"] -- "S3 API, HTTPS" --> a[("Object storage, site A<br/>bucket/snapshots/")]
  n["Every node<br/>cron, daily 07:00"] -- "S3 API, HTTPS" --> b[("bucket/logs/")]

One bucket per environment, with snapshots/ written by the agent and logs/ by the script of Audit logging and log shipping. The agent's retain is set so high that it never deletes; how long a snapshot is kept is decided by the bucket.

The platform also offers image-based backup of virtual machines. That is not a Vault backup: an image of one node taken while the cluster is writing is a copy of one replica at an arbitrary moment. It is what one would use to get a server back, not the data.

Restore

There are three situations, and they need different things.

SituationWhat is lostWhat to do
A minority of nodesNothingRebuild the nodes with the same configuration and an empty /data. They join through retry_join, unseal through COMMON and receive the data from the leader. No snapshot is involved.
Data, on a healthy clusterA deleted mount, a destroyed secret, a bad changeRestore a snapshot into the running cluster
The whole clusterEverythingBuild a new cluster, initialize it, restore a snapshot into it

Restoring into a running cluster replaces its entire state with the snapshot's.

bash
$ aws s3 cp s3://prod-vault/snapshots/<SNAPSHOT_FILE> /tmp/restore.snap
$ export VAULT_ADDR='http://127.0.0.1:8200'
$ vault login
$ vault operator raft snapshot restore /tmp/restore.snap

Everything written after the snapshot is gone, including tokens and SecretIDs issued since. With a two-hour interval that is the exposure.

Restoring into a newly built cluster is the case to rehearse. The new cluster is initialized first and so has its own keys; the snapshot was encrypted with the old ones.

bash
$ vault operator init -recovery-shares=5 -recovery-threshold=2
$ vault login
$ vault operator raft snapshot restore -force /tmp/restore.snap

-force skips the check that the snapshot belongs to this cluster. After the restore the cluster's data, including its encrypted root key, is the old cluster's. The node can read it because its seal block points at the same Transit key on COMMON, and the recovery keys and root tokens that are valid from then on are the old cluster's, not the ones the fresh initialization printed a minute before. Whoever does this needs the old recovery keys or a working administrator login.

That dependency is worth saying once more as a table.

To restoreYou need
PROD or NONPRODA snapshot, COMMON running and unsealed, the cluster's Transit key on COMMON, a valid unseal token
COMMONA snapshot of COMMON, and two of COMMON's five unseal shares

COMMON's snapshot is small and changes almost never, and it is the one everything else depends on.

What was not done

The notes hold the build of the backup in full and no record of a restore test. The procedure above is the documented one, and it is the design's intention, not a log of something I ran against these clusters. A backup that has not been restored is a hope. A quarterly restore of the latest PROD snapshot into a scratch cluster, unsealing through COMMON, would have tested the snapshots, the Transit key, the token and the runbook in one exercise.

Reading it today

← solutionz