LINUXOR.SK ... open source notes ...

Vault 10 - Initialization and Transit auto-unseal

category: solutionz · date: 2024-12-31 · updated: 2026-10-02 · author: LALA

Vault Solution · Previous: Vault server configuration · Next: Authentication and policies

A Vault that has just started holds encrypted data and no key to it. Unsealing gives it the key. This article brings up the three clusters in the only order that works, COMMON first, and then looks hard at the one credential the whole arrangement depends on.

Sealed, unsealed, and who holds the key

Vault encrypts everything it stores with a data encryption key. That key is stored next to the data, encrypted with the root key, and the root key is stored too, encrypted by the seal. What the seal is decides who has to be present when a node starts.

SealThe root key is protected byA node startsUsed for
ShamirA key split into shares, held by peopleSealed, until enough shares are enteredCOMMON
TransitA key in the Transit engine of another VaultUnsealed, if the other Vault answersPROD, NONPROD
mermaid
sequenceDiagram
  participant N as prod-vault node
  participant L as common-vault load balancer
  participant C as common-vault, active node
  Note over N: Starts sealed and reads the encrypted root key from /data
  N->>L: POST transit/decrypt/autounseal-prod-vault, with the token from vault.hcl
  L->>C: Passed through
  Note over C: Checks the token's policy, decrypts with the Transit key, writes an audit entry
  C-->>N: Root key in plain text
  Note over N: Decrypts the keyring, is unsealed, joins or leads the cluster

With a Transit seal there are no unseal keys. Initializing such a cluster yields recovery keys instead. They cannot unseal anything; they authorise the few operations that need a quorum of people, such as generating a new root token.

Order of work

StepWhereWhat
1COMMONStart the three nodes, initialize with Shamir shares, unseal each node
2COMMONEnable the Transit engine and create one key per client cluster
3COMMONWrite one policy per client cluster and issue one token for each
4PROD, NONPRODPut the token and key name into the seal block on every node
5PROD, NONPRODStart the nodes, initialize once, watch the others join and unseal

COMMON

COMMON has no seal block, so it uses the default, Shamir. The cluster is initialized once, on one node.

bash
$ export VAULT_ADDR='http://127.0.0.1:8200'
$ vault operator init -key-shares=5 -key-threshold=2
ParameterValueMeaning
key-shares5The unseal key is split into five shares
key-threshold2Any two of them unseal a node

The command prints the five shares and the initial root token, once. Then every node is unsealed by two of them; the second and third node join through retry_join and are unsealed the same way.

bash
$ vault operator unseal
$ vault operator unseal
$ vault status

This is the manual work the design keeps for exactly one cluster: after any restart of a COMMON node, two share holders are needed.

Transit engine and keys

bash
$ vault secrets enable -default-lease-ttl=26280h -max-lease-ttl=26280h transit
$ vault write -f transit/keys/autounseal-prod-vault
$ vault write -f transit/keys/autounseal-nonprod-vault
KeyPathWraps the root key of
autounseal-prod-vaulttransit/keys/autounseal-prod-vaultprod-vault
autounseal-nonprod-vaulttransit/keys/autounseal-nonprod-vaultnonprod-vault

The mount is enabled with a lease limit of 26,280 hours, three years. That setting does not govern the unseal token, which is issued by the token auth method and not by this mount. What lifts the token over the server-wide ten-hour limit of vault.hcl is its explicit maximum lifetime, in the next step.

Policies and tokens

Each client cluster gets a policy that allows two operations on its own key and nothing else. The policy is a Config document: autounseal-prod-vault.hcl.

bash
$ mkdir /etc/vault.d/policy
$ vault policy write autounseal-prod-vault /etc/vault.d/policy/autounseal-prod-vault.hcl
$ vault policy write autounseal-nonprod-vault /etc/vault.d/policy/autounseal-nonprod-vault.hcl

The token is created response-wrapped. What the command prints is not the token but a single-use wrapping token that lives for two minutes; the real token is fetched with it once.

bash
$ vault token create -policy="autounseal-prod-vault" -wrap-ttl=120 -ttl=26280h -explicit-max-ttl=26280h
output 7 lines
Key                              Value
---                              -----
wrapping_token:                  <WRAPPING_TOKEN>
wrapping_accessor:               <ACCESSOR>
wrapping_token_ttl:              2m
wrapping_token_creation_path:    auth/token/create
wrapped_accessor:                <ACCESSOR>
bash
$ VAULT_TOKEN="<WRAPPING_TOKEN>" vault unwrap
output 6 lines
Key                  Value
---                  -----
token                <TRANSIT_UNSEAL_TOKEN>
token_duration       26280h
token_policies       ["autounseal-prod-vault" "default"]
policies             ["autounseal-prod-vault" "default"]

Wrapping means the token never appears on a screen or in a shell history on the way from the administrator of COMMON to the administrator of PROD, and if someone else had unwrapped it first, the second attempt would fail and say so.

PROD and NONPROD

The token goes into the seal "transit" block of vault.hcl on every node, with the address of COMMON's load balancer and the name of the key. Then the first node is started and the cluster initialized there, once.

bash
$ export VAULT_ADDR='http://127.0.0.1:8200'
$ vault operator init -recovery-shares=5 -recovery-threshold=2
output 12 lines
Recovery Key 1: <RECOVERY_KEY_1>
Recovery Key 2: <RECOVERY_KEY_2>
Recovery Key 3: <RECOVERY_KEY_3>
Recovery Key 4: <RECOVERY_KEY_4>
Recovery Key 5: <RECOVERY_KEY_5>

Initial Root Token: <ROOT_TOKEN>

Success! Vault is initialized

Recovery key initialized with 5 key shares and a key threshold of 2. Please
securely distribute the key shares printed above.

Nobody unseals anything. The node is unsealed the moment it is initialized, and each further node, once started, finds the cluster through retry_join, joins, and unseals itself through COMMON.

bash
$ vault status
output 16 lines
Key                      Value
---                      -----
Recovery Seal Type       shamir
Initialized              true
Sealed                   false
Total Recovery Shares    5
Threshold                2
Version                  1.14.0
Storage Type             raft
Cluster Name             prod-vault
HA Enabled               true
HA Cluster               https://prod-vault-node4.example.net:8201
HA Mode                  standby
Active Node Address      https://prod-vault.example.net:8200
Raft Committed Index     295
Raft Applied Index       295
bash
$ vault login
$ vault operator raft list-peers
output 7 lines
Node                            Address                              State       Voter
----                            -------                              -----       -----
prod-vault-node1.example.net    prod-vault-node1.example.net:8201    follower    true
prod-vault-node2.example.net    prod-vault-node2.example.net:8201    follower    true
prod-vault-node3.example.net    prod-vault-node3.example.net:8201    follower    true
prod-vault-node4.example.net    prod-vault-node4.example.net:8201    leader      true
prod-vault-node5.example.net    prod-vault-node5.example.net:8201    follower    true

Autopilot gives the same picture with health in it.

bash
$ vault operator raft autopilot state
output 9 lines
Healthy:                         true
Failure Tolerance:               2
Leader:                          prod-vault-node4.example.net
Voters:
   prod-vault-node4.example.net
   prod-vault-node1.example.net
   prod-vault-node2.example.net
   prod-vault-node3.example.net
   prod-vault-node5.example.net

And the health endpoint on a standby, as HAProxy sees it:

bash
$ curl http://127.0.0.1:8200/v1/sys/health
output 1 line
{"initialized":true,"sealed":false,"standby":true,"performance_standby":false,"replication_performance_mode":"disabled","replication_dr_mode":"disabled","version":"1.14.0","cluster_name":"prod-vault"}

The initial root token was used for the configuration in the next article. The notes do not show it being revoked afterwards, and an audit sample shows a root token issued in July 2023 still in use that November. It should be revoked as soon as the first administrator account works; a new one can be generated with the recovery keys when it is really needed.

What depends on what

If this is unavailableA running PROD clusterA PROD node that restarts
COMMON, sealed or downUnaffectedStays sealed
The network path to COMMON's virtual addressUnaffectedStays sealed
COMMON's certificate, expiredUnaffectedStays sealed
The unseal token, expired or revokedUnaffectedStays sealed
The Transit key, deletedUnaffectedStays sealed, for good

A running Vault keeps its keys in memory and does not ask the seal again. That makes every row of this table quiet: nothing fails when the dependency breaks, only later, when a node restarts for an unrelated reason. Each row therefore needs its own check, independent of whether Vault is serving.

The last row is the reason the Transit keys are never rotated carelessly and never deleted, and why COMMON has its own snapshots: without that key, the encrypted root key in PROD's storage cannot be opened by anyone, recovery keys included.

The token that runs out

The unseal tokens were created on 24 July 2023 with -ttl=26280h -explicit-max-ttl=26280h. An explicit maximum lifetime is a hard limit that no renewal can pass. disable_renewal = "false" in the seal block lets Vault renew the token, and renewal cannot extend it beyond that limit.

Twenty-six thousand two hundred and eighty hours after 24 July 2023 is 23 July 2026.

From that day the token is gone. Nothing happens to the running clusters, by the first row of reasoning above. The first PROD node to restart afterwards, for a kernel update or a certificate, stays sealed. If the nodes are restarted one after another, as a careful rolling update does, the cluster loses its majority at the third.

The explicit maximum was the way around the ten-hour limit in vault.hcl: it overrides the server-wide maximum, and three years is long enough to be forgotten. A credential that a system needs in order to start must either never expire or be renewed by something that is watched, and its expiry must be on a calendar that outlives the people who created it.

BetterHow
A periodic tokenvault token create -orphan -period=24h -policy=autounseal-prod-vault. A periodic token is not subject to the maximum lifetime; it lives as long as it is renewed, which the seal does by itself. If PROD is switched off for more than the period, the token dies; that is a feature.
Not in vault.hclThe token in an environment file readable by vault only, or referenced as file://, so that the configuration file can be world-readable and kept in version control
A checkA monitoring item that looks the token up by its accessor on COMMON and alerts on remaining lifetime
A written rotationThe steps to issue a new token and restart the nodes one by one, tested on NONPROD

Reading it today

← solutionz