Vault 11 - Authentication and policies
Vault Solution · Previous: Initialization and Transit auto-unseal · Next: Audit logging and log shipping
Vault answers two questions on every request: who is asking, and may they do this to that path. This article sets up the three ways of proving identity that the Solution used and the six policies that decide the rest, including the one design idea I would keep in any rebuild: the system that writes secrets cannot read them.
Who authenticates how
| Persona | What it is | Auth method | Policy |
|---|---|---|---|
| Administrator | A named person on the Vault team | userpass | admin-policy |
service-portal | The customer portal | AppRole | service-portal-policy-kv2 |
orchestrator | The orchestration platform | AppRole | orchestrator-policy-kv2 |
orchestrator-ui | Its operator interface | AppRole | orchestrator-ui-policy-kv2 |
| Snapshot agent | A process on every Vault node | AppRole | snapshots-policy |
prod-vault, nonprod-vault | The clusters themselves, as clients of COMMON | Token | autounseal-prod-vault, autounseal-nonprod-vault |
What is switched on in each cluster is as little as that table needs.
| Environment | Auth methods enabled | Secrets engines enabled |
|---|---|---|
| PROD | userpass/, approle/ | kv2/, KV version 2 |
| NONPROD | userpass/, approle/ | kv2/, KV version 2 |
| COMMON | userpass/, approle/ | transit/ |
The token method is built in and always present. Nothing else was enabled: no LDAP, no cloud or Kubernetes auth, no other engine.
From login to answer
sequenceDiagram participant A as Application participant V as Vault A->>V: POST auth/approle/login with role_id and secret_id Note over V: Finds the role, checks the secret_id V-->>A: Token, with the role's policies attached A->>V: Request to a path, with the token Note over V: Looks up the token, then the policies on it Note over V: Is there a rule for this path with this capability? alt A rule allows it V-->>A: Result else No rule, or deny V-->>A: 403 permission denied end Note over V: Both outcomes are written to the audit log
Policies are deny by default. A token with no policy can do nothing, and a path that no rule mentions does not exist as far as that token is concerned.
AppRole
AppRole is the method for machines. A role is a named bundle of policies. An application proves it may use the role with two values: the RoleID, which identifies the role and is not very secret, and a SecretID, which is.
$ vault auth enable approle $ vault write auth/approle/role/orchestrator-kv2 policies=orchestrator-policy-kv2 $ vault write auth/approle/role/orchestrator-kv2 token_no_default_policy=true $ vault read auth/approle/role/orchestrator-kv2/role-id $ vault write -f auth/approle/role/orchestrator-kv2/secret-id
| Role | Policy | On |
|---|---|---|
orchestrator-kv2 | orchestrator-policy-kv2 | PROD, NONPROD |
orchestrator-ui-kv2 | orchestrator-ui-policy-kv2 | PROD, NONPROD |
service-portal-kv2 | service-portal-policy-kv2 | PROD, NONPROD |
snapshots | snapshots-policy | PROD, NONPROD, COMMON |
NONPROD has further pairs of orchestrator roles, one pair per additional test environment of the orchestration platform, all bound to the same two policies. Separate roles mean separate SecretIDs, so one test environment can be cut off without touching the others.
token_no_default_policy=true keeps the built-in default policy off the tokens. default is harmless, mostly self-service on the token itself, and leaving it off means a consumer token carries exactly one policy and its abilities can be read from one file.
How the RoleID and SecretID reach the consuming system is outside Vault. From there on they are the consumer's to protect.
What the roles left at their defaults
The roles were created with a policy list and nothing else. Every other property of an AppRole kept its default.
| Property | Default, as built | What it means |
|---|---|---|
secret_id_ttl | 0 | A SecretID never expires |
secret_id_num_uses | 0 | A SecretID can be used any number of times |
secret_id_bound_cidrs, token_bound_cidrs | empty | A SecretID and its tokens work from any address |
token_ttl, token_max_ttl | 0 | Tokens fall back to the server-wide ten hours |
For the PROD roles the network already did what the CIDR bindings would have: the only way to the API was through the load balancer, from a handful of known addresses. It would still have cost one line per role to say so in Vault as well, and the audit log would then have shown a refused login instead of nothing.
Userpass
People log in with a user name and a password that Vault stores itself.
$ vault auth enable userpass $ set +o history $ vault write auth/userpass/users/admin01 password=<INITIAL_PASSWORD> policies=admin-policy $ set -o history $ vault login -method=userpass username=admin01
Three named administrators per environment, one account each, all with admin-policy. Switching off shell history around the command keeps the initial password out of .bash_history; it does not keep it out of the process list, so reading the password from standard input with password=- is the better form.
There was no directory integration and, in the Community edition, no multi-factor login. What stood in for both was the path: Vault's API is reachable for administrators only from the bastion host, the bastion only from the organisation's access gateway, and every login is in the audit trail the security operations centre watches.
The policies
Each is a Config document.
| Policy | On | Grants |
|---|---|---|
| admin-policy | all three | Cluster operations, auth methods, mounts, policies. No rule for kv2/. |
| service-portal-policy-kv2 | PROD, NONPROD | Write, without read |
| orchestrator-policy-kv2 | PROD, NONPROD | Read; manage one type |
| orchestrator-ui-policy-kv2 | PROD, NONPROD | Manage everything |
| snapshots-policy | all three | Read one path |
| autounseal-prod-vault and its NONPROD twin | COMMON | Encrypt and decrypt with one key |
Who may do what to which secret
| Object type | service-portal | orchestrator | orchestrator-ui |
|---|---|---|---|
webGatewayCredentials | write | read | manage |
webGatewayVpnManual | write | read | manage |
webGatewayVpnAuto | none | manage | manage |
cloudAwsCredentials | write | read | manage |
cloudAzureCredentials | write | read | manage |
cloudGcpCredentials | write | read | manage |
certificateCustomer | write | read | manage |
routingAuthKey | write | read | manage |
ipsecPreSharedKey | write | read | manage |
"Write" is create, update, delete, list and patch. "Read" is read and list. "Manage" is all of them. All three policies also allow list on kv2/*, so a consumer can walk the tree and see that objects exist without being able to open them.
A writer that cannot read
The portal is where a customer types a secret, and so it is the system customers can reach. Its policy has no read on any object. If the portal is compromised, the attacker can overwrite or delete customer secrets, which is visible and recoverable from KV versions and snapshots, and cannot collect them. Collecting requires the orchestrator's identity.
The same thought runs the other way: the orchestrator can read and, with one exception, cannot change what a customer entered. The exception is the type it creates itself.
The operator interface breaks the pattern on purpose. Support staff need to see and repair what is stored, so orchestrator-ui can do everything, and that makes its SecretID the most valuable one in the system.
The wildcard that grants too much
Every rule for an object type has this shape.
path "kv2/+/+/cloudAwsCredentials/+" { capabilities = ["create", "update", "delete", "list", "patch"] }
The second and fourth + stand for the customer and the UUID. The first + stands where KV version 2 puts its own prefix.
| Path | What it does | Matched by kv2/+/… |
|---|---|---|
kv2/data/<customer>/<type>/<uuid> | Read and write the secret | Yes |
kv2/metadata/<customer>/<type>/<uuid> | List versions; delete the secret and all its versions for good | Yes |
kv2/delete/… | Soft-delete chosen versions | Yes |
kv2/undelete/… | Restore soft-deleted versions | Yes |
kv2/destroy/… | Destroy chosen versions for good | Yes |
kv2/subkeys/… | Read the key names of a secret without the values | Yes |
Two consequences follow. A policy that was meant to let the portal delete a secret also lets it destroy every stored version, which removes the rollback that Secret model was designed around. And the reader policy's read applies to subkeys and metadata as well, which is harmless, while its list is granted far more widely than needed.
One rule in the portal's policy has data spelled out, kv2/data/+/cloudGcpCredentials/+. Carried through, every type would have two rules, for example:
path "kv2/data/+/cloudGcpCredentials/+" { capabilities = ["create", "update", "patch"] } path "kv2/metadata/+/cloudGcpCredentials/+" { capabilities = ["list", "delete"] }
Two rules per type instead of one, each naming the prefix it means. That is the form I would write today.
Applying a policy
Policy files are kept on the first node of each cluster under /etc/vault.d/policy, with a version suffix in the file name, and loaded under a stable name.
$ mkdir -p /etc/vault.d/policy $ vault policy write orchestrator-policy-kv2 /etc/vault.d/policy/orchestrator-policy-kv2_v03.hcl $ vault policy read orchestrator-policy-kv2 $ vault token capabilities kv2/data/customer-a/cloudAwsCredentials/1234
The last command answers "what may my current token do here" and is the quickest way to test a rule. The files had reached their third version by February 2024, as object types were added. A directory on one node is not version control; these files belong in a repository, with the policy applied from it.
Reading it today
- The syntax is unchanged in Vault 2.1. Paths,
+and*, and the capability names all mean what they meant. AppRole and userpass parameters are unchanged too. - A key written twice in one rule is a hard error since 1.21. Before that the last one won silently.
- Parameter constraints still do not work on KV v2, so "this consumer may write only these keys" still cannot be said in a policy.
- Wildcards from identity templates are refused since 2.0.1. That does not touch these policies, which use no templates, and it closes a hole in a per-customer design that would: a customer identifier containing
+or*no longer widens a templated path. - User lockout. Since 1.13 Vault locks an account for a while after repeated failed logins, for userpass and AppRole, and a
user_lockoutblock in the server configuration tunes it. The alert of the security operations centre for five failed logins in two minutes sat on top of that. - Root-level operations need a token since 2.0. See Initialization and Transit auto-unseal.
- Policies
- AppRole auth method
- KV v2 API
- User lockout