Vault 01 - Use cases and requirements
Vault Solution · Next: Secret model
The platform this Vault served sells managed network services to business customers. To configure a customer's service it needs things the customer would rather not give away: the keys of a cloud account, the administrator login of a security service, the pre-shared key of a VPN tunnel. This article says who handles those values, what each party may do with them, and what the secrets store had to guarantee.
The problem
Before Vault, a sensitive value was one more attribute of a service order. It was typed into a portal, stored in the portal's database, sent to the orchestration platform inside the order, stored again there, and finally pushed to a device. Every system on the way held a copy, every backup of every system held a copy, and nobody could say who had read one.
The rule that replaced this is short: a secret is written to one place, by the system the customer types it into, and read from that place, by the system that needs it, at the moment it needs it. The order that travels between the two carries a reference, never the value.
The parties
| Party | What it is | What it does with secrets |
|---|---|---|
service-portal | The customer portal | Creates, updates and deletes secrets on the customer's behalf. Never reads one back. |
orchestrator | The orchestration platform | Reads secrets while it configures devices and cloud accounts. Creates one kind itself, when it generates VPN credentials in automatic mode. |
orchestrator-ui | The operator interface of the orchestration platform | Manages every kind of secret, for support cases |
| Administrators | The team that runs Vault | Run the clusters, the auth methods and the policies. Have no rule that lets them read a customer secret. |
| Snapshot agent | A process on every Vault node | Downloads Raft snapshots and nothing else |
How a secret is consumed
sequenceDiagram autonumber participant C as Customer participant P as service-portal participant V as Vault participant O as orchestrator participant D as Device or cloud API C->>P: Enters a credential P->>V: AppRole login V-->>P: Token with the writer policy P->>V: Write secret at a new UUID V-->>P: Version number P->>O: Order with UUID and version, no secret O->>V: AppRole login V-->>O: Token with the reader policy O->>V: Read secret at UUID and version V-->>O: Secret O->>D: Configure with the secret Note over V: Every request and response is written to the audit log
The portal can also ask the orchestrator to test a stored credential. The test runs in the orchestrator, which is allowed to read; the portal only learns whether the credential worked.
What is stored
Six kinds of customer secret, each for one feature of the platform.
| Feature | Secret | Written by | Read by |
|---|---|---|---|
| Internet breakout through a secure web gateway | Partner administrator account, API key and password of the gateway service | service-portal | orchestrator |
| The same, manual mode | VPN names and pre-shared keys the customer created at the gateway service | service-portal | orchestrator |
| The same, automatic mode | VPN names and pre-shared keys the orchestrator creates there itself | orchestrator | orchestrator |
| Virtual firewalls in a public cloud | Account credentials for AWS, Azure or GCP | service-portal | orchestrator |
| TLS inspection | Private key of a customer certificate, and its passphrase | service-portal | orchestrator |
| Dynamic routing | BGP and OSPF authentication keys | service-portal | orchestrator |
| IPsec tunnels to third parties | IKE pre-shared key | service-portal | orchestrator |
The exact paths and keys are in Secret model.
Functional requirements
| ID | Requirement |
|---|---|
| FR1 | Manage and protect secrets: tokens, passwords, certificates, encryption keys |
| FR2 | Offer three ways in: a web interface, a command line and an API |
| FR3 | Integrate with third-party systems through the API; at the start, the orchestration platform and the customer portal |
Non-functional requirements
| ID | Requirement | How the design answers it |
|---|---|---|
| NFR1 | Open standards | HTTP API, JSON, TLS, X.509 |
| NFR2 | Management only over secure channels | HTTPS for the API and the web interface, SSH for the command line, through one bastion host per environment |
| NFR3 | Certificates signed by a trusted authority | Every node certificate is issued by the organisation's internal CA |
| NFR4 | High availability, 99.95 % | Five-node Raft clusters, two load balancers with a floating address, two availability zones |
| NFR5 | The organisation's security baseline | CIS-hardened operating system, audit trail to the security operations centre |
| NFR6 | Reliability: other systems stop when this one does | Auto-unseal, so a restarted node rejoins without a person |
| NFR7 | Maintainability: as few components as possible, and ones the operations team already knows | Integrated Storage instead of a separate Consul cluster; HAProxy and Keepalived instead of a new load-balancing product |
| NFR8 | Network latency between availability zones below 8 ms | A single region |
| NFR9 | Operated in one legal region | Both zones are in the same country |
| NFR10 | Data protection law | Access control per consumer, complete audit log, defined retention |
The constraint that shaped the design
The system had to be built with the Community edition of Vault. That one sentence removes a list of features, and most design decisions in the following articles are answers to an item on it.
| Not available | Consequence |
|---|---|
| Performance standby nodes | Only the active node answers requests; a cluster scales up, not out |
| Performance and disaster-recovery replication | One cluster per environment, protected by snapshots |
| HSM support | No hardware module for the unseal key; a third Vault cluster does that job through the Transit engine |
| Automated snapshots | A separate snapshot agent, see Raft snapshots, backup and restore |
| Namespaces | Tenants are separated by path and policy inside one mount, see Secret model |
| Multi-factor authentication | Administrators reach Vault only through the bastion host, which is itself behind the organisation's access gateway |
The second constraint was Linux as the operating system.
Decisions
| ID | Decision | Reason |
|---|---|---|
| D1 | Every Vault node is a virtual machine | The organisation's private cloud was the given platform |
| D2 | The private cloud, with its two availability zones | It met the region requirement; a third zone did not exist |
| D3 | Use both zones | Availability |
| D4 | One region | Raft needs low latency between voters |
| D5 | Integrated Storage (Raft) as the storage backend | Fewer components, fewer machines, nothing to learn besides Linux and Vault, ordinary backups |
| D6 | Five nodes in PROD and NONPROD, three in COMMON | Five is the recommended size and survives two failures; COMMON holds two keys and no data, so three is enough |
Reading it today
The Community-edition list is unchanged in Vault 2.1: replication, performance standbys, HSM seals, namespaces, automated snapshots and seal high availability are still Enterprise features. What changed is the licence. Vault 1.14, the version this system ran, was the last under the MPL; 1.15 and later are under the Business Source License 1.1. A team with the same constraint today would also look at OpenBao, the MPL-licensed fork.