Vault 12 - Audit logging and log shipping
Vault Solution · Previous: Authentication and policies · Next: Raft snapshots, backup and restore
Every operation in Vault is an API request, and Vault can write every request and every response to an audit device before it answers. This article sets up that trail, the operating system's own logs beside it, and the three places the logs go: a local file, the security operations centre, and object storage.
Where a log line goes
flowchart LR subgraph node["Vault node"] v["Vault"] ad["auditd"] os["sshd, sudo, cron, systemd"] rs["rsyslog"] f1[("/var/log/vault/audit.log")] f2[("/var/log/audit/audit.log")] f3[("/var/log/secure, messages, cron")] lr["logrotate, daily"] sync["vault-logs-s3-sync.sh, 07:00"] v -- "file audit device" --> f1 ad --> f2 ad -- "local6" --> rs os --> rs rs --> f3 f1 --> lr lr -- "node-audit-date.log.gz" --> sync end v -- "socket audit device, TCP 50543" --> soc1["SOC audit collector"] rs -- "TCP 50515" --> soc2["Central log collector"] sync -- "S3 over HTTPS" --> s3[("Bucket: logs/")]
| Source | Local file | Sent to | Kept locally |
|---|---|---|---|
| Vault audit devices | /var/log/vault/audit.log | The audit collector of the security operations centre, live; object storage, daily | 14 days |
| Vault server messages | The systemd journal | Nowhere else | Journal defaults |
| Linux audit daemon | /var/log/audit/audit.log | Central log collector, live | 14 days |
| Authentication, sudo, cron | /var/log/secure, /var/log/cron | Central log collector, live | 14 days |
| HAProxy connections | /var/log/haproxy.log on the load balancers | Nowhere else | 14 days |
Vault audit devices
Audit devices are enabled once per cluster, on the active node, and apply to all nodes.
$ mkdir /var/log/vault $ chown vault:vault /var/log/vault $ vault audit enable file file_path=/var/log/vault/audit.log $ vault audit enable socket address="10.40.7.5:50543" socket_type="tcp" $ vault audit list -detailed
The directory has to exist on every node, because whichever node is active writes the file.
| Device | Destination | Purpose |
|---|---|---|
file | /var/log/vault/audit.log | The record on the node, for troubleshooting and for the archive |
socket | A dedicated collector of the security operations centre, TCP | Live analysis and alerting |
What an entry looks like
Each line is one JSON object, either a request or the response to it. This is a request, shortened and laid out for reading; the values are examples.
{
"time": "2023-11-14T09:55:09.675895359Z",
"type": "request",
"auth": {
"client_token": "hmac-sha256:5f2c…",
"accessor": "hmac-sha256:9a41…",
"display_name": "userpass-admin01",
"policies": ["admin-policy"],
"token_type": "service"
},
"request": {
"id": "0b1e6f6e-…",
"operation": "update",
"mount_point": "sys/",
"mount_type": "system",
"path": "sys/policies/acl/orchestrator-policy-kv2",
"remote_address": "127.0.0.1",
"remote_port": 64790
}
}Three things about it matter in practice.
Secrets are hashed, not written. Token values, and the values of secrets in requests and responses, appear as HMAC-SHA256 digests with a key that is specific to the audit device. A known value can be checked against the log with sys/audit-hash; the log itself does not reveal it.
The remote address is whoever opened the TCP connection. For an administrator on a node that is 127.0.0.1. For an API client it is a load-balancer node, as explained in Load balancer, and the client's identity has to be taken from the auth block instead.
A few paths are never audited. Health, seal status, leader, initialization, unseal and the Raft join and bootstrap endpoints are answered before the audit system is involved. The health checks of HAProxy, one every five seconds from each load balancer to each node, leave no trace here.
The rule that makes audit a dependency
Vault does not answer a request it could not log. If at least one audit device is enabled, an entry has to be written to at least one of them before the response goes out. With two devices, either may fail. If both fail, Vault stops answering.
That is the reason for two devices of different kinds. The file fails when the disk is full; the socket fails when the collector or the network is down. It is also the reason log rotation and free space on /var have to be watched as if they were part of Vault, because they are.
The socket device has a sharper edge than the file. TCP to a collector that has stopped reading does not fail at once; it blocks until a timeout.
What the security operations centre looks for
The collector feeds an analytics platform with alert rules written for Vault's audit format. These were the events that raised an alert.
| Event | Why it matters |
|---|---|
| A secrets engine enabled or disabled | New capability, or a mount and its data removed |
| An auth method enabled or disabled | A new way in |
| A userpass account created, changed or deleted | A new administrator |
| An AppRole role created, changed or deleted | A new machine identity, or a changed policy list |
| A token created or revoked outside login | Tokens made by hand bypass the auth methods |
| A policy changed or deleted | Permissions widened |
| An audit device disabled | Someone is turning off the lights |
| An administrator login outside 06:00 to 19:00 | Unusual hours |
| A request with a missing or improper token | Probing |
| Five failed logins within two minutes | Guessing |
The list is short on purpose. Every row is something an administrator can do with admin-policy, and the alert is what turns "the administrator has no rule for customer secrets" from a convention into something a second team would notice being changed.
Operating-system logs
rsyslog writes the usual local files and forwards the security-relevant facilities to the central collector over TCP, with a disk-assisted queue so that an unreachable collector delays messages and does not block the host. The configurations are Config documents:
The collector does not appear in DNS for these servers; its name is in /etc/hosts. Its port is not one SELinux knows as a syslog port, so it is added once per server.
$ semanage port -a -t syslogd_port_t -p tcp 50515 $ systemctl restart rsyslog
The Linux audit daemon
auditd writes its own file and is not a syslog client by default. Its syslog plugin is switched on and pointed at a facility of its own, which rsyslog then forwards.
active = yes direction = out path = builtin_syslog type = builtin args = LOG_LOCAL6 format = string
That is /etc/audit/plugins.d/syslog.conf. The audit backlog and the audit=1 kernel parameter from Virtual machines and OS build make sure events from early boot are not lost.
Rotation
One policy for every log on every server: rotate daily, keep fourteen, compress, date in the name. Three mechanisms implement it, because three kinds of software write logs.
| Mechanism | Rotates | Configuration |
|---|---|---|
| logrotate | Everything that is a plain file | logrotate.conf and one file per package in /etc/logrotate.d |
| A daily cron script | The Linux audit log, which auditd must rotate itself | auditd-logrotate, with max_log_file_action = ignore in auditd.conf |
| DNF's own settings | The package manager's logs | log_rotate=14 and log_compress=True in /etc/dnf/dnf.conf |
The drop-in files in /etc/logrotate.d for the system's own logs (syslog, btmp, wtmp, chrony, firewalld, sssd, aide, bootlog, dnf) are the distribution's, reduced to what the main file does not already say. Two are specific to this Solution:
- logrotate.d/vault, which also renames the rotated audit log so that the file name carries the node
- logrotate.d/haproxy
$ rm --force /etc/logrotate.d/*.rpmnew $ chmod 644 /etc/logrotate.d/vault $ logrotate --debug /etc/logrotate.conf
Shipping audit logs to object storage
The organisation's security baseline asks for 90 days of audit logs on storage that cannot be rewritten. Fourteen days on a node's disk is neither. The rotated, compressed files are copied once a day to the environment's bucket in the private cloud's S3-compatible object storage, which is where retention and write-once protection have to be configured.
Preparation, once per Vault node: the storage endpoints in /etc/hosts, the internal CA in the trust store, and the AWS command line client.
$ curl "https://awscli.amazonaws.com/awscli-exe-linux-x86_64.zip" -o awscliv2.zip $ unzip awscliv2.zip $ ./aws/install $ export AWS_ENDPOINT_URL=https://s3-a.example.net $ export AWS_CA_BUNDLE=/etc/ssl/certs/ca-bundle.crt $ aws s3 ls $ aws s3 sync /var/log/vault s3://prod-vault/logs/ --exclude "*" --include "*.log.gz" $ aws s3 ls s3://prod-vault/logs/
output 3 lines
2023-12-20 09:34:23 5628 vault1-audit-2023-12-12.log.gz 2023-12-20 09:40:20 6305 vault1-audit-2023-12-13.log.gz 2023-12-20 09:40:20 6310 vault1-audit-2023-12-14.log.gz
The daily run is a script, vault-logs-s3-sync.sh, started from crontab at 07:00. Each node uploads its own files. The active node writes nearly all audit entries, so on most days four of the five have little or nothing new; the node name in the file name keeps them apart when leadership moves.
The bucket has two top-level directories, logs for this and snapshots for Raft snapshots, backup and restore. The storage has an endpoint in each availability zone; the script names the first.
Reading it today
- The commands are unchanged in Vault 2.1. File and socket devices are enabled the same way, and there is still no way to declare audit devices in the server configuration file.
- Two devices of different types is now the written recommendation, with one of them shipping off the host. This build had that.
copytruncateshould go. Vault reopens its audit file onSIGHUP. Rotating by rename and then signalling Vault loses nothing; copying and truncating a file that is being written can.- File mode matters more than it did. Current releases refuse a file audit device whose mode has an execute bit, and will not unseal if that is the only device. The default mode,
0600, is fine. - Filtering, exclusion and fallback devices are Enterprise. The options are accepted and do nothing in the Community edition.
- The secrets in the sync script are the weak spot. An access key in a root-only file is ordinary practice and still a long-lived credential on five nodes. A bucket policy that lets this key write under
logs/and neither read nor delete would limit what it is worth. - The socket device and an unresponsive collector. The documentation says plainly that a blocked TCP destination can make Vault unresponsive. If the off-host copy matters less than availability, shipping the file with a log forwarder instead of a second audit device moves that risk out of Vault's request path.
- Audit devices
- File audit device
- Socket audit device
- Audit best practices