LINUXOR.SK ... open source notes ...

Vault 12 - Audit logging and log shipping

category: solutionz · date: 2024-12-31 · updated: 2026-10-02 · author: LALA

Vault Solution · Previous: Authentication and policies · Next: Raft snapshots, backup and restore

Every operation in Vault is an API request, and Vault can write every request and every response to an audit device before it answers. This article sets up that trail, the operating system's own logs beside it, and the three places the logs go: a local file, the security operations centre, and object storage.

Where a log line goes

mermaid
flowchart LR
  subgraph node["Vault node"]
    v["Vault"]
    ad["auditd"]
    os["sshd, sudo, cron, systemd"]
    rs["rsyslog"]
    f1[("/var/log/vault/audit.log")]
    f2[("/var/log/audit/audit.log")]
    f3[("/var/log/secure, messages, cron")]
    lr["logrotate, daily"]
    sync["vault-logs-s3-sync.sh, 07:00"]
    v -- "file audit device" --> f1
    ad --> f2
    ad -- "local6" --> rs
    os --> rs
    rs --> f3
    f1 --> lr
    lr -- "node-audit-date.log.gz" --> sync
  end
  v -- "socket audit device, TCP 50543" --> soc1["SOC audit collector"]
  rs -- "TCP 50515" --> soc2["Central log collector"]
  sync -- "S3 over HTTPS" --> s3[("Bucket: logs/")]
SourceLocal fileSent toKept locally
Vault audit devices/var/log/vault/audit.logThe audit collector of the security operations centre, live; object storage, daily14 days
Vault server messagesThe systemd journalNowhere elseJournal defaults
Linux audit daemon/var/log/audit/audit.logCentral log collector, live14 days
Authentication, sudo, cron/var/log/secure, /var/log/cronCentral log collector, live14 days
HAProxy connections/var/log/haproxy.log on the load balancersNowhere else14 days

Vault audit devices

Audit devices are enabled once per cluster, on the active node, and apply to all nodes.

bash
$ mkdir /var/log/vault
$ chown vault:vault /var/log/vault
$ vault audit enable file file_path=/var/log/vault/audit.log
$ vault audit enable socket address="10.40.7.5:50543" socket_type="tcp"
$ vault audit list -detailed

The directory has to exist on every node, because whichever node is active writes the file.

DeviceDestinationPurpose
file/var/log/vault/audit.logThe record on the node, for troubleshooting and for the archive
socketA dedicated collector of the security operations centre, TCPLive analysis and alerting

What an entry looks like

Each line is one JSON object, either a request or the response to it. This is a request, shortened and laid out for reading; the values are examples.

json
{
  "time": "2023-11-14T09:55:09.675895359Z",
  "type": "request",
  "auth": {
    "client_token": "hmac-sha256:5f2c…",
    "accessor": "hmac-sha256:9a41…",
    "display_name": "userpass-admin01",
    "policies": ["admin-policy"],
    "token_type": "service"
  },
  "request": {
    "id": "0b1e6f6e-…",
    "operation": "update",
    "mount_point": "sys/",
    "mount_type": "system",
    "path": "sys/policies/acl/orchestrator-policy-kv2",
    "remote_address": "127.0.0.1",
    "remote_port": 64790
  }
}

Three things about it matter in practice.

Secrets are hashed, not written. Token values, and the values of secrets in requests and responses, appear as HMAC-SHA256 digests with a key that is specific to the audit device. A known value can be checked against the log with sys/audit-hash; the log itself does not reveal it.

The remote address is whoever opened the TCP connection. For an administrator on a node that is 127.0.0.1. For an API client it is a load-balancer node, as explained in Load balancer, and the client's identity has to be taken from the auth block instead.

A few paths are never audited. Health, seal status, leader, initialization, unseal and the Raft join and bootstrap endpoints are answered before the audit system is involved. The health checks of HAProxy, one every five seconds from each load balancer to each node, leave no trace here.

The rule that makes audit a dependency

Vault does not answer a request it could not log. If at least one audit device is enabled, an entry has to be written to at least one of them before the response goes out. With two devices, either may fail. If both fail, Vault stops answering.

That is the reason for two devices of different kinds. The file fails when the disk is full; the socket fails when the collector or the network is down. It is also the reason log rotation and free space on /var have to be watched as if they were part of Vault, because they are.

The socket device has a sharper edge than the file. TCP to a collector that has stopped reading does not fail at once; it blocks until a timeout.

What the security operations centre looks for

The collector feeds an analytics platform with alert rules written for Vault's audit format. These were the events that raised an alert.

EventWhy it matters
A secrets engine enabled or disabledNew capability, or a mount and its data removed
An auth method enabled or disabledA new way in
A userpass account created, changed or deletedA new administrator
An AppRole role created, changed or deletedA new machine identity, or a changed policy list
A token created or revoked outside loginTokens made by hand bypass the auth methods
A policy changed or deletedPermissions widened
An audit device disabledSomeone is turning off the lights
An administrator login outside 06:00 to 19:00Unusual hours
A request with a missing or improper tokenProbing
Five failed logins within two minutesGuessing

The list is short on purpose. Every row is something an administrator can do with admin-policy, and the alert is what turns "the administrator has no rule for customer secrets" from a convention into something a second team would notice being changed.

Operating-system logs

rsyslog writes the usual local files and forwards the security-relevant facilities to the central collector over TCP, with a disk-assisted queue so that an unreachable collector delays messages and does not block the host. The configurations are Config documents:

The collector does not appear in DNS for these servers; its name is in /etc/hosts. Its port is not one SELinux knows as a syslog port, so it is added once per server.

bash
$ semanage port -a -t syslogd_port_t -p tcp 50515
$ systemctl restart rsyslog

The Linux audit daemon

auditd writes its own file and is not a syslog client by default. Its syslog plugin is switched on and pointed at a facility of its own, which rsyslog then forwards.

ini
active = yes
direction = out
path = builtin_syslog
type = builtin
args = LOG_LOCAL6
format = string

That is /etc/audit/plugins.d/syslog.conf. The audit backlog and the audit=1 kernel parameter from Virtual machines and OS build make sure events from early boot are not lost.

Rotation

One policy for every log on every server: rotate daily, keep fourteen, compress, date in the name. Three mechanisms implement it, because three kinds of software write logs.

MechanismRotatesConfiguration
logrotateEverything that is a plain filelogrotate.conf and one file per package in /etc/logrotate.d
A daily cron scriptThe Linux audit log, which auditd must rotate itselfauditd-logrotate, with max_log_file_action = ignore in auditd.conf
DNF's own settingsThe package manager's logslog_rotate=14 and log_compress=True in /etc/dnf/dnf.conf

The drop-in files in /etc/logrotate.d for the system's own logs (syslog, btmp, wtmp, chrony, firewalld, sssd, aide, bootlog, dnf) are the distribution's, reduced to what the main file does not already say. Two are specific to this Solution:

bash
$ rm --force /etc/logrotate.d/*.rpmnew
$ chmod 644 /etc/logrotate.d/vault
$ logrotate --debug /etc/logrotate.conf

Shipping audit logs to object storage

The organisation's security baseline asks for 90 days of audit logs on storage that cannot be rewritten. Fourteen days on a node's disk is neither. The rotated, compressed files are copied once a day to the environment's bucket in the private cloud's S3-compatible object storage, which is where retention and write-once protection have to be configured.

Preparation, once per Vault node: the storage endpoints in /etc/hosts, the internal CA in the trust store, and the AWS command line client.

bash
$ curl "https://awscli.amazonaws.com/awscli-exe-linux-x86_64.zip" -o awscliv2.zip
$ unzip awscliv2.zip
$ ./aws/install
$ export AWS_ENDPOINT_URL=https://s3-a.example.net
$ export AWS_CA_BUNDLE=/etc/ssl/certs/ca-bundle.crt
$ aws s3 ls
$ aws s3 sync /var/log/vault s3://prod-vault/logs/ --exclude "*" --include "*.log.gz"
$ aws s3 ls s3://prod-vault/logs/
output 3 lines
2023-12-20 09:34:23       5628 vault1-audit-2023-12-12.log.gz
2023-12-20 09:40:20       6305 vault1-audit-2023-12-13.log.gz
2023-12-20 09:40:20       6310 vault1-audit-2023-12-14.log.gz

The daily run is a script, vault-logs-s3-sync.sh, started from crontab at 07:00. Each node uploads its own files. The active node writes nearly all audit entries, so on most days four of the five have little or nothing new; the node name in the file name keeps them apart when leadership moves.

The bucket has two top-level directories, logs for this and snapshots for Raft snapshots, backup and restore. The storage has an endpoint in each availability zone; the script names the first.

Reading it today

← solutionz