Balabit 11 - Backup, archive and retention
Balabit SCB Solution · Previous: Audit trails: encryption and replay · Next: Connection and channel policies
An SCB that records every administrative session fills its disks, and the recordings are the reason it exists. If the appliance dies, its configuration (certificates, keys, policies, the AD binding) has to come back; if an incident is investigated half a year later, the audit trail of that night has to be found. This part describes how both clusters copied their configuration and their audit trails to the NetApp storage every night, how long the trails stayed on the appliance, what cleaned up the connection database, and how the protocol changed from CIFS to NFS on the way.
Three kinds of copy
The SCB has three mechanisms, and the design uses all of them. A system backup copies the configuration. A data backup copies the audit trails to a remote server and leaves them on the appliance. An archive/cleanup policy moves audit trails older than a retention time to a remote server and deletes them locally. Backup and archive policies are objects of their own (Policies > Backup & Archive/Cleanup); the system backup is chosen on Basic Settings > Management, and every connection names one backup and one archive policy.
| Copy | Policy | Starts | Export on 10.11.18.236 | Protection |
|---|---|---|---|---|
| Configuration | SYSTEM-BACKUP | 00:00 | DC1_S_VCVSM001_data/scb_system_backup | Encrypted with the GPG public key SCB-BACKUP |
| Trails of the organisation | ORG-BACKUP | 01:00 | DC1_S_VCVSM001_data/scb_org_backup | The trails are already encrypted and signed by the audit policy |
| Trails of partner 1 | PARTNER1-BACKUP | 02:00 | DC1_S_VCVSM001_data/scb_partner1_backup | as above |
| Trails of partner 2 | PARTNER2-BACKUP | 03:00 | DC1_S_VCVSM001_data/scb_partner2_backup | as above |
| Archive, organisation | ORG-ARCHIVE | 04:00 | DC1_S_VCVSM001_data/scb_org_archive | as above, 90 days on the SCB |
| Archive, partner 1 | PARTNER1-ARCHIVE | 06:00 | DC1_S_VCVSM001_data/scb_partner1_archive | as above, 90 days on the SCB |
| Archive, partner 2 | PARTNER2-ARCHIVE | 06:00 | DC1_S_VCVSM001_data/scb_partner2_archive | as above, 90 days on the SCB |
The values are from chapter 7 of my design. The policies are Config documents: Backup policies, site 1 and Archive/cleanup policies, site 1.
The configuration backup
SYSTEM-BACKUP writes the configuration of the cluster at midnight. The system backup page switches on "encrypt the configuration" and holds the public half of a GPG key pair made for nothing else: SCB-BACKUP, RSA 2048, comment "SCB backup encryption", the mail address scb_mail@ad.example.net, no expiry. The page itself is in Management, site 1 and the key in XCA and GPG keys.
Why encrypt it: the configuration is the box. The site 2 config.xml shows what is in it, with the secrets replaced by the appliance when it built a support bundle: the private keys of the internal CA, the TSA and the web server, the private key that signs audit trails, the LDAP bind password and the root password. As I understand it, a system backup carries the same file with those secrets in it, and the NFS export is readable by anyone who is root on a host the export rule admits. With GPG only the holder of the private key can read it. My notes keep the key's parameters and its passphrase (a placeholder here); where the private key itself was kept is not written down.
Data backup per partner
Every night at 01:00, 02:00 and 03:00 the trails of the organisation, partner 1 and partner 2 were copied, each to its own export. The design's chapter 4.5.1 lists encryption with SCB-AUDIT-ENCRYPT, timestamping and signing with SCB-AUDIT-SIGN for these backups; the backup policy page has no such fields, and the protection comes from the audit policy, which encrypts and signs each trail while it is being written (Audit trails: encryption and replay). What lands on the NetApp is ciphertext.
Why one policy per company: the design does not give a reason. My reading is that a separate qtree per company makes each company's recordings a unit that can be restored, handed over, kept or deleted without touching the others, and that the one-hour steps keep the runs from competing for the same NFS link. The cost is that every new partner needs two more policies, two more qtrees and a new line in the export plan, and the documents show that this was not always done (below).
All policies notify only on errors.
Archive, cleanup and retention
At 04:00 and 06:00 the three archive/cleanup policies moved trails older than 90 days to the _archive exports, in a tree of archive date, protocol and connection, and deleted them from the SCB. The design's section "Audit trails cleanup" states the same 90 days.
A second cleanup concerns the connection database. Archiving moves the files, but the metadata of each connection (who, when, from where, which channels) stays in the SCB's database and keeps it growing. The design sets the database cleanup "for every protocol" to 180 days; in the as-built pages it is the "channel database cleanup" of the SSH and RDP global options, 180 days, with the per-connection field left empty. As I read the two settings together, a connection can be found in the search for half a year, and its recording is on the appliance for the first three months of that.
The last safety net is disk fill-up prevention: when the disks are 80 % used the SCB disconnects clients, and "automatically start archiving" is off. As I read it, a full disk would therefore stop the administrators' work rather than start an unscheduled archive run. Both values are on the management page and are the same in the site 2 config.xml.
What happens on the archive exports after the 90 days is not defined anywhere. The policies only write; my notes have a "TODO - Retention policy" with two questions, how long to keep a session record for a quick audit and how long to archive it, and no answer.
The nightly order in both sites:
flowchart TB subgraph s1["Site 1, design chapter 7"] a0["00:00 SYSTEM-BACKUP"] a1["01:00 ORG-BACKUP"] a2["02:00 PARTNER1-BACKUP"] a3["03:00 PARTNER2-BACKUP"] a4["04:00 ORG-ARCHIVE"] a6["06:00 PARTNER1-ARCHIVE and PARTNER2-ARCHIVE"] a0 --> a1 --> a2 --> a3 --> a4 --> a6 end subgraph s2["Site 2, config.xml of 2018-09-17"] b0["00:00 SYSTEM-BACKUP"] b1["01:00 ORG-BACKUP"] b2["02:00 PARTNER1-BACKUP"] b3["03:00 PARTNER2-BACKUP"] b4["04:00 PARTNER4-BACKUP"] b5["04:30 ORG-ARCHIVE"] b6["05:00 PARTNER1-ARCHIVE"] b7["05:30 PARTNER2-ARCHIVE"] b8["06:00 PARTNER4-ARCHIVE"] b0 --> b1 --> b2 --> b3 --> b4 --> b5 --> b6 --> b7 --> b8 end
The arrows show the order of the start times, not a dependency: as I understand it, each policy has its own timer, and nothing makes a run wait for the previous one.
From CIFS to NFS
The protocol is where the documents disagree with each other. The pre-design analysis asked which protocol to use and recorded three answers: rsync at first, to a server called "archive"; SMB/CIFS with the NetApp storage as the target solution, which I recommended; NFS listed without comment. The hand-drawn HA sketch in the analysis has all three on the line from port 3 to the "Archive" cloud. My notes kept one sentence, apparently from a release announcement, that made CIFS look like the natural choice: the new release "supports the NetApp CIFS" so that audit trails "can be natively backed up and archived to NetApp NAS storage appliances". The conceptual and logical design pictures (the lines from the SCB to the network attached storage are labelled "CIFS") and the text of chapters 3.5 and 4.5 ("CIFS shared directory") still say so.
The tables of the same chapter 4.5 and everything in chapter 7 say NFS, and so does the site 2 configuration. The change happened on the storage side: the SCB could not mount the SMB share of the SVM and the protocol was switched to NFSv3. That story, with the error messages and the one NFS setting the SVM needed, is told in SVMs and NFS and not repeated here.
The network path stayed as designed. The analysis wanted backup traffic on port 3 for the time being and later on the two SFP+ ports 5 and 6; the build kept it on port 3, as the tagged VLAN 1020 "BCK" (10.11.18.224/28) next to the redundant heartbeat. The cluster's address there, 10.11.18.225 (DNS name dc1-s-xbla001.adm.example.net), moves with the master, so the export rule on the NetApp needs only that one address. The NFS LIF 10.11.18.236 is in the same subnet; the design's communication matrix has no row for this traffic, which I read as: nothing between the two to open. The interfaces are in Hardware, cabling and networks.
flowchart LR scb["SCB cluster dc1-s-xblb001, BCK 10.11.18.225, VLAN 1020"] lif["NFS LIF 10.11.18.236, SVM DC1-S-VCVSM001"] subgraph vol["Volume DC1_S_VCVSM001_data"] q0["scb_system_backup"] q1["scb_org_backup"] q2["scb_partner1_backup"] q3["scb_partner2_backup"] q4["scb_org_archive"] q5["scb_partner1_archive"] q6["scb_partner2_archive"] q7["scb_partner3_backup and scb_partner3_archive"] end scb -- "NFSv3" --> lif lif -- "SYSTEM-BACKUP" --> q0 lif -- "ORG-BACKUP" --> q1 lif -- "PARTNER1-BACKUP" --> q2 lif -- "PARTNER2-BACKUP" --> q3 lif -- "ORG-ARCHIVE" --> q4 lif -- "PARTNER1-ARCHIVE" --> q5 lif -- "PARTNER2-ARCHIVE" --> q6 lif -. "no policy documented" .-> q7
Site 2
The site 2 cluster has no design document, but its exported configuration of 2018-09-17 holds the policies as the appliance wrote them: Backup and archive policies in config.xml (site 2). They target 10.12.18.236, the site 2 twin of the NetApp SVM, with the same export names under DC2_S_VCVSM001_data. The system backup page refers to SYSTEM-BACKUP and has configuration encryption on with a GPG public key; disk fill-up prevention is 80 % without automatic archiving.
Three things are new. Partner 4, whose connections are only in site 2, has PARTNER4-BACKUP at 04:00 and PARTNER4-ARCHIVE at 06:00, and its RDP connection references them. The archives start at half-hour steps from 04:30, so no two run at the same minute any more. And each policy carries fields the site 1 tables do not show: whether the notification lists the files (yes for backups, no for archives), a limit of 10240 files in it, and the path template as a number, 2 for three archives and 5 for partner 4's. My documents do not translate the numbers, so partner 4's archive layout is unknown.
Loose ends
- CIFS in the conceptual and logical design (pictures and text of chapters 3.5 and 4.5), NFS in the tables of 4.5 and in chapter 7. The analysis recommended SMB/CIFS after rsync.
- The third archive table of chapter 4.5.2 is headed "for PARTNER1 connections" but writes to
scb_partner2_archive, and it starts at 07:00; chapter 7 hasPARTNER2-ARCHIVEat 06:00, in the same minute asPARTNER1-ARCHIVE. - Chapter 4.5.2 describes the directory structure as "Connection Date/Protocol/Connection/"; every table says "Archive date / Protocol / Connection".
- Partner 3 had connections in site 1 (the connection summary marks them configured) and the NetApp has
scb_partner3_backupandscb_partner3_archive, but no partner 3 policy appears in chapter 7 or in the site 2 file. Which backup and archive policies the partner 3 and cloud team connections used is not recorded. - Site 2 has policies for partner 2, while the connection summary's partner 2 sheet is empty; the site 2
PARTNER2_RDP_JumpServerexists in theconfig.xml. - In the menu checklist of my notes, "Backup & Archive/Cleanup" and "System backup" are marked
-, as are Syslog and Mail settings, while LDAP Server, Trusted CA Lists and SSL certificate are+; the notes do not say what the marks meant. - No restore of a system backup and no read-back of an archived trail is recorded, in either site.
Checked against One Identity Safeguard for Privileged Sessions 9.0
| As built | Today |
|---|---|
| SCB 5 LTS on T-10 appliances | 5.0.x discontinued 2020-05-28; T-Series end of support 2024-07-31 |
| NFS after SMB failed | Rsync, SMB/CIFS and NFS still offered; NFS up to version 4, detected automatically |
| GPG-encrypted configuration backup | Unchanged; system backups contain no audit-trail data |
| Path template "Archive date / Protocol / Connection" | Not among the five templates of 9.0, while the design's prose "Connection Date/Protocol/Connection/" is |
| 90 days retention, 180 days database cleanup, 80 % disk limit | All still exist; the disk limit must be 50 to 98 % |
The details are in the Config documents of this part.
What I would do differently
- Answer the retention TODO before go-live: how long a trail stays on the SCB, how long on the archive, and who deletes it there. The 90 days were set; the rest was never decided.
- Restore the configuration backup once, with the GPG private key from where it is really kept, and open one archived trail from the export. Without that, "encrypted backup" means "a file nobody has proved can be read".
- Add the backup and archive policies to the checklist for a new partner. Partner 3 shows what happens otherwise.