NetApp 08 - Encryption and certificates
NetApp Solution · Previous: Access and directory integration · Next: Logging, monitoring and AutoSupport
Two kinds of cryptography were configured on the clusters: encryption of the data at rest, and TLS for the management interface. Both are a handful of commands. Both also leave something behind that has to be looked after for as long as the system lives: a passphrase and a key backup in the first case, an expiry date in the second.
Why NVE and not NSE
ONTAP has two technologies for data at rest. NetApp Storage Encryption (NSE) uses self-encrypting disks, which encrypt everything written to them and release it only to a node that authenticates with a key. NetApp Volume Encryption (NVE) is software: each volume gets its own XTS-AES-256 key, and data, Snapshot copies and metadata of that volume are encrypted before they reach the aggregate.
The choice was not made on merit. Constraint C1 of the design says that in a MetroCluster configuration only NVE can be used. The design gives no further reasoning, and no comparison of the two was made.
NVE needs three things: an ONTAP build that contains it, a licence and a key manager. The build was checked first, on DC1-A-XNAS001:
$ version -voutput 1 line
NetApp Release 9.1P8: Wed Aug 30 13:33:41 UTC 2017 <1O>
A build without encryption carries the marker 1no-DARE in that line; this one does not. The licence is ve, volume encryption, one per node; it is listed in the as-built report next to the protocol licences. The notes do not show it being installed, so it either came with the systems or was added in System Manager.
The Onboard Key Manager on a MetroCluster
The volume keys have to be stored somewhere. The alternatives are an external KMIP server or the Onboard Key Manager (OKM), which keeps the keys on the cluster itself, wrapped with a key derived from a cluster-wide passphrase. The design chose OKM, enabled on both clusters; an external key server is not mentioned anywhere.
On a MetroCluster the wizard is run once on each cluster, and the same passphrase must be entered on both. After a switchover the surviving cluster serves the volumes of the other one and needs their keys.
$ security key-manager setupoutput 4 lines
Would you like to configure onboard key management? {yes, no} [yes]: yes
Enter the cluster-wide passphrase for onboard key management. To continue the configuration, enter the
passphrase, otherwise type "exit": <KEY_MANAGER_PASSPHRASE>
Re-enter the cluster-wide passphrase: <KEY_MANAGER_PASSPHRASE>First on DC1-A-XNAS001, then on DC1-B-XNAS001. Afterwards each cluster was asked for its keys.
$ security key-manager key showoutput 14 lines
Node: DC1-A-ANAS001 Key Store: onboard Key ID Used By ---------------------------------------------------------------- -------- <KEY_ID> NSE-AK <KEY_ID> NSE-AK Node: DC1-A-ANAS002 Key Store: onboard Key ID Used By ---------------------------------------------------------------- -------- <KEY_ID> NSE-AK <KEY_ID> NSE-AK 4 entries were displayed.
Two keys per node, four entries per cluster, and the same picture on cluster B. The listing was taken right after setup, before any volume was encrypted.
What has to leave the system
Two things must be kept outside the storage system, and the procedure names both. The first is the passphrase. The second is the key manager backup:
$ security key-manager backup showThe command prints one long block of encoded text between a begin and an end marker. It contains the wrapped key hierarchy of the cluster and is what a recovery procedure asks for, together with the passphrase, when a node has lost its key database, for example after a boot media replacement. It is not reproduced here in any form. One without the other is useless, which is exactly the point of keeping them: the backup can be stored with the system documentation, the passphrase must not be stored next to it.
The site 1 notes state, twice, that the passphrase and the backup taken on cluster A are enough "because both clusters use the same keys and passphrase". The site 2 notes, written when that site was built, record the output of security key-manager backup show from both DC2-A-XNAS001 and DC2-B-XNAS001. The second practice is the safe one: a backup per cluster, taken again whenever keys change, costs nothing and removes the need to be sure about the first statement.
The notes are silent on where the two items were finally deposited. The original working notes contained both in clear text, which is how such material usually ends up in more places than intended; the instruction in the procedure, "manual copy to a secure location outside the storage system", is only as good as the place it names.
Encrypting the volumes
The data volumes existed before the key manager did. A listing of DC1_S_VCVSM002_data from October 2017 ends with Enable Encryption: false. In ONTAP 9.1 an existing volume is encrypted by moving it, and the destination of the move may be the aggregate it already lives on:
$ volume show -vserver DC1-S-VCVSM001 -volume DC1_S_VCVSM001_data -field aggregate $ volume move start -vserver DC1-S-VCVSM001 -volume DC1_S_VCVSM001_data -destination-aggregate DC1_A_ANAS002_data1 -encrypt-destination true $ volume move show -vserver DC1-S-VCVSM001 -volume DC1_S_VCVSM001_data -field percent-complete $ volume show -vserver DC1-S-VCVSM001 -volume DC1_S_VCVSM001_data -field is-encrypted
output 3 lines
vserver volume is-encrypted -------------- ------------------- ------------ DC1-S-VCVSM001 DC1_S_VCVSM001_data true
The move is a real move: a new, encrypted volume is built next to the old one, the data is copied, the clients are cut over without interruption and the source is deleted. For the 100 GB volume of the SCB there was room. For the KVM volume there was not:
output 1 line
Error: command failed: There is 2.77TB of available space on the aggregate DC1_A_ANAS001_data1 which is not enough to accommodate a volume.
An in-place move needs the size of the volume free in the aggregate a second time, and the volume was thick provisioned. The solution in the notes is one line: add disks to the aggregate. After that the same four commands ran for DC1_S_VCVSM002_data. The command set is a Config document: ONTAP: Onboard Key Manager and volume encryption.
| Volume | SVM | Encrypted |
|---|---|---|
DC1_S_VCVSM001_data | DC1-S-VCVSM001 | yes |
DC1_S_VCVSM002_data | DC1-S-VCVSM002 | yes |
DC1_S_VCVSM005_data | DC1-S-VCVSM005 | yes |
DC1_S_VCVSM006_data | DC1-S-VCVSM006 | yes |
SVM root volumes, DC1_S_VCVSM003_data, node root volumes | no |
"All data volumes" in the design means the four volumes that hold client data. The SVM root volumes hold only junction points, and the data volume of the tunnel SVM is empty. The notes cover the commands for the first two volumes only; for 005 and 006, which were created later, there is the design table and no command, so whether they were moved or created encrypted from the start cannot be said.
Site 2 lists the corresponding four volumes. Its note file is the site 1 procedure copied, still with the site 1 cluster names in it, and in front of it the three things that were really recorded for site 2: the passphrase, the key listing of both clusters and the key manager backup of both clusters. The design table for site 2 spells the volume names with underscores (DC2_S_VCVSM002_data), the volume table of the same document without (DC2SVCVSM002_data).
Certificates for the cluster management interface
After setup a cluster presents a self-signed certificate with its own name as the issuer. The design requires certificates of the internal CA on everything that has a management web interface and can take one: the two admin SVMs per site and Unified Manager. The bridges cannot (constraints C3 and C4).
The request is generated on the cluster, here DC1-A-XNAS001:
$ security certificate generate-csr -common-name dc1-a-xnas001.adm.example.net -size 2048 -country <COUNTRY_CODE> -locality <LOCALITY> -organization "Example Org" -email-addr admin01@example.net
The command prints two blocks, the signing request and the private key, and stores neither. Both have to be copied from the terminal at that moment: the request goes to the CA, the key is needed again at installation. A private key that passes through a terminal scrollback and a clipboard is the weak point of this procedure, and there is no way around it in this ONTAP version.
When the signed certificate came back, it was installed together with the key:
$ security certificate install -vserver DC1-A-XNAS001 -type serveroutput 3 lines
Please enter Certificate: Press <Enter> when done
Please enter Private Key: Press <Enter> when done
Do you want to continue entering root and/or intermediate certificates {y|n}: nThe question about the chain was answered with no, so the cluster serves the server certificate alone and clients must already trust the internal CA. Installing a certificate does not activate it. The cluster now holds two server certificates:
$ security certificate showoutput 11 lines
Vserver Serial Number Common Name Type
---------- --------------- -------------------------------------- ------------
DC1-A-XNAS001
<CERT_SERIAL> DC1-A-XNAS001 server
Certificate Authority: DC1-A-XNAS001
Expiration Date: Fri Aug 24 12:36:28 2018
DC1-A-XNAS001
<CERT_SERIAL> dc1-a-xnas001.adm.example.net server
Certificate Authority: Internal CA
Expiration Date: Thu Nov 15 00:59:59 2018The web server is pointed at the new one by CA, serial number and common name, and the result checked:
$ security ssl modify -vserver DC1-A-XNAS001 -ca "Internal CA" -serial <CERT_SERIAL> -common-name dc1-a-xnas001.adm.example.net -server-enabled true -client-enabled false $ security ssl show
The same was done on DC1-B-XNAS001 with its own name. The command set is a Config document: ONTAP: CA-signed certificate for the cluster. The certificate of Unified Manager is handled in Unified Manager and API Services.
Data SVMs
The data SVMs also have a web service and a certificate, although nobody logs in there. They kept self-signed certificates, replaced by new ones with the organisation's attributes and a lifetime of ten years:
$ security certificate create -common-name DC1-S-VCVSM001 -type server -size 2048 -country <COUNTRY_CODE> -locality <LOCALITY> -organization "Example Org" -email-addr admin01@example.net -expire-days 3652 -vserver DC1-S-VCVSM001 $ security ssl modify -server-enabled true -vserver DC1-S-VCVSM001 -common-name DC1-S-VCVSM001 -serial <CERT_SERIAL> -ca DC1-S-VCVSM001 $ security ssl show -vserver DC1-S-VCVSM001 -instance
The old certificate was deleted by serial number in between. The notes show this for SVMs 001, 003 and 004, and, dated 8 July 2019, for 002 and 006. There is no entry for 005.
One year
| Cluster | Certificate | Valid until |
|---|---|---|
DC1-A-XNAS001 | self-signed, from cluster setup | 24 August 2018 |
DC1-A-XNAS001, DC1-B-XNAS001 | internal CA, first | 15 November 2018 |
DC1-A-XNAS001, DC1-B-XNAS001 | internal CA, second | 4 January 2020 |
DC2-A-XNAS001 | internal CA | 3 July 2020 |
DC2-B-XNAS001 | internal CA | 28 June 2020 |
The internal CA issues for one year. Read from top to bottom the table is a renewal calendar with a hole in it: the first CA-signed certificates of site 1 ended on 15 November 2018 and their successors start on 3 January 2019. What the clusters presented in the seven weeks between is not recorded; a browser warning on a management interface is easy to click through, which is how an expired certificate stays unnoticed. The final design document is dated 29 November 2019, five weeks before the next expiry, and contains no renewal procedure.
A renewal is the whole procedure again, because generate-csr always creates a new key: request, signature, security certificate install, security ssl modify with the new serial number, and then deletion of the old certificate. Four clusters, three different dates.
TLS hardening
As delivered, the clusters offered TLS 1.0, 1.1 and 1.2 on the management interface with a wide cipher list. The setting is cluster-wide and lives at the advanced privilege level:
$ set advanced $ security config show
output 5 lines
Cluster Cluster Security
Interface FIPS Mode Supported Protocols Supported Ciphers Config Ready
--------- ---------- ----------------------- ----------------- ----------------
SSL false TLSv1.2, TLSv1.1, TLSv1 ALL:!LOW:!aNULL: yes
!EXP:!eNULLIt was reduced to TLS 1.2 with ephemeral key exchange and AES-GCM or AES-256:
$ security config modify -interface SSL -supported-protocols TLSv1.2 -supported-ciphers EECDH+AESGCM:EDH+AESGCM:AES256+EECDH:AES256+EDH
output 4 lines
Warning: When this command completes, reboot all nodes in the cluster. This is necessary to prevent components from failing due to an inconsistent
security configuration state in the cluster. To avoid a service outage, reboot one node at a time and wait for it to completely initialize
before rebooting the next node.
Do you want to continue? {y|n}: YThe warning turns a one-line change into a maintenance window: every node of every cluster has to be rebooted, one at a time. The note ends at the confirmation. It does not name the clusters it was run on, does not show the reboots and does not show security config show afterwards, so the state after hardening is documented only as intended. The command is kept as a Config document: ONTAP: TLS protocols and ciphers.
The setting covers the cluster's own HTTPS service. It does not make the bridges speak HTTPS, and SNMPv3 for monitoring was configured with SHA-1 and DES (see Logging, monitoring and AutoSupport); the weakest protocol on the management network was never the web interface of the clusters.
Lessons
- The passphrase and the key manager backup are the only two things that cannot be regenerated. They should be deposited in two named places before the first volume is encrypted, and the place written into the design; here it was written into a note file together with the secrets themselves.
- Take the key manager backup on both clusters and again after every change of keys. The assumption that one cluster's backup covers the other was never tested by a restore.
- An in-place
volume moveneeds the volume's size free in the aggregate. Encrypt before the volumes are filled or grown, or create them encrypted. - A one-year certificate on four clusters is four calendar entries, and the calendar must belong to a team, not a person. The gap at the end of 2018 is what happens otherwise.
- Changes that need a reboot of every node belong in the initial build, before there is anything on the cluster.