LINUXOR.SK ... open source notes ...

NetApp 08 - Encryption and certificates

category: solutionz · date: 2019-12-31 · updated: 2026-10-02 · author: LALA

NetApp Solution · Previous: Access and directory integration · Next: Logging, monitoring and AutoSupport

Two kinds of cryptography were configured on the clusters: encryption of the data at rest, and TLS for the management interface. Both are a handful of commands. Both also leave something behind that has to be looked after for as long as the system lives: a passphrase and a key backup in the first case, an expiry date in the second.

Why NVE and not NSE

ONTAP has two technologies for data at rest. NetApp Storage Encryption (NSE) uses self-encrypting disks, which encrypt everything written to them and release it only to a node that authenticates with a key. NetApp Volume Encryption (NVE) is software: each volume gets its own XTS-AES-256 key, and data, Snapshot copies and metadata of that volume are encrypted before they reach the aggregate.

The choice was not made on merit. Constraint C1 of the design says that in a MetroCluster configuration only NVE can be used. The design gives no further reasoning, and no comparison of the two was made.

NVE needs three things: an ONTAP build that contains it, a licence and a key manager. The build was checked first, on DC1-A-XNAS001:

bash
$ version -v
output 1 line
NetApp Release 9.1P8: Wed Aug 30 13:33:41 UTC 2017 <1O>

A build without encryption carries the marker 1no-DARE in that line; this one does not. The licence is ve, volume encryption, one per node; it is listed in the as-built report next to the protocol licences. The notes do not show it being installed, so it either came with the systems or was added in System Manager.

The Onboard Key Manager on a MetroCluster

The volume keys have to be stored somewhere. The alternatives are an external KMIP server or the Onboard Key Manager (OKM), which keeps the keys on the cluster itself, wrapped with a key derived from a cluster-wide passphrase. The design chose OKM, enabled on both clusters; an external key server is not mentioned anywhere.

On a MetroCluster the wizard is run once on each cluster, and the same passphrase must be entered on both. After a switchover the surviving cluster serves the volumes of the other one and needs their keys.

bash
$ security key-manager setup
output 4 lines
Would you like to configure onboard key management? {yes, no} [yes]: yes
Enter the cluster-wide passphrase for onboard key management. To continue the configuration, enter the
passphrase, otherwise type "exit": <KEY_MANAGER_PASSPHRASE>
Re-enter the cluster-wide passphrase: <KEY_MANAGER_PASSPHRASE>

First on DC1-A-XNAS001, then on DC1-B-XNAS001. Afterwards each cluster was asked for its keys.

bash
$ security key-manager key show
output 14 lines
Node: DC1-A-ANAS001
Key Store: onboard
Key ID                                                           Used By
---------------------------------------------------------------- --------
<KEY_ID>                                                         NSE-AK
<KEY_ID>                                                         NSE-AK

Node: DC1-A-ANAS002
Key Store: onboard
Key ID                                                           Used By
---------------------------------------------------------------- --------
<KEY_ID>                                                         NSE-AK
<KEY_ID>                                                         NSE-AK
4 entries were displayed.

Two keys per node, four entries per cluster, and the same picture on cluster B. The listing was taken right after setup, before any volume was encrypted.

What has to leave the system

Two things must be kept outside the storage system, and the procedure names both. The first is the passphrase. The second is the key manager backup:

bash
$ security key-manager backup show

The command prints one long block of encoded text between a begin and an end marker. It contains the wrapped key hierarchy of the cluster and is what a recovery procedure asks for, together with the passphrase, when a node has lost its key database, for example after a boot media replacement. It is not reproduced here in any form. One without the other is useless, which is exactly the point of keeping them: the backup can be stored with the system documentation, the passphrase must not be stored next to it.

The site 1 notes state, twice, that the passphrase and the backup taken on cluster A are enough "because both clusters use the same keys and passphrase". The site 2 notes, written when that site was built, record the output of security key-manager backup show from both DC2-A-XNAS001 and DC2-B-XNAS001. The second practice is the safe one: a backup per cluster, taken again whenever keys change, costs nothing and removes the need to be sure about the first statement.

The notes are silent on where the two items were finally deposited. The original working notes contained both in clear text, which is how such material usually ends up in more places than intended; the instruction in the procedure, "manual copy to a secure location outside the storage system", is only as good as the place it names.

Encrypting the volumes

The data volumes existed before the key manager did. A listing of DC1_S_VCVSM002_data from October 2017 ends with Enable Encryption: false. In ONTAP 9.1 an existing volume is encrypted by moving it, and the destination of the move may be the aggregate it already lives on:

bash
$ volume show -vserver DC1-S-VCVSM001 -volume DC1_S_VCVSM001_data -field aggregate
$ volume move start -vserver DC1-S-VCVSM001 -volume DC1_S_VCVSM001_data -destination-aggregate DC1_A_ANAS002_data1 -encrypt-destination true
$ volume move show -vserver DC1-S-VCVSM001 -volume DC1_S_VCVSM001_data -field percent-complete
$ volume show -vserver DC1-S-VCVSM001 -volume DC1_S_VCVSM001_data -field is-encrypted
output 3 lines
vserver        volume              is-encrypted
-------------- ------------------- ------------
DC1-S-VCVSM001 DC1_S_VCVSM001_data true

The move is a real move: a new, encrypted volume is built next to the old one, the data is copied, the clients are cut over without interruption and the source is deleted. For the 100 GB volume of the SCB there was room. For the KVM volume there was not:

output 1 line
Error: command failed: There is 2.77TB of available space on the aggregate DC1_A_ANAS001_data1 which is not enough to accommodate a volume.

An in-place move needs the size of the volume free in the aggregate a second time, and the volume was thick provisioned. The solution in the notes is one line: add disks to the aggregate. After that the same four commands ran for DC1_S_VCVSM002_data. The command set is a Config document: ONTAP: Onboard Key Manager and volume encryption.

VolumeSVMEncrypted
DC1_S_VCVSM001_dataDC1-S-VCVSM001yes
DC1_S_VCVSM002_dataDC1-S-VCVSM002yes
DC1_S_VCVSM005_dataDC1-S-VCVSM005yes
DC1_S_VCVSM006_dataDC1-S-VCVSM006yes
SVM root volumes, DC1_S_VCVSM003_data, node root volumesno

"All data volumes" in the design means the four volumes that hold client data. The SVM root volumes hold only junction points, and the data volume of the tunnel SVM is empty. The notes cover the commands for the first two volumes only; for 005 and 006, which were created later, there is the design table and no command, so whether they were moved or created encrypted from the start cannot be said.

Site 2 lists the corresponding four volumes. Its note file is the site 1 procedure copied, still with the site 1 cluster names in it, and in front of it the three things that were really recorded for site 2: the passphrase, the key listing of both clusters and the key manager backup of both clusters. The design table for site 2 spells the volume names with underscores (DC2_S_VCVSM002_data), the volume table of the same document without (DC2SVCVSM002_data).

Certificates for the cluster management interface

After setup a cluster presents a self-signed certificate with its own name as the issuer. The design requires certificates of the internal CA on everything that has a management web interface and can take one: the two admin SVMs per site and Unified Manager. The bridges cannot (constraints C3 and C4).

The request is generated on the cluster, here DC1-A-XNAS001:

bash
$ security certificate generate-csr -common-name dc1-a-xnas001.adm.example.net -size 2048 -country <COUNTRY_CODE> -locality <LOCALITY> -organization "Example Org" -email-addr admin01@example.net

The command prints two blocks, the signing request and the private key, and stores neither. Both have to be copied from the terminal at that moment: the request goes to the CA, the key is needed again at installation. A private key that passes through a terminal scrollback and a clipboard is the weak point of this procedure, and there is no way around it in this ONTAP version.

When the signed certificate came back, it was installed together with the key:

bash
$ security certificate install -vserver DC1-A-XNAS001 -type server
output 3 lines
Please enter Certificate: Press <Enter> when done
Please enter Private Key: Press <Enter> when done
Do you want to continue entering root and/or intermediate certificates {y|n}: n

The question about the chain was answered with no, so the cluster serves the server certificate alone and clients must already trust the internal CA. Installing a certificate does not activate it. The cluster now holds two server certificates:

bash
$ security certificate show
output 11 lines
Vserver    Serial Number   Common Name                            Type
---------- --------------- -------------------------------------- ------------
DC1-A-XNAS001
           <CERT_SERIAL>   DC1-A-XNAS001                          server
    Certificate Authority: DC1-A-XNAS001
          Expiration Date: Fri Aug 24 12:36:28 2018

DC1-A-XNAS001
           <CERT_SERIAL>   dc1-a-xnas001.adm.example.net          server
    Certificate Authority: Internal CA
          Expiration Date: Thu Nov 15 00:59:59 2018

The web server is pointed at the new one by CA, serial number and common name, and the result checked:

bash
$ security ssl modify -vserver DC1-A-XNAS001 -ca "Internal CA" -serial <CERT_SERIAL> -common-name dc1-a-xnas001.adm.example.net -server-enabled true -client-enabled false
$ security ssl show

The same was done on DC1-B-XNAS001 with its own name. The command set is a Config document: ONTAP: CA-signed certificate for the cluster. The certificate of Unified Manager is handled in Unified Manager and API Services.

Data SVMs

The data SVMs also have a web service and a certificate, although nobody logs in there. They kept self-signed certificates, replaced by new ones with the organisation's attributes and a lifetime of ten years:

bash
$ security certificate create -common-name DC1-S-VCVSM001 -type server -size 2048 -country <COUNTRY_CODE> -locality <LOCALITY> -organization "Example Org" -email-addr admin01@example.net -expire-days 3652 -vserver DC1-S-VCVSM001
$ security ssl modify -server-enabled true -vserver DC1-S-VCVSM001 -common-name DC1-S-VCVSM001 -serial <CERT_SERIAL> -ca DC1-S-VCVSM001
$ security ssl show -vserver DC1-S-VCVSM001 -instance

The old certificate was deleted by serial number in between. The notes show this for SVMs 001, 003 and 004, and, dated 8 July 2019, for 002 and 006. There is no entry for 005.

One year

ClusterCertificateValid until
DC1-A-XNAS001self-signed, from cluster setup24 August 2018
DC1-A-XNAS001, DC1-B-XNAS001internal CA, first15 November 2018
DC1-A-XNAS001, DC1-B-XNAS001internal CA, second4 January 2020
DC2-A-XNAS001internal CA3 July 2020
DC2-B-XNAS001internal CA28 June 2020

The internal CA issues for one year. Read from top to bottom the table is a renewal calendar with a hole in it: the first CA-signed certificates of site 1 ended on 15 November 2018 and their successors start on 3 January 2019. What the clusters presented in the seven weeks between is not recorded; a browser warning on a management interface is easy to click through, which is how an expired certificate stays unnoticed. The final design document is dated 29 November 2019, five weeks before the next expiry, and contains no renewal procedure.

A renewal is the whole procedure again, because generate-csr always creates a new key: request, signature, security certificate install, security ssl modify with the new serial number, and then deletion of the old certificate. Four clusters, three different dates.

TLS hardening

As delivered, the clusters offered TLS 1.0, 1.1 and 1.2 on the management interface with a wide cipher list. The setting is cluster-wide and lives at the advanced privilege level:

bash
$ set advanced
$ security config show
output 5 lines
          Cluster                                              Cluster Security
Interface FIPS Mode  Supported Protocols     Supported Ciphers Config Ready
--------- ---------- ----------------------- ----------------- ----------------
SSL       false      TLSv1.2, TLSv1.1, TLSv1 ALL:!LOW:!aNULL:  yes
                                             !EXP:!eNULL

It was reduced to TLS 1.2 with ephemeral key exchange and AES-GCM or AES-256:

bash
$ security config modify -interface SSL -supported-protocols TLSv1.2 -supported-ciphers EECDH+AESGCM:EDH+AESGCM:AES256+EECDH:AES256+EDH
output 4 lines
Warning: When this command completes, reboot all nodes in the cluster. This is necessary to prevent components from failing due to an inconsistent
         security configuration state in the cluster. To avoid a service outage, reboot one node at a time and wait for it to completely initialize
         before rebooting the next node.
Do you want to continue? {y|n}: Y

The warning turns a one-line change into a maintenance window: every node of every cluster has to be rebooted, one at a time. The note ends at the confirmation. It does not name the clusters it was run on, does not show the reboots and does not show security config show afterwards, so the state after hardening is documented only as intended. The command is kept as a Config document: ONTAP: TLS protocols and ciphers.

The setting covers the cluster's own HTTPS service. It does not make the bridges speak HTTPS, and SNMPv3 for monitoring was configured with SHA-1 and DES (see Logging, monitoring and AutoSupport); the weakest protocol on the management network was never the web interface of the clusters.

Lessons

← solutionz