NetApp 01 - Requirements and concept
NetApp Solution · Next: Logical design
The management infrastructure of a network service provider needed one primary storage system: a place for the disks of its virtual machines, for the backups and audit trails of its session-recording appliances, and later for backups of everything else. This article says what that storage had to do, what it was not allowed to do, and how the pieces relate before any of them is configured.
What was built
The same design was built twice, about a year apart, in two sites that do not share storage. Site 1 (DC1) was built in the second half of 2017 and documented in version 0.1 of the design document in March 2018. Site 2 (DC2) followed during 2018 and 2019, and version 0.2 of the document, from November 2019, describes both.
Each site is a NetApp MetroCluster: two clusters in two datacenters of the same campus, each cluster a FAS8200 HA pair, each mirroring its data synchronously to the other over Fibre Channel. A site survives the loss of a controller, of a disk shelf, of a whole datacenter. It does not survive the loss of the site, and it was never meant to; nothing is replicated between DC1 and DC2. The only thing the sites do for each other is host one small virtual machine, the MetroCluster Tiebreaker, which watches the other site's two clusters from the outside.
A short vocabulary
The rest of the Solution uses ONTAP terms without stopping to explain them. These are the ones that matter here.
| Term | What it is |
|---|---|
| Node | One storage controller with its copy of ONTAP |
| HA pair | Two nodes in one chassis that can take over each other's disks and addresses |
| Cluster | One or more HA pairs managed as one system. Here a cluster is exactly one HA pair |
| Aggregate | RAID groups of disks owned by one node; the raw capacity |
| Plex | One copy of an aggregate. A mirrored aggregate has two plexes, and in a MetroCluster they are in different datacenters |
| SVM | Storage virtual machine, also called a Vserver: the thing a client connects to, with its own volumes, addresses, protocols and users |
| LIF | Logical interface: an IP address that belongs to an SVM or to the cluster and can move between ports and nodes |
| Volume, qtree | A FlexVol volume is a file system inside an aggregate; a qtree is a directory in it with its own security style and export rules |
| MetroCluster | Two clusters that mirror each other's aggregates and SVM configuration, so that either can serve everything |
| Switchover, switchback | One cluster taking over the other's SVMs, and giving them back |
| NVE, NSE | NetApp Volume Encryption, done in software per volume, and NetApp Storage Encryption, done by self-encrypting disks |
Functional requirements
The design document numbered them. The wording here is shortened; the numbers are the original ones.
| ID | Requirement | What it meant in practice |
|---|---|---|
| FR1 | Network file services | Clients create, read, change and delete files on the storage over the network |
| FR2 | Multiprotocol access, NFS and CIFS | Both protocols licensed and available. Only NFS ever carried data; CIFS was used for something else entirely, see below |
| FR3 | Encryption | Data at rest must be unreadable if a disk is repurposed, returned, misplaced or stolen |
| FR4 | Remote management protocols | SSH and HTTPS on everything that can speak them |
| FR5 | Secure web management | Management web interfaces reachable over HTTPS only |
| FR6 | Active Directory integration | Administrator and operator accounts and groups live in Active Directory, not on the devices |
| FR7 | Secure Active Directory integration | The devices talk to the directory over LDAPS |
| FR8 | Certification authority | Every certificate is signed by the internal CA of the management infrastructure |
The numbering has a history. Version 0.1 counted from FR1 to FR9 and skipped FR5. Version 0.2 closed the gap, and was left with an FR9 that repeats FR1 word for word.
Non-functional requirements
| ID | Requirement | How the design answers it |
|---|---|---|
| NR1 | Openness | Standard protocols only: NFS, SSH, HTTPS, LDAP, SNMP, syslog |
| NR2 | High availability | HA pair inside a datacenter, MetroCluster between datacenters, a Tiebreaker for automatic switchover |
| NR3 | Scalability | Shelves can be added to the stacks, SVMs and volumes are created on demand |
Constraints
The interesting part of the requirements is where the hardware said no. Four constraints were written down, and each of them shaped an article of this Solution.
| ID | Requirement it limits | Constraint |
|---|---|---|
| C1 | FR3, encryption | In a MetroCluster only NetApp Volume Encryption could be used. Self-encrypting disks were not an option, so encryption is software, per volume, with the Onboard Key Manager |
| C2 | FR6, FR7, directory | The Brocade FC switches could bind to LDAP only anonymously, which the Active Directory servers refuse. The switches were integrated through TACACS+ instead. The ATTO bridges support no central authentication at all: no LDAP, no RADIUS, no TACACS+ |
| C3 | FR5, HTTPS | The ATTO bridges have no HTTPS. Their web interface is plain HTTP |
| C4 | FR8, certificates | Following from C3, the bridges carry no certificate |
So the weakest device of the system is the one that sits between the controllers and every disk: it is managed over HTTP with one local account. The mitigation was network placement, not configuration. The bridges live only in the out-of-band management network, which is reachable from the administrators' VPN and from nothing else. Access and directory integration and Encryption and certificates come back to all four.
There was one written assumption, A1: everything is correctly licensed, operating systems of the management virtual machines included.
The concept
flowchart TB admins["Administrators and operators"] subgraph site["Management infrastructure, site 1"] subgraph mc["NetApp MetroCluster"] ca["Cluster A, datacenter A"] cb["Cluster B, datacenter B"] ca --- cb end ocum["OnCommand Unified Manager"] scb["Balabit SCB appliances"] kvm["KVM hosts, management"] bck["Backup and archive server"] sensu["Sensu"] ad["Active Directory"] dns["DNS and NTP"] syslog["Central syslog"] end subgraph other["Other platforms, site 1"] kvmb["KVM hosts, platform B"] kvmc["KVM hosts, platform C"] end subgraph s2["Management infrastructure, site 2"] tb["MetroCluster Tiebreaker"] end admins -- "SSH, HTTPS" --> mc admins -- "SSH, HTTPS" --> ocum ocum -- "HTTPS" --> mc scb -- "NFS" --> mc kvm -- "NFS" --> mc bck -- "NFS" --> mc kvmb -- "NFS" --> mc kvmc -- "NFS" --> mc sensu -- "SNMPv3" --> mc sensu -- "HTTPS" --> ocum mc -- "LDAPS" --> ad mc -- "DNS, NTP" --> dns mc -- "syslog" --> syslog tb -- "SSH" --> mc
Site 2 is the mirror image: its own MetroCluster, its own Unified Manager, its own clients, and a Tiebreaker that runs in site 1.
The storage
The MetroCluster is the primary storage of the management infrastructure. Its two clusters are both active. Each owns its SVMs and serves them from its own datacenter, and each holds a dormant copy of the other's.
The clients
Five groups of clients were planned, all of them over NFS.
| Client | What it stores | State at the end |
|---|---|---|
| KVM hosts of the management infrastructure | Disks of the virtual machines that run the infrastructure itself | In service |
| KVM hosts of platform B | Disks of that platform's virtual machines | In service, added in version 0.2 |
| KVM hosts of platform C | Disks of that platform's virtual machines | In service, added in version 0.2 |
| Balabit SCB appliances | Backups of the appliance configuration, backups and archives of recorded sessions | In service |
| Backup and archive server | Backups of the other systems | SVM created, never filled: the backup solution was not implemented while I worked on the system |
The first row hides a circular dependency worth knowing about. The virtual machines on that KVM cluster include the domain controllers, the monitoring servers and the Unified Manager that the storage itself depends on. The storage therefore had to be able to start, and be managed, with none of them running: local accounts as a last resort, addresses instead of names where it matters, and no boot-time dependency on the directory.
The management software
Three products were run on virtual machines.
- OnCommand Unified Manager watches both clusters of a site, raises events, and is where the operations team looks first.
- OnCommand API Services puts a REST interface in front of the clusters, for the monitoring system. It appears in version 0.1 of the design and is gone from version 0.2.
- MetroCluster Tiebreaker watches the two clusters of the other site and starts a switchover when one datacenter disappears. The design gives it a recovery time objective of 120 seconds.
Unified Manager and API Services and MetroCluster switchover and Tiebreaker describe them.
The infrastructure services
The storage consumes what the management infrastructure already offers: Active Directory for accounts, a pair of Infoblox appliances for DNS and NTP, a central syslog server, the Sensu monitoring system, a mail relay and an HTTP proxy for AutoSupport, and the internal CA.
Site 2 had no directory of its own when its storage was built. Its clusters were joined to the domain controllers of site 1, and the design document says plainly that four SVMs would have to be reconfigured once site 2 had its own. That reconfiguration is not in the Source material.
The users
Two roles exist, and they map to two Active Directory groups.
| Role | Rights | Used by |
|---|---|---|
| Administrators | Full | The storage team |
| Operators | Read-only | The operations team, and the monitoring system |
What the design left open
A design document is also a record of what was not finished. Version 0.2 still contains these, and the Solution does not pretend otherwise.
- The integration of Sensu with Unified Manager is a heading followed by "TODO".
- The backup and archive SVM has a section that says it will be completed when the backup solution is implemented.
- The chapter on physical cabling and the network was not carried over from version 0.1, so for site 2 there are spreadsheets but no reviewed text.
- Several tables for site 2 were copied from site 1 and not fully corrected. Where that matters, the articles use the as-built reports instead and say so.