NetApp 02 - Logical design
NetApp Solution · Previous: Requirements and concept · Next: Physical design and cabling
The logical design names every building block and says how the blocks belong together: nodes into clusters, clusters into a MetroCluster, storage virtual machines onto clusters, and all of it into the infrastructure around. The later articles each take one block apart. This one is the map.
Naming
Every name follows the host naming plan of the organisation, and once the plan is known a name says where a thing is and what it does.
output 1 line
<SITE>-<Node>-<Type><Usage><DeviceID><Purpose><PurposeID>
| Part | Values | Meaning |
|---|---|---|
| Site | DC1, DC2 | The site |
| Node | A, B, S | Datacenter A, datacenter B, or stretched: something that can live in either |
| Type | A, VC, S, B, X | Appliance, virtual compute, switch, bridge, cluster |
| Usage | NAS, VSM, OCM, NTB | NetApp storage, storage virtual machine, OnCommand management software, Tiebreaker |
| DeviceID | 001, 002 | First, second |
| Purpose | M, CL, ICL | Optional. Out-of-band management, cluster link, intercluster link |
| PurposeID | 01, 02 | Optional. First port, second port |
So DC1-A-ANAS002 is the second storage appliance in datacenter A of site 1, DC1-A-XNAS001 is the cluster those appliances form, DC1-B-SNAS001 is the first storage switch in datacenter B, DC1-A-BNAS001M02 is the second management port of the first bridge, and DC1-S-VCVSM002 is a storage virtual machine that can run in either datacenter.
In-band names live in adm.example.net and out-of-band names in mgmt.example.net. A node therefore has two names: dc1-a-anas001.adm.example.net is its node management LIF, and dc1-a-anas001m.mgmt.example.net is its service processor.
ONTAP does not allow a hyphen in the name of an aggregate or a volume, so there the same names are written with underscores: DC1_A_ANAS001_data1, DC1_S_VCVSM002_data.
The components of one site
flowchart TB subgraph a["Datacenter A"] subgraph xa["Cluster DC1-A-XNAS001"] a1["DC1-A-ANAS001"] a2["DC1-A-ANAS002"] end sa["FC switches DC1-A-SNAS001, 002"] ba["Bridges DC1-A-BNAS001, 002"] da["6 shelves, 136 disks"] xa --- sa sa --- ba ba --- da end subgraph b["Datacenter B"] subgraph xb["Cluster DC1-B-XNAS001"] b1["DC1-B-ANAS001"] b2["DC1-B-ANAS002"] end sb["FC switches DC1-B-SNAS001, 002"] bb["Bridges DC1-B-BNAS001, 002"] db["6 shelves, 136 disks"] xb --- sb sb --- bb bb --- db end sa == "ISL, fabric 1 and fabric 2" === sb xa -. "intercluster LIFs, cluster peering" .- xb
| Component | Model | Per datacenter | Names in site 1, datacenter A |
|---|---|---|---|
| Storage node | NetApp FAS8200 | 2, one HA pair | DC1-A-ANAS001, DC1-A-ANAS002 |
| Cluster | ONTAP | 1 | DC1-A-XNAS001 |
| Disk shelf | NetApp DS224C | 6 | Shelf IDs 10, 11, 20, 21, 30, 31 |
| Disk | 2.5 inch SAS, 10,000 rpm | 136 | Five full shelves of 24 and one of 16 |
| FC switch | Brocade 6505 | 2, one per fabric | DC1-A-SNAS001, DC1-A-SNAS002 |
| FC-to-SAS bridge | ATTO FibreBridge 7500N | 2 | DC1-A-BNAS001, DC1-A-BNAS002 |
| Data switch | Cisco Nexus 9396PX | 2, existing | DC1-A-SPRO001, DC1-A-SPRO002 |
| Out-of-band switch | Cisco Catalyst 2960-X | 2, existing | DC1-A-SOOB002, DC1-A-SOOB003 |
| Console router | Cisco 2901 | 1, existing | DC1-A-ROOB003 |
Datacenter B has the same set with B in the name and shelf IDs 50, 51, 60, 61, 70, 71. Site 2 has the same set again with DC2. Shelf IDs are unique across the two datacenters because after a switchover one cluster sees all twelve shelves.
The console routers reach only the FC switches. They had no free lines for the controllers or the bridges, so the last resort for a controller is its service processor, and for a bridge its serial port and somebody standing next to it.
Physical design and cabling follows every cable.
Nodes, HA pairs and clusters
A FAS8200 chassis holds two controllers. The two form an HA pair: each logs the other's writes in its own NVRAM and can take over the other's disks and addresses. In this design the HA pair is also the whole cluster, connected back to back over two 10 GbE twinax cables, with no cluster switches.
| Property | Value |
|---|---|
| Controller | FAS8200, Intel Xeon D-1587, 128 GB memory |
| Flash Cache | 2,048 GB NVMe per node |
| Service processor | One per node, in the out-of-band network |
| Licence | ONTAP 9 Premium bundle: all protocols, SnapMirror, SnapRestore, FlexClone, Volume Encryption |
| ONTAP at build, site 1 | 9.1, patch P2 at installation and P11 from February 2018 |
| ONTAP at the end, both sites | 9.3P12, according to version 0.2 of the design |
A cluster of one HA pair protects against the loss of a controller. It does not protect against the loss of the room.
The MetroCluster
The two clusters of a site form a MetroCluster. Three things are mirrored.
- Data. Every aggregate has two plexes: one on the shelves of its own datacenter, one on the shelves of the other. A write is acknowledged when both have it. This is SyncMirror, and it runs over the two FC fabrics.
- NVRAM. Each node mirrors its write log to its HA partner and to its disaster-recovery partner in the other cluster, over FC-VI ports on the same fabrics.
- Configuration. Each data SVM, with its volumes, LIFs, export policies and users, is replicated to the other cluster over the cluster peering network, where it sits stopped under the same name with the suffix
-mc.
In normal operation both clusters serve data. In a switchover the surviving cluster starts the -mc copies of the other's SVMs on its own ports, with the same addresses, from the local plexes. Clients see a pause, not a different server. MetroCluster switchover and Tiebreaker shows it happening.
Storage virtual machines
Six data SVMs exist per site. Five belong to cluster A and one to cluster B, and every one of them has its stopped twin on the other side.
flowchart LR subgraph ca["DC1-A-XNAS001, datacenter A"] a1["VCVSM001, active"] a2["VCVSM002, active"] a3["VCVSM003, active, in AD"] a4["VCVSM004-mc, replica"] a5["VCVSM005, active"] a6["VCVSM006, active"] end subgraph cb["DC1-B-XNAS001, datacenter B"] b1["VCVSM001-mc, replica"] b2["VCVSM002-mc, replica"] b3["VCVSM003-mc, replica"] b4["VCVSM004, active, in AD"] b5["VCVSM005-mc, replica"] b6["VCVSM006-mc, replica"] end a1 -- "mirror" --> b1 a2 -- "mirror" --> b2 a3 -- "mirror" --> b3 b4 -- "mirror" --> a4 a5 -- "mirror" --> b5 a6 -- "mirror" --> b6
| SVM | Home cluster | Serves | Protocol |
|---|---|---|---|
DC1-S-VCVSM001 | A | Balabit SCB: configuration backups, backups and archives of audit trails | NFSv3 |
DC1-S-VCVSM002 | A | KVM hosts of the management infrastructure | NFSv3, v4, v4.1 |
DC1-S-VCVSM003 | A | Backup and archive, planned. Active Directory logins to cluster A | CIFS server only |
DC1-S-VCVSM004 | B | Active Directory logins to cluster B | CIFS server only |
DC1-S-VCVSM005 | A | KVM hosts of platform C | NFSv3, v4, v4.1 |
DC1-S-VCVSM006 | A | KVM hosts of platform B | NFSv3, v4, v4.1 |
Two things in this table need a sentence each.
All the data sits on cluster A. Cluster B owns one SVM with a root volume and nothing else, so in normal operation the controllers of datacenter B do no client work at all. They mirror. The design document does not argue for this layout, it only records it. Its effect is plain enough: there is one place where data lives and one direction of switchover that matters, and half of the controller capacity is idle until the day it is needed.
VCVSM003 and VCVSM004 are members of the Active Directory domain. Neither serves a single share. An ONTAP cluster cannot join a domain itself; it authenticates domain users by passing the request through a data SVM that has a CIFS server in the domain, a domain tunnel. Each cluster needs its own, because in a switchover the other cluster's tunnel is not there. VCVSM003 was already planned for backups on cluster A, so it does both jobs; VCVSM004 exists on cluster B for this purpose alone. Access and directory integration has the details, and SVMs and NFS the data SVMs.
Besides the data SVMs every cluster has its admin SVM, which carries the cluster's own name and the cluster management LIF, and one node SVM per node. Those are what an administrator logs in to.
Management software
| Host | Software | Runs in | Watches |
|---|---|---|---|
DC1-A-VCOCM001 | OnCommand Unified Manager 9.4 | Site 1 | Both clusters of site 1 |
DC1-A-VCOCM002 | OnCommand API Services, in a systemd-nspawn container | Site 1 | Both clusters of site 1, for Sensu |
DC2-A-VCOCM001 | OnCommand Unified Manager | Site 2 | Both clusters of site 2 |
DC1-A-VCNTB001 | MetroCluster Tiebreaker 1.21P2 | Site 1 | The MetroCluster of site 2 |
DC2-A-VCNTB001 | MetroCluster Tiebreaker 1.21P2 | Site 2 | The MetroCluster of site 1 |
All of them are RHEL 7 virtual machines on the KVM cluster of the management infrastructure, which stores its disks on this very storage. For Unified Manager that is an inconvenience: when the storage is down, so is the tool that would say so. For a Tiebreaker it would be a design error, because a Tiebreaker that dies with the datacenter it watches cannot act. That is why each Tiebreaker runs in the other site.
What the storage depends on
| Service | Provided by | Names in site 1 | Used for |
|---|---|---|---|
| Directory | Active Directory on Windows Server 2016 and 2019 | dc1-a-vcad001, dc1-a-vcad002, and two more | Administrator and operator logins |
| DNS and NTP | Infoblox appliances in HA | dc1-a-adnsvip001, dc1-b-adnsvip001 | Name resolution, time |
| Syslog | Central syslog server | dc1-s-xcsys001 | EMS events of the clusters, switch logs |
| Monitoring | Sensu, with RabbitMQ and Graphite | dc1-a-vcsns001 and others | SNMPv3 polling, checks through Unified Manager |
| Internal mail relay | dc1-a-vcmsx001 | AutoSupport and alert mail | |
| HTTP proxy | Internal proxy | dc1-a-vcprx001 | AutoSupport over HTTPS to the vendor |
| TACACS+ | Internal AAA servers | dc1-a-vcrad001, dc1-b-vcrad001 | Logins to the FC switches |
| Certificates | Internal CA | HTTPS certificates of the clusters and of Unified Manager |
Three certificates per site are signed by the internal CA: one for each cluster and one for Unified Manager. The data SVMs keep the self-signed certificates ONTAP gives them, and the bridges have none.
Site 2 has its own DNS names, syslog and monitoring servers, but when it was built it used the domain controllers of site 1. The design notes that as a temporary state.
Network design lists the networks, the addresses and every flow that had to be opened on the firewalls between them.
Software versions
| Component | Version | Note |
|---|---|---|
| ONTAP | 9.1 at build, 9.3P12 at the end | ONTAP upgrade has the one upgrade the notes record |
| Brocade Fabric OS | 8.0.1 | With the NetApp reference configuration file, version 9.1 |
| ATTO FibreBridge firmware | 2.85 | |
| OnCommand Unified Manager | 9.4 | On RHEL 7.5 |
| MetroCluster Tiebreaker | 1.21P2 | On RHEL 7.5 |