NetApp 05 - Storage design
NetApp Solution · Previous: Network design · Next: SVMs and NFS
A site has 272 disks in twelve shelves, half of them in each datacenter. This article is about who owns which disk, how the disks become mirrored aggregates, and what volumes sit on them. The layout is lopsided on purpose: one cluster owns almost everything, the other lives on a few disks of the shortest shelf. The numbers are those of site 1 (DC1) in the as-built report of 2018, with the changes of the final design noted.
Shelves and disks
| Datacenter | Shelf IDs | Disks | Seen by cluster A as | Seen by cluster B as |
|---|---|---|---|---|
| A | 10, 11 | 24 + 24 | 2.10, 2.11 | 4.10, 4.11 |
| A | 20, 21 | 24 + 24 | 3.20, 3.21 | 8.20, 8.21 |
| A | 30, 31 | 24 + 16 | 1.30, 1.31 | 5.30, 5.31 |
| B | 50, 51 | 24 + 24 | 6.50, 6.51 | 1.50, 1.51 |
| B | 60, 61 | 24 + 24 | 5.60, 5.61 | 2.60, 2.61 |
| B | 70, 71 | 24 + 16 | 4.70, 4.71 | 3.70, 3.71 |
The disks are 10K SAS drives that ONTAP reports as 1.2TB_SAS_10k, with 1.09 TB usable each. A disk is named stack, shelf, bay: 2.10.13 is bay 13 of shelf 10. The shelf ID is set on the shelf and is the same for everybody. The stack number in front of it is assigned by each cluster for itself, which is why the same shelf has two names in the table. It did not stay the same over time either: the table is from September 2017, and a later version of the layout spreadsheet has 7.10, 5.30, 4.50 and 1.70. In the working notes the short shelves appear under four names, 1.31, 5.31, 4.71 and 3.71, each with the remark "16 disks".
Both clusters see all twelve shelves, because the shelves hang on the two fabrics and not on a controller. See Physical design and cabling.
Pools and ownership
A mirrored aggregate needs disks in two places. ONTAP keeps them apart with pools. For every node, pool 0 holds the disks in its own datacenter and pool 1 the disks in the other one. An aggregate has one plex in each pool, and SyncMirror writes to both plexes synchronously.
The design spreadsheet gives every shelf a pool, seen from the cluster that owns it:
| Shelves | Owner | Pool for the owner |
|---|---|---|
| 10, 11, 20, 21, 30 | Nodes of cluster A | 0, local |
| 50, 51, 60, 61, 70 | Nodes of cluster A | 1, remote |
| 71, 11 of its 16 disks | Nodes of cluster B | 0, local |
| 31, 11 of its 16 disks | Nodes of cluster B | 1, remote |
| 31, the other 5 disks | Node 1 of cluster A | 0, local |
| 71, the other 5 disks | Node 1 of cluster A | 1, remote |
The spreadsheet writes the last four rows as pool 1/0 for shelf 31 and 0/1 for shelf 71: a shelf with two owners that look at it from opposite sides.
The as-built report counts it up per node:
| Node | Disks owned | In aggregates | Spare |
|---|---|---|---|
DC1-A-ANAS001 | 130 | 58 | 72 |
DC1-A-ANAS002 | 120 | 14 | 106 |
DC1-B-ANAS001 | 16 | 14 | 2 |
DC1-B-ANAS002 | 6 | 4 | 2 |
Cluster A owns 250 of the 272 disks of the site, cluster B owns 22. The ten full shelves are split evenly between the two nodes of cluster A, 60 disks each in each pool.
This is not what the system looked like after its first setup. A listing from September 2017 shows the two short shelves owned completely by cluster B, with a data aggregate on each of its nodes. In the final layout node 2 of cluster B has no data aggregate, and five disks of each short shelf have gone to node 1 of cluster A. The commands of that rearrangement were not kept in the notes. The notes file on disk assignment holds two links to the NetApp documentation and one sysconfig -A, followed by material that belongs to the directory integration; the file on physical storage holds the four shelf names above. What can be shown is the state before and the state after.
The commands that were used to look at ownership are in the notes, under a heading that translates as "hmm, questions". On cluster A:
$ storage shelf show $ storage disk option show $ storage disk show -ownership $ storage disk show -ownership -container-name DC1_A_ANAS001_root $ storage disk show -container-type spare $ storage disk show -container-type aggregate $ storage aggregate show -disk
Aggregates
| Aggregate | Owner | RAID | Disks | Per plex | Size | Holds |
|---|---|---|---|---|---|---|
DC1_A_ANAS001_root | DC1-A-ANAS001 | RAID4 | 4 | 2 | 953.80 GB | vol0 of the node |
DC1_A_ANAS001_data1 | DC1-A-ANAS001 | RAID-DP | 54 | 27 | 21.42 TB | KVM and backup SVMs |
DC1_A_ANAS002_root | DC1-A-ANAS002 | RAID4 | 4 | 2 | 953.80 GB | vol0 of the node |
DC1_A_ANAS002_data1 | DC1-A-ANAS002 | RAID-DP | 10 | 5 | 2.79 TB | Balabit SVM |
DC1_B_ANAS001_root | DC1-B-ANAS001 | RAID4 | 4 | 2 | 953.80 GB | vol0 of the node |
DC1_B_ANAS001_data1 | DC1-B-ANAS001 | RAID-DP | 10 | 5 | 2.79 TB | Root volume of DC1-S-VCVSM004, metadata |
DC1_B_ANAS002_root | DC1-B-ANAS002 | RAID4 | 4 | 2 | 953.80 GB | vol0 of the node |
Every aggregate is mirrored, normal. The root aggregates have the high-availability policy cfo, the data aggregates sfo. The maximum RAID group size is 16 for the data aggregates and 8 for the root aggregates.
A root aggregate is two plexes of two whole disks each, one data and one parity. The notes contain a reading list on advanced disk partitioning, which would have put the root on slices of shared disks, and nothing else on the subject. It was never an option: NetApp's documentation lists root-data partitioning as not supported in a fabric-attached MetroCluster. The as-built state shows whole disks: sixteen disks of the site hold nothing but four copies of vol0 and their mirrors, and vol0 fills each root aggregate to 95 per cent.
The large aggregate has 27 disks per plex. With a group size of 16 that is one full RAID group and one of 11, and after two parity disks per group 23 data disks, which matches the 21.42 TB. The final design lists the aggregate with 64 disks, 32 per plex and so two full groups; site 2 has the same aggregate with 38.
Why node 1 holds almost everything
The Source material states the layout and not the reasoning. What the layout itself shows:
- The site is run active/passive. All SVMs that serve clients are on cluster A. Cluster B has one SVM of its own,
DC1-S-VCVSM004, which exists only to tunnel administrator logins to Active Directory and has a root volume of 1 GB. - Inside cluster A the workloads are not balanced either. The KVM datastore, more than ten terabytes, is on node 1. Node 2 has a 100 GB volume for the Balabit SCB.
- A cluster in a MetroCluster needs a data aggregate even when it serves nothing, because MetroCluster keeps its own metadata volumes on one. Cluster B has exactly one, with the ten disks the first setup gave it.
- Nothing is lost by the imbalance. 178 of the 250 disks of cluster A are spares, 106 of them owned by node 2, so either node can grow an aggregate or get a second one without touching ownership.
The layout on the shelves
The design has a spreadsheet with one cell per disk, coloured by aggregate. Counted per shelf:
flowchart LR subgraph fa["Datacenter A: plex 0 of cluster A, mirror plexes of cluster B"] s10["Shelf 10: data1 of A1 x3, root of A1 x1, data1 of A2 x2, root of A2 x1, spare x17"] s11["Shelf 11: data1 of A1 x7, root of A1 x1, data1 of A2 x2, root of A2 x1, spare x13"] s20["Shelf 20: data1 of A1 x6, data1 of A2 x1, spare x17"] s21["Shelf 21: data1 of A1 x5, spare x19"] s30["Shelf 30: data1 of A1 x4, spare x20"] s31["Shelf 31: data1 of B1 x5, root of B1 x2, root of B2 x2, spare of B x2, data1 of A1 x2, spare of A x3"] end subgraph fb["Datacenter B: mirror plexes of cluster A, plex 0 of cluster B"] s50["Shelf 50: data1 of A1 x5, root of A1 x1, data1 of A2 x1, root of A2 x2, spare x15"] s51["Shelf 51: data1 of A1 x3, root of A1 x1, data1 of A2 x3, spare x17"] s60["Shelf 60: data1 of A1 x4, data1 of A2 x1, spare x19"] s61["Shelf 61: data1 of A1 x6, spare x18"] s70["Shelf 70: data1 of A1 x5, spare x19"] s71["Shelf 71: data1 of B1 x5, root of B1 x2, root of B2 x2, spare of B x2, data1 of A1 x4, spare of A x1"] end fa == "SyncMirror over the two fabrics" === fb
A1 and A2 are the nodes of cluster A, B1 and B2 those of cluster B. The plex names in the spreadsheet follow the same split:
| Aggregate | Plex in datacenter A | Plex in datacenter B |
|---|---|---|
DC1_A_ANAS001_data1, DC1_A_ANAS002_data1 | plex0 | plex1 |
DC1_A_ANAS001_root, DC1_A_ANAS002_root | plex0 | plex4 |
DC1_B_ANAS001_data1 | plex1 | plex0 |
DC1_B_ANAS001_root, DC1_B_ANAS002_root | plex4 | plex0 |
Two things in the picture are not tidy, and both have a history.
The disks of an aggregate are not next to each other. The root aggregates and the data aggregate of node 2 sit in single bays spread over shelves 10, 11 and 20, and their mirrors over 50, 51 and 60. The listing of September 2017 has the same bays under the names the first setup gave them, root_DC1_A_ANAS001 and so on. They were renamed and left where they were.
The two plexes of the large aggregate are not symmetrical. In datacenter A it has 25 disks in the full shelves and 2 in shelf 31; in datacenter B, 23 and 4. Those are among the disks taken over from cluster B: several of the bays are ones the removed data aggregate of DC1-B-ANAS002 had used.
Volumes
Cluster A in the as-built report:
| Volume | SVM | Aggregate | Size | Snapshot policy |
|---|---|---|---|---|
vol0 | Node DC1-A-ANAS001 | DC1_A_ANAS001_root | 902.54 GB | None |
vol0 | Node DC1-A-ANAS002 | DC1_A_ANAS002_root | 902.54 GB | None |
MDV_CRS_<ID>_A | Cluster | DC1_A_ANAS001_data1 | 10 GB | none |
MDV_CRS_<ID>_B | Cluster | DC1_A_ANAS002_data1 | 10 GB | none |
DC1_S_VCVSM001_root | DC1-S-VCVSM001 | DC1_A_ANAS002_data1 | 1 GB | default |
DC1_S_VCVSM001_data | DC1-S-VCVSM001 | DC1_A_ANAS002_data1 | 100 GB | default |
DC1_S_VCVSM002_root | DC1-S-VCVSM002 | DC1_A_ANAS001_data1 | 1 GB | default |
DC1_S_VCVSM002_data | DC1-S-VCVSM002 | DC1_A_ANAS001_data1 | 10.10 TB | none |
DC1_S_VCVSM003_root | DC1-S-VCVSM003 | DC1_A_ANAS001_data1 | 1 GB | default |
DC1_S_VCVSM003_data | DC1-S-VCVSM003 | DC1_A_ANAS001_data1 | 100 GB | default |
Cluster B has the two vol0, the root volume DC1_S_VCVSM004_root of 1 GB, and its own pair of MDV_CRS volumes, both on DC1_B_ANAS001_data1 because there is no second data aggregate to put the other one on.
The MDV_CRS volumes are not created by an administrator. MetroCluster makes them when it is configured and keeps in them the metadata of its configuration replication, the mechanism by which the partner cluster learns about SVMs, interfaces, exports and policies it must be able to bring up after a switchover. The design document describes them as "root volume for cluster SVM, partition 1 and 2", which is not what they are but is why they appear in its volume table.
The final design adds the volumes of 2019 to DC1_A_ANAS001_data1: root and data volumes for DC1-S-VCVSM005 and DC1-S-VCVSM006, and DC1_S_VCVSM002_data2, described as a temporary data volume "to be removed after agreement with the virtualization team". Their sizes are not in the Source material.
The settings that are the same for every volume:
| Setting | Value |
|---|---|
| Space guarantee | volume, so thick provisioned |
| Deduplication and compression | Not enabled on any volume; the efficiency policies default and inline-only exist and are unused |
| Autosize | Disabled; the first thing to try on a full volume is set to volume_grow |
| Snapshot autodelete | Disabled |
| Encryption | NetApp Volume Encryption on the data volumes; see Encryption and certificates |
The snapshot policy default keeps six hourly, two daily and two weekly copies. The cluster also has default-1weekly, which keeps one weekly copy, and none. The KVM datastore is on none: the images of running virtual machines were not snapshotted on the storage. On the partner cluster the policies of a mirrored SVM appear with the suffix -DR.
The same state, as listings in the shape of the ONTAP commands, is a Config document: ONTAP: disk ownership, aggregates and volumes as built. How the volumes are exported is in SVMs and NFS.
Site 2
Site 2 (DC2) has the same shelves with the same IDs, the same four small aggregates, and the same split of the short shelves. The differences in the layout spreadsheet of site 2:
| Item | Site 1 | Site 2 |
|---|---|---|
| Large aggregate | 54 disks as built, 64 in the final design | DC2_A_ANAS001_data1, 38 disks |
| Mirror plex names | plex1, plex4 | plex1, plex6, plex8 |
| Disk states in the legend | In an aggregate or spare | Also "unassigned" |
| Stack numbers of the shelves | As in the first table | Different: 1.10, 7.20, 6.30, 2.50, 8.60, 3.70 |
MDV_CRS volumes of cluster A | One on each data aggregate | Both on DC2_A_ANAS001_data1 |
The higher plex numbers suggest that mirror plexes at site 2 were removed and created again more than once; the notes do not say when or why. The spreadsheet of site 2 also still labels its rows with the datacenters of site 1, and the aggregate table of site 2 in the design contains one row with a site 1 name. No as-built report of site 2 is in the Source material, so the counts of site 2 are those of the design and not of the running system.
Lessons
- Write down the rearrangement, not only the result. The step from the factory layout to the final one, with disks changing owner across clusters, is the one piece of this build for which there are two pictures and no commands.
- Leave the short shelf to the small cluster. Putting the passive cluster entirely on the two 16-disk shelves kept the ten full shelves uniform: one owner cluster, one pool each.
- Spares are a decision too. Two thirds of the disks were spare when the as-built report was made. That is capacity bought and mirrored but not provisioned, and the design nowhere says how many spares were meant to stay.