NetApp 03 - Physical design and cabling
NetApp Solution · Previous: Logical design · Next: Network design
A fabric-attached MetroCluster is mostly cabling. The controllers never touch a disk shelf directly: every disk is reached through two Fibre Channel fabrics and a pair of FC-to-SAS bridges, and the same fabrics carry the mirroring between the two datacenters. This article is the hardware of one datacenter, where every cable goes, and what was configured on the switches and bridges. It also records the place where the kept switch configuration file and the cabling tables disagree.
What stands in one datacenter
Both datacenters of a site hold the same set of devices. Site 1 (DC1) is the worked example.
| Device | Count | Model | Names in datacenter A | Software |
|---|---|---|---|---|
| Storage controllers, one HA pair | 2 | NetApp FAS8200 | DC1-A-ANAS001, DC1-A-ANAS002 | ONTAP 9.1P2 at installation, 9.1P11 in the as-built report, 9.3P12 at the end |
| FC switches | 2 | Brocade 6505 | DC1-A-SNAS001, DC1-A-SNAS002 | Fabric OS 8.0.1 |
| FC-to-SAS bridges | 2 | ATTO FibreBridge 7500N | DC1-A-BNAS001, DC1-A-BNAS002 | Firmware 2.85 |
| Disk shelves | 6 | NetApp DS224C with IOM12 modules | Shelf IDs 10, 11, 20, 21, 30, 31 | Shelf firmware 0202 and 0210 |
| Out-of-band switches | 2 | Cisco Catalyst 2960-X | DC1-A-SOOB002, DC1-A-SOOB003 | Left as ??? in the design |
| Console router | 1 | Cisco 2901 | DC1-A-ROOB003 | IOS 15.4 |
| Data-centre switches, shared | 2 | Cisco Nexus 9396PX | DC1-A-SPRO001, DC1-A-SPRO002 | NX-OS 7.0 |
Datacenter B has the same with B in the names and shelf IDs 50, 51, 60, 61, 70, 71. The Catalyst, Nexus and console devices were not bought for the storage; they are the existing network of the management infrastructure, and the storage took ports on them.
The hardware of one HA pair, from the design document and the as-built report:
| Item | Value |
|---|---|
| Controller | FAS8200, two nodes in one chassis, redundant power supplies |
| Processor | 1 x Intel Xeon D-1587, 1.70 GHz |
| Memory | 128 GB |
| Flash Cache | 2 x 1 TB NVMe M.2 per node in slots 3 and 4, 4096 GB per HA pair |
| FC adapter | Quad-port 16 Gb FC adapter in slot 1, ports 1a to 1d; slot 2 empty |
| Disks | 136 x SFF 10K SAS per datacenter; five shelves with 24 and one with 16 |
| Disk size | The design says 1 TB; ONTAP reports them as 1.2TB_SAS_10k with 1.09 TB usable |
| Brocade 6505 | One power supply each: the design lists redundant power as "No" |
| FibreBridge 7500N | Redundant power supplies, two management ports |
Six shelves and 136 disks per datacenter do not mean 136 disks per cluster. How the disks of both datacenters are divided between the two clusters is the subject of Storage design.
The port faces
The design carried three pictures of port faces: the rear of the FAS8200, the vendor's photograph of the Brocade 6505 and the connector side of the FibreBridge. Redrawn as one diagram with what each port was used for:
flowchart TB subgraph c["FAS8200 controller, each of the two"] c1["e0a, e0b: cluster interconnect, twinax"] c2["0e, 0f: UTA2 ports in FC mode, FC-VI"] c3["1a to 1d: 16 Gb FC HBA in slot 1"] c4["e0g, e0h: 10 GbE optical, data and management"] c5["e0M: 1 GbE, service processor"] c6["0a to 0d SAS, e0c, e0d 10GBase-T, serial: not cabled"] end subgraph s["Brocade 6505"] s1["FC ports 0 to 23, 12 licensed by default"] s2["Management, RJ-45, 10/100 Mb"] s3["Serial, RJ-45"] s4["USB"] end subgraph b["FibreBridge 7500N"] b1["FC1, FC2: 16 Gb FC"] b2["SAS A to D: 12 Gb, four lanes each"] b3["Ethernet management 1 and 2"] b4["Serial, RJ-45"] end c --- s --- b
Two things about the controller are worth a sentence. The ports 0e and 0f are unified target adapter ports: they can be 10 GbE or 16 Gb FC, and here they run as FC and carry FC-VI, the interconnect over which the two clusters mirror their NVRAM. And the on-board SAS ports stay empty, because in this topology no shelf is attached to a controller.
On cluster B during the installation in September 2017, when the node names still carried an M suffix that was dropped later:
$ metrocluster interconnect adapter showoutput 9 lines
Adapter Link Node Adapter Name Type Status IP Address Port Number -------------- --------------- ------- ------ ----------- ----------- DC1-B-ANAS001M fcvi_device_0 FC-VI Up 2.0.0.8 0f DC1-B-ANAS001M fcvi_device_1 FC-VI Up 1.0.0.7 0e DC1-B-ANAS001M gop_0 GOP Up 192.0.1.4 ic0a DC1-B-ANAS002M fcvi_device_0 FC-VI Up 2.0.3.8 0f DC1-B-ANAS002M fcvi_device_1 FC-VI Up 1.0.3.7 0e DC1-B-ANAS002M gop_0 GOP Up 192.0.1.5 ic0a
How a fabric MetroCluster is wired
The whole of site 1, without port numbers:
flowchart TB subgraph fa["Datacenter A, cluster DC1-A-XNAS001"] a1["DC1-A-ANAS001"] a2["DC1-A-ANAS002"] as1["DC1-A-SNAS001"] as2["DC1-A-SNAS002"] ab1["DC1-A-BNAS001"] ab2["DC1-A-BNAS002"] ash[("Shelves 10 to 31")] a1 -- "cluster interconnect" --- a2 a1 --> as1 a1 --> as2 a2 --> as1 a2 --> as2 as1 --> ab1 as1 --> ab2 as2 --> ab1 as2 --> ab2 ab1 -- "SAS" --> ash ab2 -- "SAS" --> ash end subgraph fb["Datacenter B, cluster DC1-B-XNAS001"] b1["DC1-B-ANAS001"] b2["DC1-B-ANAS002"] bs1["DC1-B-SNAS001"] bs2["DC1-B-SNAS002"] bb1["DC1-B-BNAS001"] bb2["DC1-B-BNAS002"] bsh[("Shelves 50 to 71")] b1 -- "cluster interconnect" --- b2 b1 --> bs1 b1 --> bs2 b2 --> bs1 b2 --> bs2 bs1 --> bb1 bs1 --> bb2 bs2 --> bb1 bs2 --> bb2 bb1 -- "SAS" --> bsh bb2 -- "SAS" --> bsh end as1 == "fabric 1, two ISLs" === bs1 as2 == "fabric 2, two ISLs" === bs2
There are two fabrics. Switch 1 of datacenter A and switch 1 of datacenter B form fabric 1; the two switches numbered 2 form fabric 2. The fabrics are not connected to each other anywhere. Every controller and every bridge has one leg in each, so a switch, a bridge, an inter-switch link or a whole fabric can be lost without losing a disk. The design's drawings call the fabrics A and B and colour them red and blue; the switch configuration files call them FAB1 and FAB2. I use the numbers, because A and B are already taken by the datacenters.
Because the bridges and the shelves hang on fabrics that span both datacenters, every controller sees all twelve shelves of the site. That is what makes the mirror possible: a node in datacenter A writes one half of each aggregate to shelves next to it and the other half to shelves in datacenter B, over the inter-switch links.
Controllers, switches and bridges in datacenter A
flowchart LR n1["DC1-A-ANAS001"] n2["DC1-A-ANAS002"] s1["DC1-A-SNAS001, fabric 1"] s2["DC1-A-SNAS002, fabric 2"] br1["DC1-A-BNAS001"] br2["DC1-A-BNAS002"] r1["DC1-B-SNAS001"] r2["DC1-B-SNAS002"] n1 -- "e0a to e0a, e0b to e0b, twinax" --- n2 n1 -- "0e to 0, 1a to 1, 1c to 2" --> s1 n1 -- "0f to 0, 1b to 1, 1d to 2" --> s2 n2 -- "0e to 3, 1a to 4, 1c to 5" --> s1 n2 -- "0f to 3, 1b to 4, 1d to 5" --> s2 s1 -- "6 to FC1" --> br1 s1 -- "7 to FC1" --> br2 s2 -- "6 to FC2" --> br1 s2 -- "7 to FC2" --> br2 s1 == "8 to 8, 9 to 9" === r1 s2 == "8 to 8, 9 to 9" === r2
The cluster interconnect between the two nodes of an HA pair is two twinax cables, port to port, with no switch: a two-node switchless cluster. Everything else in the picture is 16 Gb short-wave FC optics. The same as a table, which is how it was handed to the people pulling cables:
| Switch port | On DC1-A-SNAS001, fabric 1 | On DC1-A-SNAS002, fabric 2 | Port type |
|---|---|---|---|
| 0 | Node 1, 0e, FC-VI | Node 1, 0f, FC-VI | F-Port |
| 1 | Node 1, 1a, HBA | Node 1, 1b, HBA | F-Port |
| 2 | Node 1, 1c, HBA | Node 1, 1d, HBA | F-Port |
| 3 | Node 2, 0e, FC-VI | Node 2, 0f, FC-VI | F-Port |
| 4 | Node 2, 1a, HBA | Node 2, 1b, HBA | F-Port |
| 5 | Node 2, 1c, HBA | Node 2, 1d, HBA | F-Port |
| 6 | Bridge 1, FC1 | Bridge 1, FC2 | F-Port |
| 7 | Bridge 2, FC1 | Bridge 2, FC2 | F-Port |
| 8 | ISL to DC1-B-SNAS001 port 8, trunk master | ISL to DC1-B-SNAS002 port 8, trunk master | LE E-Port |
| 9 | ISL to DC1-B-SNAS001 port 9, trunk slave | ISL to DC1-B-SNAS002 port 9, trunk slave | LE E-Port |
| 10, 11 | Not used | Not used | None |
Datacenter B is cabled identically, with its own names. Each node therefore has three ports in each fabric: one FC-VI port for the mirroring of NVRAM and two HBA ports for disk I/O. Each bridge has one port in each fabric.
Bridges and shelves
flowchart LR br1["DC1-A-BNAS001"] br2["DC1-A-BNAS002"] s10["Shelf 10"] s11["Shelf 11"] s20["Shelf 20"] s21["Shelf 21"] s30["Shelf 30"] s31["Shelf 31, 16 disks"] br1 -- "SAS A to IOM A port 1" --> s10 s10 -- "IOM A 3 to 1, IOM B 3 to 1" --> s11 s11 -- "IOM B port 3 to SAS A" --> br2 br1 -- "SAS B to IOM A port 1" --> s20 s20 -- "IOM A 3 to 1, IOM B 3 to 1" --> s21 s21 -- "IOM B port 3 to SAS B" --> br2 br1 -- "SAS C to IOM A port 1" --> s30 s30 -- "IOM A 3 to 1, IOM B 3 to 1" --> s31 s31 -- "IOM B port 3 to SAS C" --> br2
The six shelves are three stacks of two. The shelf ID says where a shelf is: the first digit is the stack, the second the position in it. Bridge 1 enters each stack at the top, on I/O module A of the first shelf. Bridge 2 enters at the bottom, on I/O module B of the last shelf. Inside the stack both module chains are daisy-chained from port 3 to port 1 of the next shelf. A stack is therefore reachable from either end, through either bridge, and each bridge is in both fabrics. Connector D of the bridges is not used; the bridge reports it as disabled. The SAS cables are 2 m passive copper.
In datacenter B the stacks are 50 and 51, 60 and 61, 70 and 71.
The switch configuration
NetApp publishes a reference configuration file (RCF) for every supported switch model and every position in the MetroCluster. Four such files were kept with the project, one per switch of site 1. They are complete Fabric OS configuration uploads of about 700 lines; all but a handful of those lines are the factory defaults of the switch. The lines that matter:
| What | Datacenter A, switch 1 | Switch 2 | Datacenter B, switch 1 | Switch 2 |
|---|---|---|---|---|
| File | ..._FAB1_SW1_D5_RCF_v9.1 | ..._FAB2_SW2_D6_RCF_v9.1 | ..._FAB1_SW3_D7_RCF_v9.1 | ..._FAB2_SW4_D8_RCF_v9.1 |
| Switch name in the file | Brcd6505-FAB1-SW1-D5 | Brcd6505-FAB2-SW2-D6 | Brcd6505-FAB1-SW3-D7 | Brcd6505-FAB2-SW4-D8 |
| Domain ID | 5 | 6 | 7 | 8 |
| Zoning | Carried in this file | Carried in this file | None in the file | None in the file |
The domain IDs are what make the two switches of one fabric distinguishable, and the zones are written as pairs of domain and port. The file sets defzone:noaccess, so a port that is in no zone talks to nobody, and then enables one zone configuration with seventeen zones: one for FC-VI, marked for high-priority quality of service, and sixteen that each put the HBA ports of all four controllers together with one bridge port. The sections of the file that define the installation, and how the four files differ, are a Config document: Brocade 6505: reference configuration file, fabric 1, switch 1.
Where the file and the cabling disagree
The zones in the kept files do not describe the cabling above.
| Port role | Kept RCF, version 9.1 | Cabling tables and diagrams of the design |
|---|---|---|
| FC-VI ports | 0, 1, 4, 5 | 0 and 3 |
| HBA ports | 2, 3, 6, 7 | 1, 2, 4, 5 |
| Bridge ports | 8 to 15 | 6 and 7 |
| Inter-switch links | Not in a zone; ports 8 to 11 and 20 to 23 carry a different port flag from the rest | 8 and 9 |
The cabling side is confirmed twice by the working notes. In the Tiebreaker test at site 2 the inter-switch links were cut by disabling ports 8 and 9 on both switches of a datacenter. And the FC-VI adapter addresses in the listing above end in 7 and 8, the domain IDs of the two switches in datacenter B, with 0 for node 1 and 3 for node 2 in the third position, which are the switch ports the tables give to FC-VI.
So the switches ran with the port assignment of the cabling tables, and with the domain IDs of the files. What their active zoning looked like is not in the Source material: there is no switchshow, cfgshow or zoneshow output anywhere in the notes, and nothing records how or when a file was loaded. NetApp's documentation says that these files go with the port layout of ONTAP 9.1 and later. A likely explanation is therefore that the switches were set up for an older port assignment, and the 9.1 files were downloaded and kept beside the documentation without being what the switches ran. I cannot prove it. The Config document shows the file as kept and says so.
On top of the RCF
What was configured on the switches by hand is in the notes, and belongs to other articles:
- Management address and DNS name on the out-of-band network, for example
dc1-a-snas001m.mgmt.example.net. See Network design. - Login through TACACS+ with the local database as fallback, because Fabric OS could only bind to LDAP anonymously and the directory did not allow that. See Access and directory integration.
- SNMP community strings changed from the factory ones, and the central syslog server. See Logging, monitoring and AutoSupport.
ONTAP itself watches the four switches of a site by SNMP. In the as-built report they appear as Brocade_10.11.15.45 and so on, with the host names as symbolic names, model Brocade6505, firmware v8.0.1 and status ok.
The bridge configuration
A FibreBridge has very little to configure: two FC ports, four SAS connectors, two Ethernet management ports and SNMP. Its limits shaped the design more than its settings did. It has no HTTPS, so its management interface is plain HTTP. It cannot use LDAP, RADIUS or TACACS+, so it has one local account. Both are written into the design as constraints.
| Setting | Value on DC1-A-BNAS001 |
|---|---|
| FC ports 1 and 2 | Up, point-to-point, 16 Gb |
| SAS connectors A, B, C | Up, 12 Gb, four lanes each |
| SAS connector D | Disabled |
| Management port 1 | 10.11.15.47, mask 255.255.255.0, gateway 10.11.15.254, DHCP disabled |
| Management port 2 | 10.11.15.48, same mask and gateway |
| DNS server | Not configured |
| SNMP | Enabled, no trap destinations |
| What it sees | 136 disks and 6 shelf enclosure devices |
The settings come from a dumpconfiguration of each bridge, which is mostly an event log. The settings alone, and the differences between the four bridges, are a Config document: ATTO FibreBridge 7500N: settings.
ONTAP monitors the bridges over SNMP as well. On cluster A:
$ storage bridge showoutput 6 lines
Bridge Symbolic Name Monitored Status Vendor Model ------------------------ ------------- --------- ------- ------ ----------------- ATTO_10.11.15.47 bridgeA1 true ok Atto FibreBridge 7500N ATTO_10.11.15.49 bridgeA2 true ok Atto FibreBridge 7500N ATTO_10.11.23.47 bridgeB1 true ok Atto FibreBridge 7500N ATTO_10.11.23.49 bridgeB2 true ok Atto FibreBridge 7500N
Each cluster monitors all four bridges and all four switches of the site, not only its own. The firewall has to allow that across the datacenters; the flow tables of the design only listed the local ones.
Out-of-band and console
Every device has a management port on the out-of-band network, VLAN 12, spread over the two Catalyst switches so that losing one of them does not take away access to both devices of a pair.
| Device | Port | DC1-A-SOOB002 | DC1-A-SOOB003 |
|---|---|---|---|
| Node 1 | e0M | Gi1/0/23 | Not connected |
| Node 2 | e0M | Not connected | Gi1/0/23 |
| FC switch 1 | Management | Gi1/0/24 | Not connected |
| FC switch 2 | Management | Not connected | Gi1/0/24 |
| Bridge 1 | Management 1 and 2 | Gi1/0/25 | Gi1/0/25 |
| Bridge 2 | Management 1 and 2 | Gi1/0/26 | Gi1/0/26 |
All of them are access ports in VLAN 12. The bridge is the only device with two management ports, and it uses both.
Serial consoles were planned for everything: the first cabling spreadsheet sends the console of the controllers, switches and bridges to the Cisco 2901. The design document then keeps only the two FC switches, on asynchronous lines 29 and 30, and says why: the console router had no more free lines. The controllers have their service processor instead, which gives a console over SSH. The bridges have nothing but their Ethernet ports.
The data ports e0g and e0h of each node go to the two Nexus switches, one cable to each. They are the subject of the next article.
Site 2
Site 2 (DC2) was built a year later from the same parts list, and the design states that the cabling drawings of site 1 apply to it unchanged. Names start with DC2, shelf IDs are again 10 to 31 and 50 to 71, bridge firmware is again 2.85. The Tiebreaker notes of site 2 confirm ports 8 and 9 as the inter-switch links there. No switch configuration files and no bridge dumps of site 2 are in the Source material.
Lessons
- Keep what the switch says, not only what was sent to it. Four vendor files were archived and not one
switchshoworcfgshow. Two years later the files and the cabling tables contradict each other and nothing settles it. Aconfiguploadtaken from each switch after the build would have. - One document, one truth. The design names the console router
DC1-A-ROOB003in the component list andDC1-A-ROOB002in the cabling tables, repeats a node name of datacenter A in the cabling table of datacenter B, and titles the Brocade hardware table "FibreBridge". None of it hurt, because the drawings were right, and the drawings were what people used. - The 16-disk shelf is not an accident. 136 is 5 x 24 + 16. The short shelf is the last one of the third stack in both datacenters, and it is the one the smaller cluster lives on. See Storage design.