NetApp 04 - Network design
NetApp Solution · Previous: Physical design and cabling · Next: Storage design
Fibre Channel carries the disks and the mirror. Everything else is Ethernet: the clients, the administrators, the peering between the two clusters and the service processors. Each node has two 10 GbE ports for all of it. This article is how eight networks were put on those two ports, which addresses and names were given out, what the firewall had to let through, and which rows of the design document were copy-paste slips.
Networks and VLANs
| VLAN | Network | Purpose | Port on the node |
|---|---|---|---|
| None | 169.254.0.0/16 | Cluster interconnect between the two nodes of an HA pair | e0a, e0b |
| 12 | 10.11.15.0/24 in datacenter A, 10.11.23.0/24 in B | Out-of-band management: service processors, FC switches, bridges | e0M |
| 1022 | 10.11.10.0/27 | Intercluster: peering of the two clusters of the MetroCluster | a0a-1022 |
| 1024 | 10.11.10.32/28 | In-band management: nodes, clusters, Unified Manager, API Services | a0a-1024 |
| 1020 | 10.11.18.224/28 | Balabit SCB backup and archive, SVM DC1-S-VCVSM001 | a0a-1020 |
| 1026 | 10.11.10.64/27 | KVM storage, SVM DC1-S-VCVSM002 | a0a-1026 |
| 1028 | 10.11.10.48/29 | Backup infrastructure, SVM DC1-S-VCVSM003 | a0a-1028 |
| 1027 | 10.11.10.56/29 | Active Directory tunnel of cluster B, SVM DC1-S-VCVSM004 | a0a-1027 |
Only the out-of-band network differs between the datacenters. Every other network is one subnet stretched over both, and it has to be: after a switchover the logical interfaces of an SVM come up in the other datacenter with the same addresses.
The KVM network was not planned this way. The first address sheet put the NFS address of the storage into the hypervisors' own network, VLAN 1006, on one of its last two free addresses, with the note that a separate network for KVM storage was "another question". The question was answered with VLAN 1026, a network that holds only the hypervisors' storage interfaces and one address of the storage, and has no gateway.
The two SVMs added in 2019 for platform B and platform C have their own networks as well. Their addresses are in the final design (10.11.40.30 and 10.11.89.94); their VLAN IDs at site 1 are not in the Source material.
From physical port to SVM
flowchart BT subgraph ph["Physical ports of one node"] e0a["e0a"] e0b["e0b"] e0m["e0M"] a0a["a0a, LACP over e0g and e0h"] end subgraph vl["VLAN ports"] v1022["a0a-1022"] v1024["a0a-1024"] v1020["a0a-1020"] v1026["a0a-1026"] v1028["a0a-1028"] end subgraph lif["Logical interfaces"] cl["clus1, clus2"] sp["Service processor"] icl["intercluster_lif"] nm["Node management"] cm["cluster_mgmt"] d1["nfs_lif1 of VCVSM001"] d2["nfs_lif1 of VCVSM002"] d3["cifs_lif1 of VCVSM003"] end e0a --> cl e0b --> cl e0m --> sp a0a --> v1022 --> icl a0a --> v1024 v1024 --> nm v1024 --> cm a0a --> v1020 --> d1 a0a --> v1026 --> d2 a0a --> v1028 --> d3 cl --> n1["Twinax to the HA partner"] sp --> n2["VLAN 12"] icl --> n3["VLAN 1022"] nm --> n4["VLAN 1024"] cm --> n4 d1 --> n5["VLAN 1020, Balabit SCB"] d2 --> n6["VLAN 1026, KVM hypervisors"] d3 --> n7["VLAN 1028, backup"]
The drawing of datacenter B in the design is the same picture with a single data VLAN, 1027, for DC1-S-VCVSM004. The as-built report shows more than the drawings: both clusters have all six VLAN ports on both nodes. Cluster B needs the ports of 1020, 1026 and 1028 for the day it takes over the SVMs of cluster A, and cluster A needs 1027 for the reverse.
The ports of one node, from the as-built report:
| Port | Type | MTU | IPspace | Broadcast domain |
|---|---|---|---|---|
e0a, e0b | Physical, 10 GbE, twinax | 9000 | Cluster | Cluster |
e0g, e0h | Physical, 10 GbE optical, members of a0a | 9000 | Default | None |
a0a | Interface group, multimode_lacp, distribution by IP | 9000 | Default | None |
a0a-1022, a0a-1024 | VLAN | 9000 | Default | vlan-1022, vlan-1024 |
a0a-1020 | VLAN | 9000 | DC1-S-VCVSM001_ipspace | vlan-1020 |
a0a-1026 | VLAN | 9000 | DC1-S-VCVSM002_ipspace | vlan-1026 |
a0a-1028 | VLAN | 9000 | DC1-S-VCVSM003_ipspace | vlan-1028 |
a0a-1027 | VLAN | 9000 | DC1-S-VCVSM004_ipspace | vlan-1027 |
e0M | Physical, 1 GbE | 1500 | Default | Default |
e0c, e0d | Physical, 10GBase-T, link down | 1500 | Default | None |
Three decisions are visible in that table.
- One interface group for everything.
e0gande0hform a dynamic LACP group, one cable to each Nexus switch. Nothing has an address ona0aitself; every network is a VLAN port on top. - Node management is in-band. The node-management interfaces sit on
a0a-1024, beside the cluster-management interface. The copper porte0Mcarries only the service processor. Administration therefore depends on the data-centre switches, and the service processor is the way in when they are gone. - One IPspace per data SVM. Each data SVM got its own IPspace with one broadcast domain and one failover group, both named after the VLAN and both holding the VLAN port of node 1 and node 2. An SVM in its own IPspace has its own routing table, so each can have a default gateway of its own, or none.
The complete state, port by port and interface by interface, is a Config document: ONTAP: network state as built. It is a listing of state and not a command set, because the commands that created the interface groups, VLANs, broadcast domains and interfaces were not kept in the notes.
Logical interfaces
| Interface | Cluster A | Cluster B | Home port |
|---|---|---|---|
clus1, clus2 of each node | Four addresses in 169.254.0.0/16 | Four addresses in 169.254.0.0/16 | e0a, e0b |
cluster_mgmt | 10.11.10.33/28 | 10.11.10.34/28 | a0a-1024 on node 1 |
| Node management, node 1 | 10.11.10.37/28 | 10.11.10.39/28 | a0a-1024 |
| Node management, node 2 | 10.11.10.38/28 | 10.11.10.40/28 | a0a-1024 |
intercluster_lif_1, _2 on node 1 | 10.11.10.1/27, 10.11.10.2/27 | 10.11.10.5/27, 10.11.10.6/27 | a0a-1022 |
intercluster_lif_3, _4 on node 2 | 10.11.10.3/27, 10.11.10.4/27 | 10.11.10.7/27, 10.11.10.8/27 | a0a-1022 |
| Service processor, node 1 | 10.11.15.42/24 | 10.11.23.42/24 | e0M |
| Service processor, node 2 | 10.11.15.44/24 | 10.11.23.44/24 | e0M |
The data interfaces, one per SVM:
| SVM | Interface | Address | Home | State on the other cluster |
|---|---|---|---|---|
DC1-S-VCVSM001 | DC1-S-VCVSM001_nfs_lif1 | 10.11.18.236/28 | Cluster A, node 2, a0a-1020 | up/down under DC1-S-VCVSM001-mc |
DC1-S-VCVSM002 | DC1-S-VCVSM002_nfs_lif1 | 10.11.10.94/27 | Cluster A, node 1, a0a-1026 | up/down under DC1-S-VCVSM002-mc |
DC1-S-VCVSM003 | DC1-S-VCVSM003_cifs_lif1 | 10.11.10.49/29 | Cluster A, node 1, a0a-1028 | up/down under DC1-S-VCVSM003-mc |
DC1-S-VCVSM004 | DC1-S-VCVSM004_cifs_lif1 | 10.11.10.57/29 | Cluster B, node 2, a0a-1027 | up/down under DC1-S-VCVSM004-mc |
Each interface exists twice: once on its own cluster, up, and once on the partner cluster under the -mc copy of the SVM, administratively up and operationally down until a switchover. The home node of each NFS data interface is the node that owns the aggregate with the SVM's data, so client traffic does not cross the cluster interconnect in normal operation.
The default routes show which networks are routed:
| SVM | Default gateway |
|---|---|
| Cluster, admin SVM | 10.11.10.46 |
DC1-S-VCVSM001 | 10.11.18.238 |
DC1-S-VCVSM002 | None: the KVM storage network is not routed |
DC1-S-VCVSM003 | 10.11.10.54 |
DC1-S-VCVSM004 | 10.11.10.62 |
The intercluster peering was verified during the installation by pinging every intercluster address of the other cluster from each node. On cluster B, at a time when the node names still ended in M:
$ network ping -node DC1-B-ANAS001M -destination 10.11.10.1 $ cluster peer show -instance
Names
In-band names are in adm.example.net, out-of-band names in mgmt.example.net. Infoblox holds an A record and a PTR record for each.
| Name | Address | What |
|---|---|---|
dc1-a-xnas001.adm.example.net | 10.11.10.33 | Cluster management, cluster A |
dc1-b-xnas001.adm.example.net | 10.11.10.34 | Cluster management, cluster B |
dc1-a-anas001.adm.example.net, dc1-a-anas002 | 10.11.10.37, .38 | Node management, cluster A |
dc1-b-anas001.adm.example.net, dc1-b-anas002 | 10.11.10.39, .40 | Node management, cluster B |
dc1-a-anas001m.mgmt.example.net, dc1-a-anas002m | 10.11.15.42, .44 | Service processors, datacenter A |
dc1-b-anas001m.mgmt.example.net, dc1-b-anas002m | 10.11.23.42, .44 | Service processors, datacenter B |
dc1-s-vcvsm001.adm.example.net to dc1-s-vcvsm004 | 10.11.18.236, 10.11.10.94, 10.11.10.49, 10.11.10.57 | Data SVMs |
dc1-s-vcvsm005.adm.example.net, dc1-s-vcvsm006 | 10.11.40.30, 10.11.89.94 | Data SVMs for platform C and platform B |
dc1-a-snas001m.mgmt.example.net, dc1-a-snas002m | 10.11.15.45, .46 | FC switches, datacenter A |
dc1-b-snas001m.mgmt.example.net, dc1-b-snas002m | 10.11.23.45, .46 | FC switches, datacenter B |
dc1-a-bnas001m01.mgmt.example.net, m02 | 10.11.15.47, .48 | Bridge 1, datacenter A, both management ports |
dc1-a-bnas002m01.mgmt.example.net, m02 | 10.11.15.49, .50 | Bridge 2, datacenter A |
dc1-b-bnas001m01.mgmt.example.net, m02 | 10.11.23.47, .48 | Bridge 1, datacenter B |
dc1-b-bnas002m01.mgmt.example.net, m02 | 10.11.23.49, .50 | Bridge 2, datacenter B |
dc1-a-vcocm001.adm.example.net | 10.11.10.44 | OnCommand Unified Manager |
dc1-a-vcocm002.adm.example.net | 10.11.10.43 | OnCommand API Services |
Cluster and intercluster interfaces have no DNS records. The address sheet also reserved a full second set of addresses in every network for a second MetroCluster that was never built: nodes 003 and 004, cluster 002, switches and bridges 003 and 004. Two of those reserved in-band addresses, .43 and .44, are the ones the two management servers ended up with.
The switch side
| Switch | Port | Mode | VLANs |
|---|---|---|---|
DC1-A-SPRO001, Nexus 9396PX | Eth1/19 to node 1 e0g, Eth1/20 to node 2 e0g | Trunk, LACP | 1020, 1022, 1024, 1026, 1027, 1028 |
DC1-A-SPRO002, Nexus 9396PX | Eth1/19 to node 1 e0h, Eth1/20 to node 2 e0h | Trunk, LACP | 1020, 1022, 1024, 1026, 1027, 1028 |
DC1-A-SOOB002, Catalyst 2960-X | Gi1/0/23 to Gi1/0/26 | Access | 12 |
DC1-A-SOOB003, Catalyst 2960-X | Gi1/0/23 to Gi1/0/26 | Access | 12 |
Datacenter B uses the same port numbers on its own four switches. The MTU is 9000 on every switch port a node is connected to; the notes list that as the first thing to check when the two clusters disagree about an interface. The Source material has the port list handed to the network team and nothing of the Nexus configuration itself, so the port-channel numbers and how the two Nexus switches present one LACP partner are not recorded here.
What the firewall has to allow
The networks of the storage are behind the firewall of the management infrastructure. The design lists the flows in four groups. Datacenter A is shown; datacenter B is the same with its own addresses.
Between the components of a cluster:
| Source | Destination | Port | Purpose |
|---|---|---|---|
Cluster management 10.11.10.33 | FC switches 10.11.15.45, .46 | UDP 161 | ONTAP monitors the switches by SNMP |
Cluster management 10.11.10.33 | Bridges 10.11.15.47, .49 | UDP 161 | ONTAP monitors the bridges by SNMP |
Service processor of node 1 10.11.15.42 | Node management of node 2 10.11.10.38 | UDP 4444 | Hardware-assisted takeover |
Service processor of node 2 10.11.15.44 | Node management of node 1 10.11.10.37 | UDP 4444 | Hardware-assisted takeover |
From the administrators, who arrive from the admin VPN 10.11.20.0/25:
| Destination | Port | Purpose |
|---|---|---|
| Nodes and cluster, in-band | TCP 22, 443 | SSH and System Manager |
DC1-A-VCOCM001 | TCP 22, 2222, 443 | SSH to the host, SSH to the container, web interface |
DC1-A-VCOCM002 | TCP 22, 2222, 8443 | SSH to the host, SSH to the container, API and its administration |
| Service processors, out-of-band | TCP 22 | SSH |
| FC switches, out-of-band | TCP 22 | SSH |
| Bridges, both management ports | TCP 80 | HTTP, the only thing a bridge offers |
To and from the infrastructure services:
| Source | Destination | Port | Purpose |
|---|---|---|---|
| All components | Active Directory 10.11.16.209, .210 | TCP 88, 135, 139, 389, 445, 464, 636, 3268, 3269; UDP 88, 389, 636 | Kerberos, LDAP, LDAPS, SMB, global catalog |
| All components | Infoblox 10.11.16.145, 10.11.18.145 | TCP and UDP 53 | DNS |
| All components | Infoblox 10.11.16.145, 10.11.18.145 | 123 | NTP |
| All components | Syslog 10.11.17.113 | TCP and UDP 514 | Syslog |
| FC switches | 10.11.17.17, 10.11.17.19 | 49 | TACACS+ |
| Nodes, cluster, management servers | Mail 10.11.19.33 | TCP 25, 465 | SMTP |
| Nodes, cluster, management servers | Proxy dc1-a-vcprx001 | TCP 3128 | AutoSupport over the web proxy |
Sensu 10.11.16.129 | Nodes and cluster | UDP 161 | SNMPv3 polling |
Sensu 10.11.16.129 | API Services 10.11.10.43 | TCP 8443 | REST API |
"All components" is long: service processors, node and cluster management, both management servers, the data interfaces of the SVMs, the FC switches and both ports of every bridge. A bridge has no use for Active Directory, but the design asked for the same rule set for everything.
The networks around the storage
flowchart LR subgraph st["Networks of the storage"] oob["Out-of-band management"] icl["Intercluster"] inb["In-band management, with Unified Manager and API Services"] bal["Balabit backup and archive, VCVSM001"] kvm["KVM storage, VCVSM002"] bck["Backup storage, VCVSM003"] adt["AD integration, VCVSM004"] end ca["Cluster DC1-A-XNAS001"] cb["Cluster DC1-B-XNAS001"] ca --- oob ca --- icl ca --- inb ca --- bal ca --- kvm ca --- bck ca --- adt cb --- oob cb --- icl cb --- inb cb --- bal cb --- kvm cb --- bck cb --- adt fw(("Router and firewall")) inb --- fw bck --- fw adt --- fw fw --- ad["Active Directory, two servers"] fw --- dns["DNS and NTP, Infoblox in A and B"] fw --- mail["Mail server"] fw --- sys["Central syslog"] fw --- mon["Monitoring, Sensu"]
The clients sit directly in their storage networks: the Balabit SCB cluster in VLAN 1020, the hypervisors of both datacenters in VLAN 1026. Nothing is routed between a client and its NFS export.
Slips in the design document
The design was written by copying the tables of datacenter A and editing them. Not every edit was made. The as-built report decides each case.
| In the design | As built |
|---|---|
DNS table gives dc1-a-anas001 and dc1-a-anas002 twice, the second time with .39 and .40 | .39 and .40 are dc1-b-anas001 and dc1-b-anas002; the access tables of the same document have them right |
Flow tables use 10.11.15.43 and 10.11.23.43 for the service processor of node 2 | .44; .43 was reserved for the cluster that was never built |
One SNMP flow has the source 10.11.10.3 | 10.11.10.33, the cluster management address |
SNMP flows to the bridges go to .47 and .48, labelled as bridge 1 and bridge 2 | .48 is the second port of bridge 1; ONTAP monitors .47 and .49 |
Flows of datacenter B give its nodes .38 and .39 | .39 and .40 |
Mail server is 10.11.19.34 in one table and 10.11.19.33 in the others; the proxy has the mail server's address | Not decided by the Source material |
| NTP as TCP 123, TACACS+ as UDP 49 | NTP is UDP; the switches list their TACACS+ servers on port 49, and TACACS+ runs over TCP |
| SNMP flows only inside a datacenter | Each cluster monitors the switches and bridges of both datacenters |
First drawing labels both interface-group members e0g | e0g and e0h |
The Source material does not contain the firewall rules that were actually installed, so it cannot be said which of these slips reached them.
Site 2
Site 2 (DC2) repeats the design with its own addresses. What the Source material holds reliably:
| Item | Site 2 |
|---|---|
| Cluster management | 10.12.10.33 for DC2-A-XNAS001, 10.12.10.34 for DC2-B-XNAS001 |
| Node management | 10.12.10.37 to 10.12.10.40 |
| Service processors | 10.12.15.42, .44 in datacenter A; 10.12.23.42, .44 in B |
DC2-S-VCVSM001 | 10.12.2.236 |
DC2-S-VCVSM005 | 10.12.40.30/27 on a0a-2020, IPspace DC2-S-VCVSM005_ipspace, broadcast domain vlan-2020 |
DC2-S-VCVSM006 | 10.12.89.94 |
The rest of the site 2 address table in the design cannot be used. Its rows for DC2-S-VCVSM002 to 004, for the FC switches and for the bridges still carry the site 1 addresses, and the two nodes of datacenter B are again named as nodes of A. The intercluster addresses and the VLAN IDs of site 2, apart from 2020, are not in the Source material at all.
Site 2 is also where the rule about identical ports on both clusters was learned in practice. In June 2019 metrocluster check reported a warning for the interface of DC2-S-VCVSM005: cluster B could not find a port with connectivity to that address. The notes for that incident start with a checklist that is the summary of this article: the same MTU and VLANs on every switch port, and the same interfaces, ports, broadcast domains and IPspaces on both clusters. The incident is in Troubleshooting.