LINUXOR.SK ... open source notes ...

NetApp MetroCluster - primary storage for a management infrastructure

category: solutionz · date: 2019-01-01 · updated: 2026-10-02 · author: LALA

Between 2017 and 2019 I designed, built and ran the primary storage of the management infrastructure of a network service provider, twice: once in each of two sites. It held the disks of the virtual machines that ran the infrastructure, the audit trails of the session-recording appliances, and it had to keep serving when a whole datacenter went dark.

This Solution is the whole of that system: in each site a fabric-attached MetroCluster of two FAS8200 HA pairs, with Brocade FC switches, ATTO FC-to-SAS bridges and DS224C shelves; NFS for KVM clusters; volume encryption; administrator logins through Active Directory; syslog, SNMP and AutoSupport; Unified Manager; a Tiebreaker for automatic switchover; one recorded upgrade; and the incidents on the way.

noteBuilt on ONTAP 9.1 and run on ONTAP 9.3. Every command set and device configuration is shown as it ran and then checked against ONTAP 9.19.1, the current release on 2026-10-02. The FAS8200 itself ends at ONTAP 9.16.1. The system is anonymized: sites, names, domains, addresses, people and serial numbers are replaced, and every secret is a placeholder.

The system in one picture

mermaid
flowchart TB
  clients["NFS clients: KVM hosts, Balabit SCB"]
  subgraph s1["Site 1, DC1"]
    subgraph a["Datacenter A"]
      ca["Cluster DC1-A-XNAS001, FAS8200 HA pair"]
      fa["2 Brocade 6505, 2 ATTO 7500N, 6 DS224C"]
      ca --- fa
    end
    subgraph b["Datacenter B"]
      cb["Cluster DC1-B-XNAS001, FAS8200 HA pair"]
      fb["2 Brocade 6505, 2 ATTO 7500N, 6 DS224C"]
      cb --- fb
    end
    fa == "two FC fabrics, SyncMirror and NVRAM mirror" === fb
    ca -. "cluster peering, SVM configuration" .- cb
    ocum["Unified Manager"]
  end
  subgraph s2["Site 2, DC2"]
    tb["Tiebreaker for site 1"]
    mc2["The same MetroCluster again"]
  end
  clients -- "NFS" --> ca
  ocum --> ca
  ocum --> cb
  tb -. "watches" .-> ca
  tb -. "watches" .-> cb

The fictional environment

Every article and every Config document uses the same names and addresses.

ThingSite 1Site 2
ClustersDC1-A-XNAS001, DC1-B-XNAS001DC2-A-XNAS001, DC2-B-XNAS001
NodesDC1-A-ANAS001, 002, DC1-B-ANAS001, 002DC2-A-ANAS001, 002, DC2-B-ANAS001, 002
Data SVMsDC1-S-VCVSM001 to 006DC2-S-VCVSM001 to 006
Cluster management10.11.10.33, 10.11.10.3410.12.10.33, 10.12.10.34
In-band namesadm.example.netadm.example.net
Out-of-band namesmgmt.example.netmgmt.example.net
Addresses10.11.0.0/1610.12.0.0/16

Site 1 is the worked example throughout. The networks of site 1 that come up most often:

NetworkVLANPurpose
10.11.15.0/24, 10.11.23.0/2412Out-of-band management, datacenter A and B
10.11.10.0/271022Intercluster
10.11.10.32/281024In-band management of nodes and clusters
10.11.10.64/271026NFS for the KVM hosts of the management infrastructure
10.11.18.224/281020NFS for the Balabit SCB appliances
10.11.10.48/29, 10.11.10.56/291028, 1027The two SVMs that tunnel Active Directory logins

The directory is ad.example.net, NetBIOS name EXAMPLE. Administrators are admin01 to admin06.

Articles

Read in this order; it goes from why, through how it was built, to how it was run.

#ArticleWhat it covers
1Requirements and conceptWhat the storage had to do, the four constraints the hardware imposed, and who talks to whom
2Logical designNaming, nodes, clusters, the MetroCluster, the six SVMs, and what the storage depends on
3Physical design and cablingControllers, fabrics, bridges, shelves, every port, and a reference file that contradicts the cabling
4Network designVLANs, the interface group, LIFs, names, and the firewall flows
5Storage designPools, disk ownership, mirrored aggregates, volumes, and a lopsided layout
6SVMs and NFSNFS for RHV and for the SCB appliances, qtrees, exports, and uid 36
7Access and directory integrationRoles, the LDAP attempt that failed, the domain tunnel, TACACS+ on the switches
8Encryption and certificatesVolume encryption with the Onboard Key Manager, CA-signed certificates, TLS hardening
9Logging, monitoring and AutoSupportEMS to syslog, SNMPv3 for Sensu, the cluster watching its own switches, AutoSupport
10Unified Manager and API ServicesUnified Manager on RHEL 7, and API Services in a hand-built nspawn container
11MetroCluster switchover and TiebreakerThe manual switchover test, the Tiebreaker, and the automatic-failover tests
12ONTAP upgradeA manual non-disruptive upgrade of a four-node MetroCluster
13TroubleshootingSeven incidents: symptom, cause, checks, fix

Configuration

ONTAP, Fabric OS and the Tiebreaker are configured by commands, not by files. A Config document for them holds the command set as it was run. Where the Source material holds only the result and not the commands, the document is a listing of the state as built, and says so.

Fabric and bridges

ONTAP, as built

ONTAP, access and security

ONTAP, logging and monitoring

ONTAP, operations

Management software on RHEL 7

Hardware and software

ComponentAs builtRole
NetApp FAS82004 per site, ONTAP 9.1 then 9.3P12Storage controllers, two HA pairs
NetApp DS224C12 per site, 272 SAS disks of 10,000 rpmDisk shelves
Brocade 65054 per site, Fabric OS 8.0.1The two FC fabrics between the datacenters
ATTO FibreBridge 7500N4 per site, firmware 2.85FC-to-SAS bridges in front of the shelves
OnCommand Unified Manager9.4 on RHEL 7.5Health, events, the operators' view
OnCommand API Services2.0 on RHEL 7.4, in a systemd-nspawn containerREST interface for monitoring, later dropped
MetroCluster Tiebreaker1.21P2 on RHEL 7.5Automatic switchover, decided from the other site
NetAppDocsPowerShell moduleThe as-built reports

What I would do differently

The articles say this where it belongs. The short list:

← solutionz