NetApp MetroCluster - primary storage for a management infrastructure
category: solutionz · date: 2019-01-01 · updated: 2026-10-02 · author: LALA
Between 2017 and 2019 I designed, built and ran the primary storage of the management infrastructure of a network service provider, twice: once in each of two sites. It held the disks of the virtual machines that ran the infrastructure, the audit trails of the session-recording appliances, and it had to keep serving when a whole datacenter went dark.
This Solution is the whole of that system: in each site a fabric-attached MetroCluster of two FAS8200 HA pairs, with Brocade FC switches, ATTO FC-to-SAS bridges and DS224C shelves; NFS for KVM clusters; volume encryption; administrator logins through Active Directory; syslog, SNMP and AutoSupport; Unified Manager; a Tiebreaker for automatic switchover; one recorded upgrade; and the incidents on the way.
noteBuilt on ONTAP 9.1 and run on ONTAP 9.3. Every command set and device configuration is shown as it ran and then checked against ONTAP 9.19.1, the current release on 2026-10-02. The FAS8200 itself ends at ONTAP 9.16.1. The system is anonymized: sites, names, domains, addresses, people and serial numbers are replaced, and every secret is a placeholder.
The system in one picture
mermaid
flowchart TB
clients["NFS clients: KVM hosts, Balabit SCB"]
subgraph s1["Site 1, DC1"]
subgraph a["Datacenter A"]
ca["Cluster DC1-A-XNAS001, FAS8200 HA pair"]
fa["2 Brocade 6505, 2 ATTO 7500N, 6 DS224C"]
ca --- fa
end
subgraph b["Datacenter B"]
cb["Cluster DC1-B-XNAS001, FAS8200 HA pair"]
fb["2 Brocade 6505, 2 ATTO 7500N, 6 DS224C"]
cb --- fb
end
fa == "two FC fabrics, SyncMirror and NVRAM mirror" === fb
ca -. "cluster peering, SVM configuration" .- cb
ocum["Unified Manager"]
end
subgraph s2["Site 2, DC2"]
tb["Tiebreaker for site 1"]
mc2["The same MetroCluster again"]
end
clients -- "NFS" --> ca
ocum --> ca
ocum --> cb
tb -. "watches" .-> ca
tb -. "watches" .-> cb
The fictional environment
Every article and every Config document uses the same names and addresses.
| Thing | Site 1 | Site 2 |
|---|
| Clusters | DC1-A-XNAS001, DC1-B-XNAS001 | DC2-A-XNAS001, DC2-B-XNAS001 |
| Nodes | DC1-A-ANAS001, 002, DC1-B-ANAS001, 002 | DC2-A-ANAS001, 002, DC2-B-ANAS001, 002 |
| Data SVMs | DC1-S-VCVSM001 to 006 | DC2-S-VCVSM001 to 006 |
| Cluster management | 10.11.10.33, 10.11.10.34 | 10.12.10.33, 10.12.10.34 |
| In-band names | adm.example.net | adm.example.net |
| Out-of-band names | mgmt.example.net | mgmt.example.net |
| Addresses | 10.11.0.0/16 | 10.12.0.0/16 |
Site 1 is the worked example throughout. The networks of site 1 that come up most often:
| Network | VLAN | Purpose |
|---|
10.11.15.0/24, 10.11.23.0/24 | 12 | Out-of-band management, datacenter A and B |
10.11.10.0/27 | 1022 | Intercluster |
10.11.10.32/28 | 1024 | In-band management of nodes and clusters |
10.11.10.64/27 | 1026 | NFS for the KVM hosts of the management infrastructure |
10.11.18.224/28 | 1020 | NFS for the Balabit SCB appliances |
10.11.10.48/29, 10.11.10.56/29 | 1028, 1027 | The two SVMs that tunnel Active Directory logins |
The directory is ad.example.net, NetBIOS name EXAMPLE. Administrators are admin01 to admin06.
Articles
Read in this order; it goes from why, through how it was built, to how it was run.
| # | Article | What it covers |
|---|
| 1 | Requirements and concept | What the storage had to do, the four constraints the hardware imposed, and who talks to whom |
| 2 | Logical design | Naming, nodes, clusters, the MetroCluster, the six SVMs, and what the storage depends on |
| 3 | Physical design and cabling | Controllers, fabrics, bridges, shelves, every port, and a reference file that contradicts the cabling |
| 4 | Network design | VLANs, the interface group, LIFs, names, and the firewall flows |
| 5 | Storage design | Pools, disk ownership, mirrored aggregates, volumes, and a lopsided layout |
| 6 | SVMs and NFS | NFS for RHV and for the SCB appliances, qtrees, exports, and uid 36 |
| 7 | Access and directory integration | Roles, the LDAP attempt that failed, the domain tunnel, TACACS+ on the switches |
| 8 | Encryption and certificates | Volume encryption with the Onboard Key Manager, CA-signed certificates, TLS hardening |
| 9 | Logging, monitoring and AutoSupport | EMS to syslog, SNMPv3 for Sensu, the cluster watching its own switches, AutoSupport |
| 10 | Unified Manager and API Services | Unified Manager on RHEL 7, and API Services in a hand-built nspawn container |
| 11 | MetroCluster switchover and Tiebreaker | The manual switchover test, the Tiebreaker, and the automatic-failover tests |
| 12 | ONTAP upgrade | A manual non-disruptive upgrade of a four-node MetroCluster |
| 13 | Troubleshooting | Seven incidents: symptom, cause, checks, fix |
Configuration
ONTAP, Fabric OS and the Tiebreaker are configured by commands, not by files. A Config document for them holds the command set as it was run. Where the Source material holds only the result and not the commands, the document is a listing of the state as built, and says so.
Fabric and bridges
ONTAP, as built
ONTAP, access and security
ONTAP, logging and monitoring
ONTAP, operations
Management software on RHEL 7
Hardware and software
| Component | As built | Role |
|---|
| NetApp FAS8200 | 4 per site, ONTAP 9.1 then 9.3P12 | Storage controllers, two HA pairs |
| NetApp DS224C | 12 per site, 272 SAS disks of 10,000 rpm | Disk shelves |
| Brocade 6505 | 4 per site, Fabric OS 8.0.1 | The two FC fabrics between the datacenters |
| ATTO FibreBridge 7500N | 4 per site, firmware 2.85 | FC-to-SAS bridges in front of the shelves |
| OnCommand Unified Manager | 9.4 on RHEL 7.5 | Health, events, the operators' view |
| OnCommand API Services | 2.0 on RHEL 7.4, in a systemd-nspawn container | REST interface for monitoring, later dropped |
| MetroCluster Tiebreaker | 1.21P2 on RHEL 7.5 | Automatic switchover, decided from the other site |
| NetAppDocs | PowerShell module | The as-built reports |
What I would do differently
The articles say this where it belongs. The short list:
- Keep what the device says, not only what was sent to it. The archived switch configuration files and the cabling tables contradict each other, and no
switchshow was kept to settle it. See article 3. - The key manager passphrase and backup need a named home before the first volume is encrypted, and a restore from them was never exercised. See article 8.
- One-year certificates on four clusters ran out once. The first ones expired in November 2018 and their successors are valid from January 2019; the notes do not explain the gap. See article 8.
- The first Active Directory integration was the wrong mechanism. Weeks went into the LDAP name service before the domain tunnel, and the diagnostic commands that would have said so came last. See article 7.
- The weakest device sits in front of every disk. The bridges are managed over HTTP with one local account, and only their network placement protects them. See article 1.
- Hardware-assisted takeover was inactive on all four nodes until somebody read the event log. See article 13.
- Test the unplanned switchover, and write down how the failure was injected. The clean test hid what the dirty one showed. See article 11.
- Half the notes stop before the answer. The installation log, the disk rearrangement and the login investigation all end without their conclusion, and these articles say so instead of inventing one.