NetApp 11 - MetroCluster switchover and Tiebreaker
NetApp Solution · Previous: Unified Manager and API Services · Next: ONTAP upgrade
A MetroCluster is bought for one event: a whole datacenter going dark while the storage keeps serving. This article covers how that protection works, the manual switchover tests that proved it, and the Tiebreaker software that was added so that nobody has to be awake when it happens.
What protects what
Each site has two clusters, one per datacenter, and each cluster is one FAS8200 HA pair. The two clusters together are the MetroCluster.
flowchart LR subgraph mc["Metro-cluster, site 1"] subgraph fa["Datacenter A, cluster DC1-A-XNAS001"] a1["DC1-A-ANAS001"] --- a2["DC1-A-ANAS002"] end subgraph fb["Datacenter B, cluster DC1-B-XNAS001"] b1["DC1-B-ANAS001"] --- b2["DC1-B-ANAS002"] end end a1 -. "DR partner" .- b2 a2 -. "DR partner" .- b1
The design document carries this picture twice, once generic and once per site; site 2 is the same with DC2 in every name. The solid lines are the HA pairs, which handle the loss of one controller locally by takeover and giveback. The dotted lines are the disaster-recovery partnerships across the datacenters, as metrocluster node show printed them: node 001 of one cluster is paired with node 002 of the other.
Three mechanisms make the second cluster able to take over.
- SyncMirror. Every aggregate has two plexes, one on the shelves of its own datacenter and one on the shelves of the other. Writes go to both synchronously over the Fibre Channel fabrics. Storage design has the aggregates.
- Mirrored SVM configuration. Each data SVM exists twice: running on its home cluster, and as a stopped copy with the suffix
-mcon the partner cluster. Volumes, LIFs, export policies and the rest are replicated to the copy. - Switchover. The surviving cluster brings the remote plexes online as its own and starts the
-mcSVMs, with the same addresses.
| SVM | Runs on | Stopped copy on |
|---|---|---|
DC1-S-VCVSM001, 002, 003, 005, 006 | DC1-A-XNAS001 | DC1-B-XNAS001, as DC1-S-VCVSM00x-mc |
DC1-S-VCVSM004 | DC1-B-XNAS001 | DC1-A-XNAS001, as DC1-S-VCVSM004-mc |
Almost everything runs in datacenter A. VCVSM004 is in B only because cluster B needs an SVM of its own to tunnel Active Directory logins, see Access and directory integration. In practice, then, "switchover" nearly always means B taking over A.
The manual switchover test
Before anything automatic was allowed, the switchover was run by hand at site 1, from cluster B. The notes are a commented transcript. The complete sequence is a Config document: ONTAP: MetroCluster switchover test.
sequenceDiagram participant B as Cluster B participant A as Cluster A participant SP as Service processors of A B->>A: metrocluster switchover Note over A: SVMs stopped, nodes shut down Note over B: mode switchover, serves the SVMs of A B->>B: metrocluster heal aggregates B->>B: metrocluster heal root-aggregates SP->>A: boot_ontap on both nodes Note over A: mode waiting-for-switchback B->>A: metrocluster switchback Note over A,B: mode normal on both
The first command is a negotiated switchover, the planned kind. On DC1-B-XNAS001:
$ metrocluster switchoveroutput 3 lines
Warning: negotiated switchover is about to start. It will stop all the data Vservers on cluster "DC1-A-XNAS001" and automatically restart them on cluster "DC1-B-XNAS001". It will finally gracefully shutdown cluster "DC1-A-XNAS001".
Do you want to continue? {y|n}: y
[Job 525] Job succeeded: Switchover is successful.The last sentence of the warning matters: a negotiated switchover halts the nodes of the other cluster. Afterwards metrocluster show on B printed the local cluster in mode switchover and the remote one as not-reachable.
output 8 lines
Local: DC1-B-XNAS001
Configuration State configured
Mode switchover
AUSO Failure Domain auso-on-cluster-disaster
Remote: DC1-A-XNAS001
Configuration State not-reachable
Mode -
AUSO Failure Domain not-reachableThen comes healing, in two phases and in this order. On MetroCluster FC both are still manual today; current documentation writes them as metrocluster heal -phase aggregates and -phase root-aggregates. The first resynchronises the mirrored data aggregates, the second hands the root aggregates back to the nodes that are about to boot.
$ metrocluster heal aggregates $ metrocluster heal root-aggregates
The nodes of cluster A sit at the boot loader after the shutdown. They are started from their service processors, one after the other.
$ system console $ boot_ontap
The first line is typed at the service processor prompt, the second at LOADER-A>. Once both nodes were up, metrocluster show on B printed what each stage should look like.
| Stage | Cluster B, local | Cluster A, remote |
|---|---|---|
| After switchover | configured, mode switchover | not-reachable |
| After healing and booting A | configured, mode switchover | configured, mode waiting-for-switchback |
| After switchback | configured, mode normal | configured, mode normal |
The switchback is one command, again on B. It stops the switched-over SVMs on B and restarts them on A.
$ metrocluster switchbackoutput 1 line
[Job 528] Job succeeded: Switchback is successful.
The notes do not record how long the clients, the KVM hosts with their NFS storage domains, paused during the switchover and the switchback. The design document states an objective of 120 seconds for an automatic switchover; no measurement of it is in the Source material.
Why a third location
In the manual test a person decided. Without a person, the two clusters have a problem that cannot be solved from inside: when cluster B stops hearing from cluster A, either A is dead or the links between them are. In the first case B must take over. In the second case A is still serving data, and a takeover by B would give two clusters writing to two halves of the same mirrors.
ONTAP does switch over on its own in one narrow case, shown above as AUSO Failure Domain auso-on-cluster-disaster. For everything else NetApp provides the MetroCluster Tiebreaker: a Java service on a Linux host in a third location that watches both clusters over the management network and orders a switchover only when one cluster is unreachable and the other one confirms that it has lost its partner too.
The Solution has no third location within a site. It has two sites, though, and each one is a third location for the other.
| Tiebreaker host | Runs in | Watches | Monitor name |
|---|---|---|---|
dc2-a-vcntb001.adm.example.net | Site 2 | DC1-A-XNAS001 at 10.11.10.33 and DC1-B-XNAS001 at 10.11.10.34 | Metro-Cluster_in_DC1 |
dc1-a-vcntb001.adm.example.net | Site 1 | DC2-A-XNAS001 at 10.12.10.33 and DC2-B-XNAS001 at 10.12.10.34 | Metro-Cluster_in_DC2 |
A fire that takes out datacenter A of site 1 cannot take the Tiebreaker with it. The price is that the Tiebreaker's view depends on the network between the sites: if that link is down, the Tiebreaker sees neither cluster and does nothing, which is the safe outcome.
Installing the Tiebreaker
Each Tiebreaker is a small highly available virtual machine on the KVM platform of its site: 2 vCPU, 4 GB of memory, a 32 GB disk, one interface in VLAN 1021, Red Hat Enterprise Linux 7.5. The software was version 1.21P2. The full listing is a Config document: Tiebreaker installation on RHEL 7.
It has two prerequisites, a Java runtime and a database.
$ yum install java-1.8.0-openjdk $ yum install mariadb-server $ systemctl start mariadb $ mysql_secure_installation
The package itself asks for two passwords during installation, one for its own database user and the MariaDB root password.
$ rpm -ivh NetApp-MetroCluster-Tiebreaker-Software-1.21P2-1.x86_64.rpm $ systemctl start netapp-metrocluster-tiebreaker-software $ systemctl enable mariadb $ systemctl enable netapp-metrocluster-tiebreaker-software
/tmp is mounted with noexec. The mount point must be without that flag on a Tiebreaker host.This is the one warning the installation notes write in capitals. A hardened Linux baseline commonly mounts /tmp with noexec, and the package installs without complaint on such a host; only the service fails to start. The notes do not record how the mount option was changed on these two machines.
Monitor, observer mode, online mode
The Tiebreaker has its own shell, started as root with netapp-metrocluster-tiebreaker-software-cli. A MetroCluster is added as a monitor, with the cluster management address and the local admin account of each cluster.
$ monitor add wizard $ monitor show -status
output 13 lines
MetroCluster: Metro-Cluster_in_DC1
Disaster: false
Monitor State: Normal
Observer Mode: true
Silent Period: 5
Override Vetoes: false
Cluster: DC1-A-XNAS001(UUID:<CLUSTER_UUID>)
Reachable: true
Intersite Connectivity Available: true
Node: DC1-A-ANAS001(UUID:<NODE_UUID>)
Reachable: true
Intersite Connectivity Available: true
State: normalA new monitor starts in observer mode: it watches and reports, and never acts. Turning it off, which the notes call switching to online mode, is the step that gives the software authority over the storage, and the CLI says so. The transcript below is from the Tiebreaker in site 2.
$ monitor modify -monitor-name Metro-Cluster_in_DC1 -observer-mode falseoutput 2 lines
Warning: If you are turning observer-mode to false, make sure to review the 'risks and limitations' as described in the MetroCluster Tiebreaker Installation and Configuration Guide. Are you sure you want to enable automatic switchover capability for monitor "Metro-Cluster_in_DC1"? [Y/N]: Y
The Tiebreaker also sends its own AutoSupport messages. It was pointed at the HTTP proxy of its own site, port 3128, without a proxy user name, and autosupport invoke answered AutoSupport transmission : success. The clusters' AutoSupport is in Logging, monitoring and AutoSupport.
The automatic-failover tests
Two diagrams in the design material show where the failures were injected. The first one cuts only Ethernet: the LACP group a0a of a node, ports e0g and e0h, which carries every VLAN the node has except the cluster interconnect and the service processor.
flowchart TB cl["Clients, Balabit SCB and KVM hosts"] --> v["VLANs 1020, 1026, 1028 data"] tb["Tiebreaker in the other site"] --> m["VLAN 1024 management"] ic["Partner cluster"] --> i["VLAN 1022 intercluster"] v --> a0a m --> a0a i --> a0a subgraph n1["Node DC1-A-ANAS001"] a0a["a0a LACP, e0g and e0h, cut here"] cli["e0a, e0b cluster interconnect"] sp["e0M service processor"] end
The second one cuts both paths out of datacenter A: the two inter-switch links of each Fibre Channel fabric, and the Ethernet ports of both nodes.
flowchart LR subgraph fa["Datacenter A"] n1["ANAS001, e0g e0h, cut 3"] n2["ANAS002, e0g e0h, cut 4"] s1["SNAS001, ISL ports 8 and 9, cut 1"] s2["SNAS002, ISL ports 8 and 9, cut 2"] n1 --- s1 n1 --- s2 n2 --- s1 n2 --- s2 end s1 -- "fabric 1 ISL" --> fb["Datacenter B"] s2 -- "fabric 2 ISL" --> fb n1 -- "Ethernet" --> lan["Data and management network"] n2 -- "Ethernet" --> lan
Site 1
The notes of the site 1 test are the Tiebreaker's status at each stage. They do not contain the commands that caused the failure, so the table says what the Tiebreaker reported, and the diagrams above are the only record of what was cut.
| Stage | Monitor state | Cluster A | Cluster B |
|---|---|---|---|
| Before | Normal | Reachable, intersite connectivity available | Reachable, intersite connectivity available |
| Cluster A lost from the network | MCCTB is unable to reach cluster "DC1-A-XNAS001" | Not reachable | Reachable, intersite connectivity still available |
| B loses A as well | Disaster detected, Disaster: true | Not reachable, node state unknown | Reachable, intersite connectivity not available |
| Decision | Triggered switchover at cluster DC1-B-XNAS001 | Not reachable | Nodes switchover in progress, then switchover completed |
| After healing | MCC in switched over state | Not listed | Nodes heal roots completed |
| A booted | MCC in switched over state | Nodes waiting for switchback recovery | heal roots completed |
| After switchback | Normal | normal | normal |
The second row is the whole point of the product. With cluster A unreachable from the Tiebreaker but still visible to cluster B over Fibre Channel, the Tiebreaker reported and did nothing. Only when B also said that it had lost its partner did it declare a disaster and trigger the switchover on B.
The Tiebreaker switches over and stops there. Healing, booting and switchback were done by hand as in the manual test, with two differences from it.
The root-aggregate healing finished with a warning, because the data aggregates were still resynchronising. On cluster B:
$ metrocluster heal root-aggregates $ storage aggregate show-resync-status -in-progress true
output 6 lines
[Job 10345] Job succeeded: Heal Root Aggregates completed with warnings. Use the "metrocluster operation show" command to view the warnings.
Aggregate Resyncing Plex Percentage
--------- ------------------------- ----------
DC1_A_ANAS001_data1
plex0 13.64The switchback was held until that command answered There are no entries matching your query. After an unplanned switchover the mirrors really have diverged, and a negotiated test never shows this.
The second node of cluster A did not come up at the first attempt. It stopped with The HA partner is currently operational and in takeover mode ... waiting for giveback, was halted from that prompt, and booted normally with a second boot_ontap.
Site 2
The site 2 test, on 14 November 2019, has its commands recorded. The inter-switch links were disabled on both Fibre Channel switches of datacenter A, as admin01 on DC2-A-SNAS001 and then on DC2-A-SNAS002.
$ portdisable 8-9
The Tiebreaker in site 1 answered with Disaster: true and Triggered switchover at cluster DC2-B-XNAS001, and a little later with MCC in switched over state and both nodes of B at switchover completed. Its output shows cluster A as not reachable at that moment, which the port commands alone do not explain; the diagram has the Ethernet ports cut as well, and the notes do not say how that was done.
The recovery was the same: portenable 8-9, heal aggregates, heal root aggregates with the same warning, a resynchronisation of DC2_A_ANAS001_data1 that the notes caught at 14.93 %, both nodes booted from their service processors without the giveback detour, switchback, and Monitor State: Normal. The notes show the portenable on the first switch only.
Lessons
- Test the unplanned case. The negotiated switchover is clean and proves the procedure. The cut-the-links test produced the two things that the clean test hid: healing that ends with a warning, and a node that waits for a giveback.
- Do not switch back on a warning.
storage aggregate show-resync-status -in-progress truemust return nothing first. - Leave observer mode on purpose. A Tiebreaker left in observer mode looks healthy in every status output and will not act.
Observer Mode: falsebelongs on the checklist. - Write down the injection, not only the reaction. The site 1 notes have every status the Tiebreaker printed and none of the commands that caused them.
- No timings were taken. The 120-second objective in the design was neither confirmed nor refuted by anything in the notes.