LINUXOR.SK ... open source notes ...

NetApp 11 - MetroCluster switchover and Tiebreaker

category: solutionz · date: 2019-12-31 · updated: 2026-10-02 · author: LALA

NetApp Solution · Previous: Unified Manager and API Services · Next: ONTAP upgrade

A MetroCluster is bought for one event: a whole datacenter going dark while the storage keeps serving. This article covers how that protection works, the manual switchover tests that proved it, and the Tiebreaker software that was added so that nobody has to be awake when it happens.

What protects what

Each site has two clusters, one per datacenter, and each cluster is one FAS8200 HA pair. The two clusters together are the MetroCluster.

mermaid
flowchart LR
  subgraph mc["Metro-cluster, site 1"]
    subgraph fa["Datacenter A, cluster DC1-A-XNAS001"]
      a1["DC1-A-ANAS001"] --- a2["DC1-A-ANAS002"]
    end
    subgraph fb["Datacenter B, cluster DC1-B-XNAS001"]
      b1["DC1-B-ANAS001"] --- b2["DC1-B-ANAS002"]
    end
  end
  a1 -. "DR partner" .- b2
  a2 -. "DR partner" .- b1

The design document carries this picture twice, once generic and once per site; site 2 is the same with DC2 in every name. The solid lines are the HA pairs, which handle the loss of one controller locally by takeover and giveback. The dotted lines are the disaster-recovery partnerships across the datacenters, as metrocluster node show printed them: node 001 of one cluster is paired with node 002 of the other.

Three mechanisms make the second cluster able to take over.

SVMRuns onStopped copy on
DC1-S-VCVSM001, 002, 003, 005, 006DC1-A-XNAS001DC1-B-XNAS001, as DC1-S-VCVSM00x-mc
DC1-S-VCVSM004DC1-B-XNAS001DC1-A-XNAS001, as DC1-S-VCVSM004-mc

Almost everything runs in datacenter A. VCVSM004 is in B only because cluster B needs an SVM of its own to tunnel Active Directory logins, see Access and directory integration. In practice, then, "switchover" nearly always means B taking over A.

The manual switchover test

Before anything automatic was allowed, the switchover was run by hand at site 1, from cluster B. The notes are a commented transcript. The complete sequence is a Config document: ONTAP: MetroCluster switchover test.

mermaid
sequenceDiagram
  participant B as Cluster B
  participant A as Cluster A
  participant SP as Service processors of A
  B->>A: metrocluster switchover
  Note over A: SVMs stopped, nodes shut down
  Note over B: mode switchover, serves the SVMs of A
  B->>B: metrocluster heal aggregates
  B->>B: metrocluster heal root-aggregates
  SP->>A: boot_ontap on both nodes
  Note over A: mode waiting-for-switchback
  B->>A: metrocluster switchback
  Note over A,B: mode normal on both

The first command is a negotiated switchover, the planned kind. On DC1-B-XNAS001:

bash
$ metrocluster switchover
output 3 lines
Warning: negotiated switchover is about to start. It will stop all the data Vservers on cluster "DC1-A-XNAS001" and automatically restart them on cluster "DC1-B-XNAS001". It will finally gracefully shutdown cluster "DC1-A-XNAS001".
Do you want to continue? {y|n}: y
[Job 525] Job succeeded: Switchover is successful.

The last sentence of the warning matters: a negotiated switchover halts the nodes of the other cluster. Afterwards metrocluster show on B printed the local cluster in mode switchover and the remote one as not-reachable.

output 8 lines
Local: DC1-B-XNAS001
         Configuration State  configured
                        Mode  switchover
         AUSO Failure Domain  auso-on-cluster-disaster
Remote: DC1-A-XNAS001
         Configuration State  not-reachable
                        Mode  -
         AUSO Failure Domain  not-reachable

Then comes healing, in two phases and in this order. On MetroCluster FC both are still manual today; current documentation writes them as metrocluster heal -phase aggregates and -phase root-aggregates. The first resynchronises the mirrored data aggregates, the second hands the root aggregates back to the nodes that are about to boot.

bash
$ metrocluster heal aggregates
$ metrocluster heal root-aggregates

The nodes of cluster A sit at the boot loader after the shutdown. They are started from their service processors, one after the other.

bash
$ system console
$ boot_ontap

The first line is typed at the service processor prompt, the second at LOADER-A>. Once both nodes were up, metrocluster show on B printed what each stage should look like.

StageCluster B, localCluster A, remote
After switchoverconfigured, mode switchovernot-reachable
After healing and booting Aconfigured, mode switchoverconfigured, mode waiting-for-switchback
After switchbackconfigured, mode normalconfigured, mode normal

The switchback is one command, again on B. It stops the switched-over SVMs on B and restarts them on A.

bash
$ metrocluster switchback
output 1 line
[Job 528] Job succeeded: Switchback is successful.

The notes do not record how long the clients, the KVM hosts with their NFS storage domains, paused during the switchover and the switchback. The design document states an objective of 120 seconds for an automatic switchover; no measurement of it is in the Source material.

Why a third location

In the manual test a person decided. Without a person, the two clusters have a problem that cannot be solved from inside: when cluster B stops hearing from cluster A, either A is dead or the links between them are. In the first case B must take over. In the second case A is still serving data, and a takeover by B would give two clusters writing to two halves of the same mirrors.

ONTAP does switch over on its own in one narrow case, shown above as AUSO Failure Domain auso-on-cluster-disaster. For everything else NetApp provides the MetroCluster Tiebreaker: a Java service on a Linux host in a third location that watches both clusters over the management network and orders a switchover only when one cluster is unreachable and the other one confirms that it has lost its partner too.

The Solution has no third location within a site. It has two sites, though, and each one is a third location for the other.

Tiebreaker hostRuns inWatchesMonitor name
dc2-a-vcntb001.adm.example.netSite 2DC1-A-XNAS001 at 10.11.10.33 and DC1-B-XNAS001 at 10.11.10.34Metro-Cluster_in_DC1
dc1-a-vcntb001.adm.example.netSite 1DC2-A-XNAS001 at 10.12.10.33 and DC2-B-XNAS001 at 10.12.10.34Metro-Cluster_in_DC2

A fire that takes out datacenter A of site 1 cannot take the Tiebreaker with it. The price is that the Tiebreaker's view depends on the network between the sites: if that link is down, the Tiebreaker sees neither cluster and does nothing, which is the safe outcome.

Installing the Tiebreaker

Each Tiebreaker is a small highly available virtual machine on the KVM platform of its site: 2 vCPU, 4 GB of memory, a 32 GB disk, one interface in VLAN 1021, Red Hat Enterprise Linux 7.5. The software was version 1.21P2. The full listing is a Config document: Tiebreaker installation on RHEL 7.

It has two prerequisites, a Java runtime and a database.

bash
$ yum install java-1.8.0-openjdk
$ yum install mariadb-server
$ systemctl start mariadb
$ mysql_secure_installation

The package itself asks for two passwords during installation, one for its own database user and the MariaDB root password.

bash
$ rpm -ivh NetApp-MetroCluster-Tiebreaker-Software-1.21P2-1.x86_64.rpm
$ systemctl start netapp-metrocluster-tiebreaker-software
$ systemctl enable mariadb
$ systemctl enable netapp-metrocluster-tiebreaker-software
noteThe daemon does not start when /tmp is mounted with noexec. The mount point must be without that flag on a Tiebreaker host.

This is the one warning the installation notes write in capitals. A hardened Linux baseline commonly mounts /tmp with noexec, and the package installs without complaint on such a host; only the service fails to start. The notes do not record how the mount option was changed on these two machines.

Monitor, observer mode, online mode

The Tiebreaker has its own shell, started as root with netapp-metrocluster-tiebreaker-software-cli. A MetroCluster is added as a monitor, with the cluster management address and the local admin account of each cluster.

bash
$ monitor add wizard
$ monitor show -status
output 13 lines
MetroCluster: Metro-Cluster_in_DC1
    Disaster: false
    Monitor State: Normal
    Observer Mode: true
    Silent Period: 5
    Override Vetoes: false
    Cluster: DC1-A-XNAS001(UUID:<CLUSTER_UUID>)
        Reachable: true
        Intersite Connectivity Available: true
            Node: DC1-A-ANAS001(UUID:<NODE_UUID>)
                Reachable: true
                Intersite Connectivity Available: true
                State: normal

A new monitor starts in observer mode: it watches and reports, and never acts. Turning it off, which the notes call switching to online mode, is the step that gives the software authority over the storage, and the CLI says so. The transcript below is from the Tiebreaker in site 2.

bash
$ monitor modify -monitor-name Metro-Cluster_in_DC1 -observer-mode false
output 2 lines
Warning: If you are turning observer-mode to false, make sure to review the 'risks and limitations' as described in the MetroCluster Tiebreaker Installation and Configuration Guide.
Are you sure you want to enable automatic switchover capability for monitor "Metro-Cluster_in_DC1"? [Y/N]: Y

The Tiebreaker also sends its own AutoSupport messages. It was pointed at the HTTP proxy of its own site, port 3128, without a proxy user name, and autosupport invoke answered AutoSupport transmission : success. The clusters' AutoSupport is in Logging, monitoring and AutoSupport.

The automatic-failover tests

Two diagrams in the design material show where the failures were injected. The first one cuts only Ethernet: the LACP group a0a of a node, ports e0g and e0h, which carries every VLAN the node has except the cluster interconnect and the service processor.

mermaid
flowchart TB
  cl["Clients, Balabit SCB and KVM hosts"] --> v["VLANs 1020, 1026, 1028 data"]
  tb["Tiebreaker in the other site"] --> m["VLAN 1024 management"]
  ic["Partner cluster"] --> i["VLAN 1022 intercluster"]
  v --> a0a
  m --> a0a
  i --> a0a
  subgraph n1["Node DC1-A-ANAS001"]
    a0a["a0a LACP, e0g and e0h, cut here"]
    cli["e0a, e0b cluster interconnect"]
    sp["e0M service processor"]
  end

The second one cuts both paths out of datacenter A: the two inter-switch links of each Fibre Channel fabric, and the Ethernet ports of both nodes.

mermaid
flowchart LR
  subgraph fa["Datacenter A"]
    n1["ANAS001, e0g e0h, cut 3"]
    n2["ANAS002, e0g e0h, cut 4"]
    s1["SNAS001, ISL ports 8 and 9, cut 1"]
    s2["SNAS002, ISL ports 8 and 9, cut 2"]
    n1 --- s1
    n1 --- s2
    n2 --- s1
    n2 --- s2
  end
  s1 -- "fabric 1 ISL" --> fb["Datacenter B"]
  s2 -- "fabric 2 ISL" --> fb
  n1 -- "Ethernet" --> lan["Data and management network"]
  n2 -- "Ethernet" --> lan

Site 1

The notes of the site 1 test are the Tiebreaker's status at each stage. They do not contain the commands that caused the failure, so the table says what the Tiebreaker reported, and the diagrams above are the only record of what was cut.

StageMonitor stateCluster ACluster B
BeforeNormalReachable, intersite connectivity availableReachable, intersite connectivity available
Cluster A lost from the networkMCCTB is unable to reach cluster "DC1-A-XNAS001"Not reachableReachable, intersite connectivity still available
B loses A as wellDisaster detected, Disaster: trueNot reachable, node state unknownReachable, intersite connectivity not available
DecisionTriggered switchover at cluster DC1-B-XNAS001Not reachableNodes switchover in progress, then switchover completed
After healingMCC in switched over stateNot listedNodes heal roots completed
A bootedMCC in switched over stateNodes waiting for switchback recoveryheal roots completed
After switchbackNormalnormalnormal

The second row is the whole point of the product. With cluster A unreachable from the Tiebreaker but still visible to cluster B over Fibre Channel, the Tiebreaker reported and did nothing. Only when B also said that it had lost its partner did it declare a disaster and trigger the switchover on B.

The Tiebreaker switches over and stops there. Healing, booting and switchback were done by hand as in the manual test, with two differences from it.

The root-aggregate healing finished with a warning, because the data aggregates were still resynchronising. On cluster B:

bash
$ metrocluster heal root-aggregates
$ storage aggregate show-resync-status -in-progress true
output 6 lines
[Job 10345] Job succeeded: Heal Root Aggregates completed with warnings. Use the "metrocluster operation show" command to view the warnings.

Aggregate Resyncing Plex            Percentage
--------- ------------------------- ----------
DC1_A_ANAS001_data1
          plex0                          13.64

The switchback was held until that command answered There are no entries matching your query. After an unplanned switchover the mirrors really have diverged, and a negotiated test never shows this.

The second node of cluster A did not come up at the first attempt. It stopped with The HA partner is currently operational and in takeover mode ... waiting for giveback, was halted from that prompt, and booted normally with a second boot_ontap.

Site 2

The site 2 test, on 14 November 2019, has its commands recorded. The inter-switch links were disabled on both Fibre Channel switches of datacenter A, as admin01 on DC2-A-SNAS001 and then on DC2-A-SNAS002.

bash
$ portdisable 8-9

The Tiebreaker in site 1 answered with Disaster: true and Triggered switchover at cluster DC2-B-XNAS001, and a little later with MCC in switched over state and both nodes of B at switchover completed. Its output shows cluster A as not reachable at that moment, which the port commands alone do not explain; the diagram has the Ethernet ports cut as well, and the notes do not say how that was done.

The recovery was the same: portenable 8-9, heal aggregates, heal root aggregates with the same warning, a resynchronisation of DC2_A_ANAS001_data1 that the notes caught at 14.93 %, both nodes booted from their service processors without the giveback detour, switchback, and Monitor State: Normal. The notes show the portenable on the first switch only.

Lessons

← solutionz