NetApp - ONTAP MetroCluster switchover test
NetApp Solution · Config document · referenced from MetroCluster switchover and Tiebreaker
The manual high-availability test of a MetroCluster as it was run: negotiated switchover from datacenter A to datacenter B, healing, booting the nodes of A, switchback, and the checks between the steps.
| Item | Value |
|---|---|
| Shown here for | Site 1: DC1-A-XNAS001 switched over to DC1-B-XNAS001 |
| Runs on | Cluster B, the surviving cluster, except the boot commands, which are typed on the service processors of the nodes of cluster A |
| Also applied to | Site 2, with DC2 in every name; there the same recovery steps followed a switchover triggered by the Tiebreaker |
| Applied with | SSH to the cluster management address of B; SSH to the service processors of A |
| ONTAP version at the time | Not recorded in the test notes; the clusters ran 9.1 when site 1 was built |
The command set
# --- On cluster B: DC1-B-XNAS001 -------------------------------------------- # Negotiated switchover. Stops the data SVMs on DC1-A-XNAS001, restarts them # on DC1-B-XNAS001 and gracefully shuts down cluster A. Asks for confirmation. # Expected: "Job succeeded: Switchover is successful." metrocluster switchover # Check: local mode "switchover", remote cluster "not-reachable" metrocluster show # Heal the data aggregates, then the root aggregates, in this order metrocluster heal aggregates metrocluster heal root-aggregates # After an unplanned switchover the second command can end with # "completed with warnings": aggregates are still resynchronising. # Show the warning, then wait until the resync query returns no entries. metrocluster operation show storage aggregate show-resync-status -in-progress true # --- On the service processor of DC1-A-ANAS001 ------------------------------ # Open the node console from the "SP>" prompt, boot from "LOADER-A>" system console boot_ontap # --- On the service processor of DC1-A-ANAS002 ------------------------------ system console boot_ontap # If the node stops at "waiting for giveback..." because its HA partner is in # takeover mode: answer y to "Do you wish to halt this node rather than wait", # then boot it again from the loader. boot_ontap # --- On cluster B: DC1-B-XNAS001 -------------------------------------------- # Check: local mode still "switchover", remote cluster "configured" with mode # "waiting-for-switchback" metrocluster show # Switchback. Stops the switched-over SVMs on B and restarts them on A. # Asks for confirmation. Expected: "Job succeeded: Switchback is successful." metrocluster switchback # Check: mode "normal" on both clusters metrocluster show # Summary of nodes, HA partners and DR partners after the test metrocluster node show -fields dr-group-id,cluster,node,ha-partner,dr-cluster,dr-partner,dr-auxiliary,node-systemid,ha-partner-systemid,dr-auxiliary-systemid
The first manual test at site 1 used only the commands without the resynchronisation check and without the second boot_ontap; both were added when the recovery was repeated after the Tiebreaker tests, where the healing of the root aggregates ended with a warning and one node waited for a giveback. They are shown in the place where they were run.
| Run | Switchover started by | Healing of root aggregates | Booting cluster A |
|---|---|---|---|
| Site 1, manual test | metrocluster switchover on B | Successful | Both nodes at the first attempt |
| Site 1, Tiebreaker test | The Tiebreaker in site 2 | Completed with warnings, resync of DC1_A_ANAS001_data1 awaited | Second node halted and booted again |
| Site 2, Tiebreaker test | The Tiebreaker in site 1 | Completed with warnings, resync of DC2_A_ANAS001_data1 awaited | Both nodes at the first attempt |
Checked against ONTAP 9.19.1
| As built | Today |
|---|---|
metrocluster switchover, negotiated | Unchanged. The options are -forced-on-disaster, -controller-replacement, -override-vetoes and, at advanced privilege, -simulate |
metrocluster heal aggregates, metrocluster heal root-aggregates | The documented form is metrocluster heal -phase aggregates, then metrocluster heal -phase root-aggregates. Both phases are still required on MetroCluster FC; automatic healing exists only for MetroCluster IP |
metrocluster switchback | Unchanged, with -override-vetoes and, at advanced privilege, -simulate |
storage aggregate show-resync-status, metrocluster check run | Both still have pages in the command reference |
| FAS8200 on ONTAP 9.1 to 9.3 | FAS8200 support was discontinued from ONTAP 9.17.1, so 9.16.1 is the last release for this hardware. Fabric-attached MetroCluster with four nodes is still a documented configuration |
The procedure stands as it was run, with the phase spelled as an option. See metrocluster heal and metrocluster switchover.