LINUXOR.SK ... open source notes ...

NetApp - ONTAP MetroCluster switchover test

category: solutionz · date: 2019-12-31 · updated: 2026-10-02 · author: LALA

NetApp Solution · Config document · referenced from MetroCluster switchover and Tiebreaker

noteThe first command stops every data SVM of the other cluster and shuts its nodes down. It is a planned test, not something to try on a cluster that is serving clients unannounced.

The manual high-availability test of a MetroCluster as it was run: negotiated switchover from datacenter A to datacenter B, healing, booting the nodes of A, switchback, and the checks between the steps.

ItemValue
Shown here forSite 1: DC1-A-XNAS001 switched over to DC1-B-XNAS001
Runs onCluster B, the surviving cluster, except the boot commands, which are typed on the service processors of the nodes of cluster A
Also applied toSite 2, with DC2 in every name; there the same recovery steps followed a switchover triggered by the Tiebreaker
Applied withSSH to the cluster management address of B; SSH to the service processors of A
ONTAP version at the timeNot recorded in the test notes; the clusters ran 9.1 when site 1 was built

The command set

bash
# --- On cluster B: DC1-B-XNAS001 --------------------------------------------

# Negotiated switchover. Stops the data SVMs on DC1-A-XNAS001, restarts them
# on DC1-B-XNAS001 and gracefully shuts down cluster A. Asks for confirmation.
# Expected: "Job succeeded: Switchover is successful."
metrocluster switchover

# Check: local mode "switchover", remote cluster "not-reachable"
metrocluster show

# Heal the data aggregates, then the root aggregates, in this order
metrocluster heal aggregates
metrocluster heal root-aggregates

# After an unplanned switchover the second command can end with
# "completed with warnings": aggregates are still resynchronising.
# Show the warning, then wait until the resync query returns no entries.
metrocluster operation show
storage aggregate show-resync-status -in-progress true

# --- On the service processor of DC1-A-ANAS001 ------------------------------

# Open the node console from the "SP>" prompt, boot from "LOADER-A>"
system console
boot_ontap

# --- On the service processor of DC1-A-ANAS002 ------------------------------

system console
boot_ontap

# If the node stops at "waiting for giveback..." because its HA partner is in
# takeover mode: answer y to "Do you wish to halt this node rather than wait",
# then boot it again from the loader.
boot_ontap

# --- On cluster B: DC1-B-XNAS001 --------------------------------------------

# Check: local mode still "switchover", remote cluster "configured" with mode
# "waiting-for-switchback"
metrocluster show

# Switchback. Stops the switched-over SVMs on B and restarts them on A.
# Asks for confirmation. Expected: "Job succeeded: Switchback is successful."
metrocluster switchback

# Check: mode "normal" on both clusters
metrocluster show

# Summary of nodes, HA partners and DR partners after the test
metrocluster node show -fields dr-group-id,cluster,node,ha-partner,dr-cluster,dr-partner,dr-auxiliary,node-systemid,ha-partner-systemid,dr-auxiliary-systemid

The first manual test at site 1 used only the commands without the resynchronisation check and without the second boot_ontap; both were added when the recovery was repeated after the Tiebreaker tests, where the healing of the root aggregates ended with a warning and one node waited for a giveback. They are shown in the place where they were run.

RunSwitchover started byHealing of root aggregatesBooting cluster A
Site 1, manual testmetrocluster switchover on BSuccessfulBoth nodes at the first attempt
Site 1, Tiebreaker testThe Tiebreaker in site 2Completed with warnings, resync of DC1_A_ANAS001_data1 awaitedSecond node halted and booted again
Site 2, Tiebreaker testThe Tiebreaker in site 1Completed with warnings, resync of DC2_A_ANAS001_data1 awaitedBoth nodes at the first attempt

Checked against ONTAP 9.19.1

As builtToday
metrocluster switchover, negotiatedUnchanged. The options are -forced-on-disaster, -controller-replacement, -override-vetoes and, at advanced privilege, -simulate
metrocluster heal aggregates, metrocluster heal root-aggregatesThe documented form is metrocluster heal -phase aggregates, then metrocluster heal -phase root-aggregates. Both phases are still required on MetroCluster FC; automatic healing exists only for MetroCluster IP
metrocluster switchbackUnchanged, with -override-vetoes and, at advanced privilege, -simulate
storage aggregate show-resync-status, metrocluster check runBoth still have pages in the command reference
FAS8200 on ONTAP 9.1 to 9.3FAS8200 support was discontinued from ONTAP 9.17.1, so 9.16.1 is the last release for this hardware. Fabric-attached MetroCluster with four nodes is still a documented configuration

The procedure stands as it was run, with the phase spelled as an option. See metrocluster heal and metrocluster switchover.

← solutionz