LINUXOR.SK ... open source notes ...

NetApp - ONTAP upgrade runbook for a four-node MetroCluster

category: solutionz · date: 2019-12-31 · updated: 2026-10-02 · author: LALA

NetApp Solution · Config document · referenced from ONTAP upgrade

noteThe commands switch between two clusters. Each block says which one; a takeover typed on the wrong cluster takes over the wrong node.

The non-disruptive upgrade of a four-node fabric MetroCluster by manual takeover and giveback, as it was run at site 1: checks, preparation, two rounds of takeover and giveback in the order of the DR partnerships, and the tasks afterwards.

ItemValue
Shown here forSite 1, clusters DC1-A-XNAS001 and DC1-B-XNAS001, on 20 February 2018
ONTAP versionFrom 9.1P8 to 9.1P11
Image sourcehttp://10.11.10.44:8080/, a web server on the Unified Manager host DC1-A-VCOCM001
Also applied toSite 2: a copy of this runbook with more pre-checks exists; no record of its run
Applied withSSH to the cluster management address of each cluster

The command set

bash
# ############################################################################
# CHECKS BEFORE THE UPGRADE
# ############################################################################

# --- Number of Snapshot copies per node, must be below 20,000 ---------------
# On DC1-A-XNAS001
set advanced
volume snapshot show -node DC1-A-ANAS001
volume snapshot show -node DC1-A-ANAS002
# On DC1-B-XNAS001
set advanced
volume snapshot show -node DC1-B-ANAS001
volume snapshot show -node DC1-B-ANAS002

# --- CPU utilisation, must stay below 50 % in all 10 samples ----------------
# On DC1-A-XNAS001
node run -node DC1-A-ANAS001 -command sysstat -c 10 -x 3
node run -node DC1-A-ANAS002 -command sysstat -c 10 -x 3
# On DC1-B-XNAS001
node run -node DC1-B-ANAS001 -command sysstat -c 10 -x 3
node run -node DC1-B-ANAS002 -command sysstat -c 10 -x 3

# --- No mirrored aggregate may be resynchronising ---------------------------
# On DC1-A-XNAS001; expected "There are no entries matching your query."
storage aggregate plex show -in-progress true

# --- LIFs, aggregates, volumes: on both clusters ----------------------------
# LIFs of the running SVMs up and on their home nodes
network interface show
# Anything listed here is not online
storage aggregate show -state !online
volume show -state !online

# --- No aggregate, volume or Snapshot jobs running or queued ----------------
# On DC1-A-XNAS001. Prefer to wait; delete only a job that will not finish.
job show
job delete -id <job_id>

# --- Back-end configuration errors and failed disks: on both clusters -------
# Expected "This table is currently empty."
storage errors show

# --- MetroCluster check -----------------------------------------------------
# On DC1-A-XNAS001; expected ok for nodes, lifs, config-replication,
# aggregates and clusters
metrocluster check run
metrocluster check show

# ############################################################################
# PRE-UPGRADE TASKS
# ############################################################################

# --- AutoSupport: announce a two-hour maintenance window --------------------
# On DC1-A-XNAS001, then the same two commands on DC1-B-XNAS001
system node autosupport invoke -node * -type all -message "MAINT=2h Starting_NDU"
autosupport history show -subject *MAINT*

# --- DR partnerships decide the order ---------------------------------------
# On DC1-A-XNAS001. Result: DC1-A-ANAS001 <-> DC1-B-ANAS002,
#                           DC1-A-ANAS002 <-> DC1-B-ANAS001
metrocluster node show -fields dr-partner

# --- Copy the image to all four nodes and make it the default ---------------
# On DC1-A-XNAS001
system node image update -node DC1-A-ANAS001 -package http://10.11.10.44:8080/91P11_q_image.tgz -setdefault true
system node image update -node DC1-A-ANAS002 -package http://10.11.10.44:8080/91P11_q_image.tgz -setdefault true
# On DC1-B-XNAS001
system node image update -node DC1-B-ANAS001 -package http://10.11.10.44:8080/91P11_q_image.tgz -setdefault true
system node image update -node DC1-B-ANAS002 -package http://10.11.10.44:8080/91P11_q_image.tgz -setdefault true

# --- Confirm: 9.1P11 is default and not current, on both clusters -----------
system node image show
system node image package show

# --- Disable automatic giveback on all four nodes ---------------------------
# On DC1-A-XNAS001
storage failover modify -node DC1-A-ANAS001 -auto-giveback false
storage failover modify -node DC1-A-ANAS002 -auto-giveback false
# On DC1-B-XNAS001
storage failover modify -node DC1-B-ANAS001 -auto-giveback false
storage failover modify -node DC1-B-ANAS002 -auto-giveback false
# Verify, on both clusters
storage failover show -fields auto-giveback

# ############################################################################
# ROUND 1: DC1-A-ANAS001 AND ITS DR PARTNER DC1-B-ANAS002
# ############################################################################

# --- Takeover of the first node; on DC1-A-XNAS001 ---------------------------
# Expected: DC1-A-ANAS001 "Waiting for giveback", DC1-A-ANAS002 "In takeover"
storage failover takeover -ofnode DC1-A-ANAS001
storage failover show

# --- Takeover of its DR partner; on DC1-B-XNAS001 ---------------------------
storage failover takeover -ofnode DC1-B-ANAS002
storage failover show

# --- Wait at least eight minutes --------------------------------------------
# Client multipathing must stabilise and clients must recover from the pause
# in I/O that the takeover caused.

# --- Giveback; on DC1-A-XNAS001 ---------------------------------------------
# Repeat the second command until "No aggregates to give back"
storage failover giveback -ofnode DC1-A-ANAS001
storage failover show-giveback

# --- Giveback; on DC1-B-XNAS001 ---------------------------------------------
storage failover giveback -ofnode DC1-B-ANAS002
storage failover show-giveback

# --- Both nodes of round 1 must now run the new image; on both clusters -----
system node image show

# ############################################################################
# ROUND 2: DC1-A-ANAS002 AND ITS DR PARTNER DC1-B-ANAS001
# ############################################################################

# --- Takeover by a partner that already runs the new version ----------------
# On DC1-A-XNAS001
storage failover takeover -ofnode DC1-A-ANAS002 -option allow-versionmismatch
storage failover show
# On DC1-B-XNAS001
storage failover takeover -ofnode DC1-B-ANAS001 -option allow-versionmismatch
storage failover show

# --- Wait at least eight minutes, as in round 1 -----------------------------

# --- Giveback; on DC1-A-XNAS001 ---------------------------------------------
storage failover giveback -ofnode DC1-A-ANAS002
storage failover show-giveback

# --- Giveback; on DC1-B-XNAS001 ---------------------------------------------
storage failover giveback -ofnode DC1-B-ANAS001
storage failover show-giveback

# --- All four nodes must now run the new image; on both clusters ------------
system node image show

# ############################################################################
# POST-UPGRADE TASKS
# ############################################################################

# --- Enable automatic giveback again on all four nodes ----------------------
# On DC1-A-XNAS001
storage failover modify -node DC1-A-ANAS001 -auto-giveback true
storage failover modify -node DC1-A-ANAS002 -auto-giveback true
# On DC1-B-XNAS001
storage failover modify -node DC1-B-ANAS001 -auto-giveback true
storage failover modify -node DC1-B-ANAS002 -auto-giveback true
# Verify, on both clusters
storage failover show -fields auto-giveback

# --- MetroCluster check; on DC1-A-XNAS001 -----------------------------------
metrocluster check run
metrocluster check show

# --- AutoSupport: the upgrade is finished -----------------------------------
# On DC1-A-XNAS001, then the same two commands on DC1-B-XNAS001
system node autosupport invoke -node * -type all -message "Finishing_NDU"
autosupport history show -subject *MAINT*

# ############################################################################
# CHECKS AFTER THE UPGRADE: on both clusters
# ############################################################################

storage errors show
vserver show
network interface show
network port show

The runbook kept for site 2 is this one with the following differences. It still carries the node names and the image of site 1, so the table lists what was added or removed, not what was run at site 2.

Place in the runbookSite 1, as runSite 2 runbook
Before the LIF checkNothingldap check -vserver <vserver>, and ldap client modify -client-config <config> -ldap-servers <ip_address> if the check fails
Before the LIF checkNothingstorage failover show: failover enabled and possible
After the volume checkNothingvolume show -is-inconsistent true
After the volume checkNothingdf -vserver <vserver> -volume <volume> for deduplicated volumes: 4 % free in the volume, 3 % in the aggregate
After the volume checkNothingnetwork interface show -role data -failover: failover policy and targets of each data LIF
After the volume checkNothingnetwork interface modify {-role data} -status-admin up and network interface revert *
After storage errors showNothingstorage disk show -state for disks in maintenance, pending or reconstructing
Round 2 takeoversWith -option allow-versionmismatchWithout the option

The reference the notes give for trouble during an upgrade is the vendor's article "FAQ: Troubleshooting Clustered Data ONTAP Upgrades".

Checked against ONTAP 9.19.1

As builtToday
Manual takeover and giveback, one DR pair at a timeDocumented only for four-node MetroCluster configurations on ONTAP 9.2 or earlier and for eight-node MetroCluster FC. The steps themselves are unchanged: automatic giveback off, DR pairs in lockstep, an eight-minute wait before giveback, automatic giveback on again
No automated method usedFor a two- or four-node MetroCluster on 9.3 or later the recommended method is the automated non-disruptive upgrade, with System Manager or with cluster image package get, cluster image validate and cluster image update. With four nodes the automated upgrade starts on the HA pairs of both sites at the same time; validation is run on cluster A, then on cluster B
-option allow-versionmismatch on the round 2 takeoversThe documented value is allow-version-mismatch, and it is "not required for upgrades from ONTAP 9.0 to ONTAP 9.1 or for any patch upgrades", so the 9.1P8 to 9.1P11 run did not need it. Whether the 9.1 command line accepted the spelling without the hyphen could not be confirmed
system node image update -node <node> -package http://... -setdefault true, node by nodeAn HTTP URL is still accepted. The manual install step is now one command for all nodes: system node image update -node * -package <location> -replace-package true -setdefault true -background true
Pre-checks by hand: volumes and aggregates online, jobs, storage errors show, metrocluster check run, the MAINT AutoSupportStill part of the documented flow. Automated upgrades run their own pre-checks and, in 9.19.1, report errors and warnings separately
Target 9.1P11, later 9.3P12The last release for the FAS8200 is 9.16.1. From 9.3 that is four stages: 9.3 to 9.7, which needs the 9.5 and 9.7 images, then 9.9.1, then 9.13.1, then 9.16.1

Anyone repeating this on a system that is on 9.3 or later should use the automated upgrade and keep this listing for understanding what it does underneath. See upgrade methods and upgrade paths.

← solutionz