NetApp - ONTAP upgrade runbook for a four-node MetroCluster
NetApp Solution · Config document · referenced from ONTAP upgrade
The non-disruptive upgrade of a four-node fabric MetroCluster by manual takeover and giveback, as it was run at site 1: checks, preparation, two rounds of takeover and giveback in the order of the DR partnerships, and the tasks afterwards.
| Item | Value |
|---|---|
| Shown here for | Site 1, clusters DC1-A-XNAS001 and DC1-B-XNAS001, on 20 February 2018 |
| ONTAP version | From 9.1P8 to 9.1P11 |
| Image source | http://10.11.10.44:8080/, a web server on the Unified Manager host DC1-A-VCOCM001 |
| Also applied to | Site 2: a copy of this runbook with more pre-checks exists; no record of its run |
| Applied with | SSH to the cluster management address of each cluster |
The command set
# ############################################################################ # CHECKS BEFORE THE UPGRADE # ############################################################################ # --- Number of Snapshot copies per node, must be below 20,000 --------------- # On DC1-A-XNAS001 set advanced volume snapshot show -node DC1-A-ANAS001 volume snapshot show -node DC1-A-ANAS002 # On DC1-B-XNAS001 set advanced volume snapshot show -node DC1-B-ANAS001 volume snapshot show -node DC1-B-ANAS002 # --- CPU utilisation, must stay below 50 % in all 10 samples ---------------- # On DC1-A-XNAS001 node run -node DC1-A-ANAS001 -command sysstat -c 10 -x 3 node run -node DC1-A-ANAS002 -command sysstat -c 10 -x 3 # On DC1-B-XNAS001 node run -node DC1-B-ANAS001 -command sysstat -c 10 -x 3 node run -node DC1-B-ANAS002 -command sysstat -c 10 -x 3 # --- No mirrored aggregate may be resynchronising --------------------------- # On DC1-A-XNAS001; expected "There are no entries matching your query." storage aggregate plex show -in-progress true # --- LIFs, aggregates, volumes: on both clusters ---------------------------- # LIFs of the running SVMs up and on their home nodes network interface show # Anything listed here is not online storage aggregate show -state !online volume show -state !online # --- No aggregate, volume or Snapshot jobs running or queued ---------------- # On DC1-A-XNAS001. Prefer to wait; delete only a job that will not finish. job show job delete -id <job_id> # --- Back-end configuration errors and failed disks: on both clusters ------- # Expected "This table is currently empty." storage errors show # --- MetroCluster check ----------------------------------------------------- # On DC1-A-XNAS001; expected ok for nodes, lifs, config-replication, # aggregates and clusters metrocluster check run metrocluster check show # ############################################################################ # PRE-UPGRADE TASKS # ############################################################################ # --- AutoSupport: announce a two-hour maintenance window -------------------- # On DC1-A-XNAS001, then the same two commands on DC1-B-XNAS001 system node autosupport invoke -node * -type all -message "MAINT=2h Starting_NDU" autosupport history show -subject *MAINT* # --- DR partnerships decide the order --------------------------------------- # On DC1-A-XNAS001. Result: DC1-A-ANAS001 <-> DC1-B-ANAS002, # DC1-A-ANAS002 <-> DC1-B-ANAS001 metrocluster node show -fields dr-partner # --- Copy the image to all four nodes and make it the default --------------- # On DC1-A-XNAS001 system node image update -node DC1-A-ANAS001 -package http://10.11.10.44:8080/91P11_q_image.tgz -setdefault true system node image update -node DC1-A-ANAS002 -package http://10.11.10.44:8080/91P11_q_image.tgz -setdefault true # On DC1-B-XNAS001 system node image update -node DC1-B-ANAS001 -package http://10.11.10.44:8080/91P11_q_image.tgz -setdefault true system node image update -node DC1-B-ANAS002 -package http://10.11.10.44:8080/91P11_q_image.tgz -setdefault true # --- Confirm: 9.1P11 is default and not current, on both clusters ----------- system node image show system node image package show # --- Disable automatic giveback on all four nodes --------------------------- # On DC1-A-XNAS001 storage failover modify -node DC1-A-ANAS001 -auto-giveback false storage failover modify -node DC1-A-ANAS002 -auto-giveback false # On DC1-B-XNAS001 storage failover modify -node DC1-B-ANAS001 -auto-giveback false storage failover modify -node DC1-B-ANAS002 -auto-giveback false # Verify, on both clusters storage failover show -fields auto-giveback # ############################################################################ # ROUND 1: DC1-A-ANAS001 AND ITS DR PARTNER DC1-B-ANAS002 # ############################################################################ # --- Takeover of the first node; on DC1-A-XNAS001 --------------------------- # Expected: DC1-A-ANAS001 "Waiting for giveback", DC1-A-ANAS002 "In takeover" storage failover takeover -ofnode DC1-A-ANAS001 storage failover show # --- Takeover of its DR partner; on DC1-B-XNAS001 --------------------------- storage failover takeover -ofnode DC1-B-ANAS002 storage failover show # --- Wait at least eight minutes -------------------------------------------- # Client multipathing must stabilise and clients must recover from the pause # in I/O that the takeover caused. # --- Giveback; on DC1-A-XNAS001 --------------------------------------------- # Repeat the second command until "No aggregates to give back" storage failover giveback -ofnode DC1-A-ANAS001 storage failover show-giveback # --- Giveback; on DC1-B-XNAS001 --------------------------------------------- storage failover giveback -ofnode DC1-B-ANAS002 storage failover show-giveback # --- Both nodes of round 1 must now run the new image; on both clusters ----- system node image show # ############################################################################ # ROUND 2: DC1-A-ANAS002 AND ITS DR PARTNER DC1-B-ANAS001 # ############################################################################ # --- Takeover by a partner that already runs the new version ---------------- # On DC1-A-XNAS001 storage failover takeover -ofnode DC1-A-ANAS002 -option allow-versionmismatch storage failover show # On DC1-B-XNAS001 storage failover takeover -ofnode DC1-B-ANAS001 -option allow-versionmismatch storage failover show # --- Wait at least eight minutes, as in round 1 ----------------------------- # --- Giveback; on DC1-A-XNAS001 --------------------------------------------- storage failover giveback -ofnode DC1-A-ANAS002 storage failover show-giveback # --- Giveback; on DC1-B-XNAS001 --------------------------------------------- storage failover giveback -ofnode DC1-B-ANAS001 storage failover show-giveback # --- All four nodes must now run the new image; on both clusters ------------ system node image show # ############################################################################ # POST-UPGRADE TASKS # ############################################################################ # --- Enable automatic giveback again on all four nodes ---------------------- # On DC1-A-XNAS001 storage failover modify -node DC1-A-ANAS001 -auto-giveback true storage failover modify -node DC1-A-ANAS002 -auto-giveback true # On DC1-B-XNAS001 storage failover modify -node DC1-B-ANAS001 -auto-giveback true storage failover modify -node DC1-B-ANAS002 -auto-giveback true # Verify, on both clusters storage failover show -fields auto-giveback # --- MetroCluster check; on DC1-A-XNAS001 ----------------------------------- metrocluster check run metrocluster check show # --- AutoSupport: the upgrade is finished ----------------------------------- # On DC1-A-XNAS001, then the same two commands on DC1-B-XNAS001 system node autosupport invoke -node * -type all -message "Finishing_NDU" autosupport history show -subject *MAINT* # ############################################################################ # CHECKS AFTER THE UPGRADE: on both clusters # ############################################################################ storage errors show vserver show network interface show network port show
The runbook kept for site 2 is this one with the following differences. It still carries the node names and the image of site 1, so the table lists what was added or removed, not what was run at site 2.
| Place in the runbook | Site 1, as run | Site 2 runbook |
|---|---|---|
| Before the LIF check | Nothing | ldap check -vserver <vserver>, and ldap client modify -client-config <config> -ldap-servers <ip_address> if the check fails |
| Before the LIF check | Nothing | storage failover show: failover enabled and possible |
| After the volume check | Nothing | volume show -is-inconsistent true |
| After the volume check | Nothing | df -vserver <vserver> -volume <volume> for deduplicated volumes: 4 % free in the volume, 3 % in the aggregate |
| After the volume check | Nothing | network interface show -role data -failover: failover policy and targets of each data LIF |
| After the volume check | Nothing | network interface modify {-role data} -status-admin up and network interface revert * |
After storage errors show | Nothing | storage disk show -state for disks in maintenance, pending or reconstructing |
| Round 2 takeovers | With -option allow-versionmismatch | Without the option |
The reference the notes give for trouble during an upgrade is the vendor's article "FAQ: Troubleshooting Clustered Data ONTAP Upgrades".
Checked against ONTAP 9.19.1
| As built | Today |
|---|---|
| Manual takeover and giveback, one DR pair at a time | Documented only for four-node MetroCluster configurations on ONTAP 9.2 or earlier and for eight-node MetroCluster FC. The steps themselves are unchanged: automatic giveback off, DR pairs in lockstep, an eight-minute wait before giveback, automatic giveback on again |
| No automated method used | For a two- or four-node MetroCluster on 9.3 or later the recommended method is the automated non-disruptive upgrade, with System Manager or with cluster image package get, cluster image validate and cluster image update. With four nodes the automated upgrade starts on the HA pairs of both sites at the same time; validation is run on cluster A, then on cluster B |
-option allow-versionmismatch on the round 2 takeovers | The documented value is allow-version-mismatch, and it is "not required for upgrades from ONTAP 9.0 to ONTAP 9.1 or for any patch upgrades", so the 9.1P8 to 9.1P11 run did not need it. Whether the 9.1 command line accepted the spelling without the hyphen could not be confirmed |
system node image update -node <node> -package http://... -setdefault true, node by node | An HTTP URL is still accepted. The manual install step is now one command for all nodes: system node image update -node * -package <location> -replace-package true -setdefault true -background true |
Pre-checks by hand: volumes and aggregates online, jobs, storage errors show, metrocluster check run, the MAINT AutoSupport | Still part of the documented flow. Automated upgrades run their own pre-checks and, in 9.19.1, report errors and warnings separately |
| Target 9.1P11, later 9.3P12 | The last release for the FAS8200 is 9.16.1. From 9.3 that is four stages: 9.3 to 9.7, which needs the 9.5 and 9.7 images, then 9.9.1, then 9.13.1, then 9.16.1 |
Anyone repeating this on a system that is on 9.3 or later should use the automated upgrade and keep this listing for understanding what it does underneath. See upgrade methods and upgrade paths.