NetApp 12 - ONTAP upgrade
NetApp Solution · Previous: MetroCluster switchover and Tiebreaker · Next: Troubleshooting
A four-node fabric MetroCluster is upgraded while it serves data, one node at a time, by taking each node over, letting it boot the new image and giving its storage back. This article walks through the upgrade as it was done at site 1 in February 2018, says what the later runbook for site 2 added, and is honest about which upgrades left no notes at all.
Which versions, when
The Source material shows four ONTAP versions, and a record of only one step between them.
| When | Version | Where it is recorded |
|---|---|---|
| 21 September 2017 | 9.1P2 | version in the console dump of the initial installation of site 1 |
| 29 September 2017 | 9.1P8 | Install date of the older image in system node image show, all four nodes of site 1 |
| 20 February 2018 | 9.1P11 | The upgrade described here; also the version in the as-built reports of site 1 |
| November 2019 | 9.3P12 | The node tables of design document 0.2, for both sites |
Only the step from 9.1P8 to 9.1P11 at site 1 is in the notes with its commands and outputs. How the clusters got from 9.1P2 to 9.1P8, and later from 9.1 to 9.3P12, is not written down anywhere in the Source material, and neither is the version site 2 was installed with. I do not fill those gaps.
The reason for the P11 upgrade is in the troubleshooting notes: on 19 February 2018 the event log held csm.sessionFailed errors between the nodes of the two clusters, and the remedy from the vendor's bug report was a newer patch release. See Troubleshooting.
The order of the nodes
A patch upgrade within 9.1 was done by hand with takeover and giveback. That was the documented way for a four-node MetroCluster on 9.2 or earlier; from 9.3 on the automated upgrade with cluster image update is the recommended one, and the runbook's last section says more. The rule that decides the order is the disaster-recovery pairing: the procedure upgrades a node and its DR partner in the other cluster together. The pairing is read from the cluster, not assumed.
$ metrocluster node show -fields dr-partner
output 6 lines
dr-group-id cluster node dr-partner ----------- ------------- ------------- ------------- 1 DC1-A-XNAS001 DC1-A-ANAS001 DC1-B-ANAS002 1 DC1-A-XNAS001 DC1-A-ANAS002 DC1-B-ANAS001 1 DC1-B-XNAS001 DC1-B-ANAS001 DC1-A-ANAS002 1 DC1-B-XNAS001 DC1-B-ANAS002 DC1-A-ANAS001
The pairs are crosswise: node 001 of A with node 002 of B. That gives two rounds.
flowchart TB p["Pre-checks on both clusters"] --> i["Image copied to all four nodes, set as default"] i --> g["Automatic giveback disabled on all four nodes"] subgraph r1["Round 1"] t1["Takeover of DC1-A-ANAS001 by DC1-A-ANAS002"] --> t2["Takeover of DC1-B-ANAS002 by DC1-B-ANAS001"] t2 --> w1["Wait at least eight minutes"] w1 --> g1["Giveback to DC1-A-ANAS001"] g1 --> g2["Giveback to DC1-B-ANAS002"] end subgraph r2["Round 2"] t3["Takeover of DC1-A-ANAS002, version mismatch allowed"] --> t4["Takeover of DC1-B-ANAS001, version mismatch allowed"] t4 --> w2["Wait at least eight minutes"] w2 --> g3["Giveback to DC1-A-ANAS002"] g3 --> g4["Giveback to DC1-B-ANAS001"] end g --> t1 g2 --> v1["Image check on both clusters"] v1 --> t3 g4 --> v2["Image check, automatic giveback enabled again, post-checks"]
A node that is taken over reboots. Because the new image was set as its default, it comes back running the new version and waits for the giveback. During round 1 each cluster therefore runs one node on 9.1P11 and one on 9.1P8, and in round 2 the already upgraded node has to take over a partner with an older version, which is why those two takeovers carry an extra option. Today's documentation says the option is not required for patch upgrades, so this run did not need it. With a single DR group there is nothing more; an eight-node MetroCluster would repeat the whole thing for its second DR group.
The full sequence is a Config document: ONTAP: upgrade runbook.
Pre-checks
Every check was run on both clusters. None of them changes anything.
| Check | Command | Expected |
|---|---|---|
| Snapshot copies per node | volume snapshot show -node <node> | Fewer than 20,000; the four nodes had 38, 28, 18 and 8 |
| CPU load | node run -node <node> -command sysstat -c 10 -x 3 | CPU below 50 % in all ten samples |
| Mirror resynchronisation | storage aggregate plex show -in-progress true | No entries |
| LIFs | network interface show | LIFs of the running SVMs up and at home |
| Aggregates | storage aggregate show -state !online | No entries |
| Volumes | volume show -state !online | No entries |
| Jobs | job show | Nothing running or queued that touches aggregates, volumes or Snapshot copies |
| Back-end errors, failed disks | storage errors show | This table is currently empty. |
| MetroCluster | metrocluster check run, then metrocluster check show | ok for all five components |
The CPU limit exists because for the length of a takeover one node carries the load of two, and the runbook adds that no load should be added to the cluster until the upgrade is complete. If a mirrored aggregate is resynchronising, the upgrade waits until it has finished.
On DC1-A-XNAS001 the MetroCluster check printed the result that every later step relies on.
output 7 lines
Component Result ------------------- --------- nodes ok lifs ok config-replication ok aggregates ok clusters ok
Pre-upgrade tasks
First the vendor was told. An AutoSupport message with a maintenance window keeps the support organisation from opening cases for the takeovers that follow. On each cluster:
$ system node autosupport invoke -node * -type all -message "MAINT=2h Starting_NDU" $ autosupport history show -subject *MAINT*
The history showed the HTTP delivery as sent-successful on all four nodes, and the SMTP delivery as re-queued.
Then the image. The package 91P11_q_image.tgz was offered by a web server on port 8080 of 10.11.10.44, the address of the Unified Manager host DC1-A-VCOCM001, and each node fetched and installed it into its alternate image slot. On cluster A, for each of its nodes, and the same on cluster B:
$ system node image update -node DC1-A-ANAS001 -package http://10.11.10.44:8080/91P11_q_image.tgz -setdefault true $ system node image show
output 6 lines
Is Is Install
Node Image Default Current Version Date
-------- ------- ------- ------- ------------------------- -------------------
DC1-A-ANAS001
image1 true false 9.1P11 2/20/2018 12:38:04
image2 false true 9.1P8 9/29/2017 09:15:38Is Default true, Is Current false is the state to look for: the node will boot 9.1P11 next time and is still running 9.1P8. The install dates of the four nodes run from 12:38 to 12:40.
Last, automatic giveback was switched off on all four nodes. Without that, a node that has booted the new version would get its aggregates back as soon as it is ready, and the operator would lose the pause in which clients settle.
$ storage failover modify -node DC1-A-ANAS001 -auto-giveback false $ storage failover show -fields auto-giveback
Takeover and giveback
Round 1 starts on cluster A.
$ storage failover takeover -ofnode DC1-A-ANAS001 $ storage failover show
output 5 lines
Takeover Node Partner Possible State Description -------------- -------------- -------- ------------------------------------- DC1-A-ANAS001 DC1-A-ANAS002 - Waiting for giveback DC1-A-ANAS002 DC1-A-ANAS001 false In takeover
The same is done on cluster B for the DR partner, DC1-B-ANAS002. Then nothing is done for at least eight minutes. The note in the runbook gives two reasons: client multipathing has to stabilise, and clients have to recover from the pause in I/O that a takeover causes. For NFS clients such as the KVM hosts the second one is what counts.
The giveback follows, first on A, then on B, each watched until it is complete.
$ storage failover giveback -ofnode DC1-A-ANAS001 $ storage failover show-giveback
output 7 lines
Partner
Node Aggregate Giveback Status
-------------- ----------------- ---------------------------------------------
DC1-A-ANAS001
No aggregates to give back
DC1-A-ANAS002
No aggregates to give backsystem node image show on both clusters confirms that the two nodes of round 1 now run 9.1P11. Round 2 repeats the steps for DC1-A-ANAS002 and DC1-B-ANAS001, with one difference in the takeover.
$ storage failover takeover -ofnode DC1-A-ANAS002 -option allow-versionmismatchPost-upgrade tasks
Automatic giveback was enabled again on all four nodes and verified, metrocluster check run was repeated, and a closing AutoSupport message was sent from both clusters.
$ storage failover modify -node DC1-A-ANAS001 -auto-giveback true $ metrocluster check run $ metrocluster check show $ system node autosupport invoke -node * -type all -message "Finishing_NDU"
The final look at both clusters was four listings: storage errors show, vserver show, network interface show and network port show.
What the notes do not prove
The site 1 notes are a runbook filled in while working, and they are filled in unevenly.
- The outputs of the pre-checks, the image installation, the first takeover and the first two givebacks are there. For the takeover of
DC1-B-ANAS002, both givebacks of round 2 and everysystem node image showafter a round, the command is there and the output is not. - The
metrocluster check showafter the upgrade carriesLast Checked On: 2/18/2018 16:43:15, two days before the timestamps of the upgrade itself. It comes from an earlier run of the check and is not evidence of the state after the upgrade. - The closing AutoSupport is verified with a history filter on the subject pattern MAINT, but its subject is
Finishing_NDUand contains noMAINT. As written, the filter can only have shown the opening message again. - No problem during the takeovers or givebacks is recorded, and no duration of the client pause.
Site 2
The file for site 2 is not a second transcript. It is the site 1 runbook copied and extended, still with the names of DC1, the 9.1P11 package and the February 2018 outputs of site 1 in it. It shows what the procedure had grown into when an upgrade of site 2 was prepared, and nothing about what happened there: no version, no date, no output from a DC2 cluster.
What it adds is all in the pre-checks.
| Added before the upgrade | Command in the runbook |
|---|---|
| LDAP connection of the SVMs, required for a non-disruptive upgrade | ldap check -vserver <vserver> |
| Storage failover enabled and possible | storage failover show |
| No inconsistent volumes | volume show -is-inconsistent true |
| Free space for deduplication metadata: 4 % in each deduplicated volume, 3 % in its aggregate | df -vserver <vserver> -volume <volume> |
| Failover policy and failover targets of every data LIF | network interface show -role data -failover |
| Data LIFs enabled and reverted to their home ports | network interface modify {-role data} -status-admin up, network interface revert * |
| No disks in maintenance or reconstruction | storage disk show -state with the states maintenance, pending, reconstructing |
The LDAP check is the one with a history in this Solution: the clusters authenticate administrators through Active Directory, and that path had its own troubleshooting sessions, see Access and directory integration.
One thing is removed: in the site 2 file the round 2 takeovers no longer carry -option allow-versionmismatch. The notes do not say whether that was a decision or an omission.
Lessons
- Keep the output with the command. A runbook that holds the command and not the answer cannot be told apart, a year later, from a runbook that was never run.
- Copy the runbook, then rename everything in it. The site 2 file with site 1 names is one paste away from a takeover on the wrong cluster.
- Write the version history down once. Three of the four versions above had to be reconstructed from install dates and a design table.