LINUXOR.SK ... open source notes ...

NetApp 12 - ONTAP upgrade

category: solutionz · date: 2019-12-31 · updated: 2026-10-02 · author: LALA

NetApp Solution · Previous: MetroCluster switchover and Tiebreaker · Next: Troubleshooting

A four-node fabric MetroCluster is upgraded while it serves data, one node at a time, by taking each node over, letting it boot the new image and giving its storage back. This article walks through the upgrade as it was done at site 1 in February 2018, says what the later runbook for site 2 added, and is honest about which upgrades left no notes at all.

Which versions, when

The Source material shows four ONTAP versions, and a record of only one step between them.

WhenVersionWhere it is recorded
21 September 20179.1P2version in the console dump of the initial installation of site 1
29 September 20179.1P8Install date of the older image in system node image show, all four nodes of site 1
20 February 20189.1P11The upgrade described here; also the version in the as-built reports of site 1
November 20199.3P12The node tables of design document 0.2, for both sites

Only the step from 9.1P8 to 9.1P11 at site 1 is in the notes with its commands and outputs. How the clusters got from 9.1P2 to 9.1P8, and later from 9.1 to 9.3P12, is not written down anywhere in the Source material, and neither is the version site 2 was installed with. I do not fill those gaps.

The reason for the P11 upgrade is in the troubleshooting notes: on 19 February 2018 the event log held csm.sessionFailed errors between the nodes of the two clusters, and the remedy from the vendor's bug report was a newer patch release. See Troubleshooting.

The order of the nodes

A patch upgrade within 9.1 was done by hand with takeover and giveback. That was the documented way for a four-node MetroCluster on 9.2 or earlier; from 9.3 on the automated upgrade with cluster image update is the recommended one, and the runbook's last section says more. The rule that decides the order is the disaster-recovery pairing: the procedure upgrades a node and its DR partner in the other cluster together. The pairing is read from the cluster, not assumed.

bash
$ metrocluster node show -fields dr-partner
output 6 lines
dr-group-id cluster       node          dr-partner
----------- ------------- ------------- -------------
1           DC1-A-XNAS001 DC1-A-ANAS001 DC1-B-ANAS002
1           DC1-A-XNAS001 DC1-A-ANAS002 DC1-B-ANAS001
1           DC1-B-XNAS001 DC1-B-ANAS001 DC1-A-ANAS002
1           DC1-B-XNAS001 DC1-B-ANAS002 DC1-A-ANAS001

The pairs are crosswise: node 001 of A with node 002 of B. That gives two rounds.

mermaid
flowchart TB
  p["Pre-checks on both clusters"] --> i["Image copied to all four nodes, set as default"]
  i --> g["Automatic giveback disabled on all four nodes"]
  subgraph r1["Round 1"]
    t1["Takeover of DC1-A-ANAS001 by DC1-A-ANAS002"] --> t2["Takeover of DC1-B-ANAS002 by DC1-B-ANAS001"]
    t2 --> w1["Wait at least eight minutes"]
    w1 --> g1["Giveback to DC1-A-ANAS001"]
    g1 --> g2["Giveback to DC1-B-ANAS002"]
  end
  subgraph r2["Round 2"]
    t3["Takeover of DC1-A-ANAS002, version mismatch allowed"] --> t4["Takeover of DC1-B-ANAS001, version mismatch allowed"]
    t4 --> w2["Wait at least eight minutes"]
    w2 --> g3["Giveback to DC1-A-ANAS002"]
    g3 --> g4["Giveback to DC1-B-ANAS001"]
  end
  g --> t1
  g2 --> v1["Image check on both clusters"]
  v1 --> t3
  g4 --> v2["Image check, automatic giveback enabled again, post-checks"]

A node that is taken over reboots. Because the new image was set as its default, it comes back running the new version and waits for the giveback. During round 1 each cluster therefore runs one node on 9.1P11 and one on 9.1P8, and in round 2 the already upgraded node has to take over a partner with an older version, which is why those two takeovers carry an extra option. Today's documentation says the option is not required for patch upgrades, so this run did not need it. With a single DR group there is nothing more; an eight-node MetroCluster would repeat the whole thing for its second DR group.

The full sequence is a Config document: ONTAP: upgrade runbook.

Pre-checks

Every check was run on both clusters. None of them changes anything.

CheckCommandExpected
Snapshot copies per nodevolume snapshot show -node <node>Fewer than 20,000; the four nodes had 38, 28, 18 and 8
CPU loadnode run -node <node> -command sysstat -c 10 -x 3CPU below 50 % in all ten samples
Mirror resynchronisationstorage aggregate plex show -in-progress trueNo entries
LIFsnetwork interface showLIFs of the running SVMs up and at home
Aggregatesstorage aggregate show -state !onlineNo entries
Volumesvolume show -state !onlineNo entries
Jobsjob showNothing running or queued that touches aggregates, volumes or Snapshot copies
Back-end errors, failed disksstorage errors showThis table is currently empty.
MetroClustermetrocluster check run, then metrocluster check showok for all five components

The CPU limit exists because for the length of a takeover one node carries the load of two, and the runbook adds that no load should be added to the cluster until the upgrade is complete. If a mirrored aggregate is resynchronising, the upgrade waits until it has finished.

On DC1-A-XNAS001 the MetroCluster check printed the result that every later step relies on.

output 7 lines
Component           Result
------------------- ---------
nodes               ok
lifs                ok
config-replication  ok
aggregates          ok
clusters            ok

Pre-upgrade tasks

First the vendor was told. An AutoSupport message with a maintenance window keeps the support organisation from opening cases for the takeovers that follow. On each cluster:

bash
$ system node autosupport invoke -node * -type all -message "MAINT=2h Starting_NDU"
$ autosupport history show -subject *MAINT*

The history showed the HTTP delivery as sent-successful on all four nodes, and the SMTP delivery as re-queued.

Then the image. The package 91P11_q_image.tgz was offered by a web server on port 8080 of 10.11.10.44, the address of the Unified Manager host DC1-A-VCOCM001, and each node fetched and installed it into its alternate image slot. On cluster A, for each of its nodes, and the same on cluster B:

bash
$ system node image update -node DC1-A-ANAS001 -package http://10.11.10.44:8080/91P11_q_image.tgz -setdefault true
$ system node image show
output 6 lines
                 Is      Is                                Install
Node     Image   Default Current Version                   Date
-------- ------- ------- ------- ------------------------- -------------------
DC1-A-ANAS001
         image1  true    false   9.1P11                    2/20/2018 12:38:04
         image2  false   true    9.1P8                     9/29/2017 09:15:38

Is Default true, Is Current false is the state to look for: the node will boot 9.1P11 next time and is still running 9.1P8. The install dates of the four nodes run from 12:38 to 12:40.

Last, automatic giveback was switched off on all four nodes. Without that, a node that has booted the new version would get its aggregates back as soon as it is ready, and the operator would lose the pause in which clients settle.

bash
$ storage failover modify -node DC1-A-ANAS001 -auto-giveback false
$ storage failover show -fields auto-giveback

Takeover and giveback

Round 1 starts on cluster A.

bash
$ storage failover takeover -ofnode DC1-A-ANAS001
$ storage failover show
output 5 lines
                              Takeover
Node           Partner        Possible State Description
-------------- -------------- -------- -------------------------------------
DC1-A-ANAS001  DC1-A-ANAS002  -        Waiting for giveback
DC1-A-ANAS002  DC1-A-ANAS001  false    In takeover

The same is done on cluster B for the DR partner, DC1-B-ANAS002. Then nothing is done for at least eight minutes. The note in the runbook gives two reasons: client multipathing has to stabilise, and clients have to recover from the pause in I/O that a takeover causes. For NFS clients such as the KVM hosts the second one is what counts.

The giveback follows, first on A, then on B, each watched until it is complete.

bash
$ storage failover giveback -ofnode DC1-A-ANAS001
$ storage failover show-giveback
output 7 lines
               Partner
Node           Aggregate         Giveback Status
-------------- ----------------- ---------------------------------------------
DC1-A-ANAS001
                                 No aggregates to give back
DC1-A-ANAS002
                                 No aggregates to give back

system node image show on both clusters confirms that the two nodes of round 1 now run 9.1P11. Round 2 repeats the steps for DC1-A-ANAS002 and DC1-B-ANAS001, with one difference in the takeover.

bash
$ storage failover takeover -ofnode DC1-A-ANAS002 -option allow-versionmismatch

Post-upgrade tasks

Automatic giveback was enabled again on all four nodes and verified, metrocluster check run was repeated, and a closing AutoSupport message was sent from both clusters.

bash
$ storage failover modify -node DC1-A-ANAS001 -auto-giveback true
$ metrocluster check run
$ metrocluster check show
$ system node autosupport invoke -node * -type all -message "Finishing_NDU"

The final look at both clusters was four listings: storage errors show, vserver show, network interface show and network port show.

What the notes do not prove

The site 1 notes are a runbook filled in while working, and they are filled in unevenly.

Site 2

The file for site 2 is not a second transcript. It is the site 1 runbook copied and extended, still with the names of DC1, the 9.1P11 package and the February 2018 outputs of site 1 in it. It shows what the procedure had grown into when an upgrade of site 2 was prepared, and nothing about what happened there: no version, no date, no output from a DC2 cluster.

What it adds is all in the pre-checks.

Added before the upgradeCommand in the runbook
LDAP connection of the SVMs, required for a non-disruptive upgradeldap check -vserver <vserver>
Storage failover enabled and possiblestorage failover show
No inconsistent volumesvolume show -is-inconsistent true
Free space for deduplication metadata: 4 % in each deduplicated volume, 3 % in its aggregatedf -vserver <vserver> -volume <volume>
Failover policy and failover targets of every data LIFnetwork interface show -role data -failover
Data LIFs enabled and reverted to their home portsnetwork interface modify {-role data} -status-admin up, network interface revert *
No disks in maintenance or reconstructionstorage disk show -state with the states maintenance, pending, reconstructing

The LDAP check is the one with a history in this Solution: the clusters authenticate administrators through Active Directory, and that path had its own troubleshooting sessions, see Access and directory integration.

One thing is removed: in the site 2 file the round 2 takeovers no longer carry -option allow-versionmismatch. The notes do not say whether that was a decision or an omission.

Lessons

← solutionz