LINUXOR.SK ... open source notes ...

NetApp 13 - Troubleshooting

category: solutionz · date: 2019-12-31 · updated: 2026-10-02 · author: LALA

NetApp Solution · Previous: ONTAP upgrade

The last article collects the incidents that left notes behind: what the system said, what it turned out to be, what was checked and what fixed it. Some of them end with a fix, two end where the notes stop, and I say which is which.

Overview

WhenWhereSymptomOutcome
September 2017Site 1, installationmetrocluster configure failed, disks shown as unknownNot in the notes
November 2017DC1-A-XNAS001A domain account could not be resolved for an administrator loginData collected; continued in Article 7
February 2018DC1-B-XNAS001vldb.aggrBladeID.missingSVM resynchronised
February 2018DC1-A-XNAS001csm.sessionFailedPatch upgrade
February 2018All four nodes of site 1cf.hwassist.socBindFailedPartner address set
February 2018DC1-A-ANAS001nvmem.battery.wrongChargeFalse positive
June 2019Site 2metrocluster check: lifs warningLIF placement repaired

Installation day: configure fails, disks unknown

The largest file in the notes is a console log of 21 September 2017, when the MetroCluster of site 1 was being configured for the first time, on ONTAP 9.1P2. It is a log of looking, with tab completions and typing mistakes, and it has no conclusion. What can be read from it is this.

On cluster A the configuration had just failed.

bash
$ metrocluster operation show
output 5 lines
  Operation: configure
      State: failed
 Start Time: 9/21/2017 09:24:35
   End Time: 9/21/2017 09:29:30
     Errors: Configure command timed out. Error: An internal error occurred. Retry the operation. For further assistance, contact technical support. Try the command again with "-refresh true".

On cluster B metrocluster show printed Configuration: invalid with configuration-error for both clusters, metrocluster check show refused to run because MetroCluster was not configured, and metrocluster node show listed the two local nodes as ready to configure and nothing for cluster A.

The search then went through every layer between the clusters, and each one was healthy.

CheckedCommandResult
Cluster peeringcluster peer show -instanceAvailable, authentication ok, all four intercluster addresses active
Intercluster and management reachabilitynetwork ping -node <node> -destination <ip_address>Alive in both directions
FC-VI interconnectmetrocluster interconnect adapter showOn cluster B, both FC-VI adapters of each node up at 16G
Bridgesstorage bridge showAll four ATTO bridges monitored, status ok
Shelvesstorage shelf showAll twelve shelves of both datacenters visible, status normal
Aggregatesstorage aggregate showRoot and data aggregates of all four nodes online, mirrored, normal

What was not healthy was the picture of the disks. All 272 disks of both datacenters were visible from both clusters, as they must be in a fabric MetroCluster, but their ownership was lopsided.

Seen fromOwned by the nodes of AOwned by the nodes of B
Cluster A239 disks: 28 in aggregates, 211 spare32 disks, container type unknown
Cluster B239 disks, container type unknown32 disks: 26 in aggregates, 6 spare

One more disk was broken. That a cluster shows the other cluster's disks as unknown is not the fault; it can see them and does not own them. What stands out is the split. The nodes of cluster A owned every working disk of the ten full shelves, including shelves 50, 51, 60, 61 and 70, which stand in datacenter B. The nodes of cluster B owned only the sixteen disks in each of the two partly filled shelves, 31 and 71, and one of their two data aggregates was smaller than the other three in the MetroCluster.

Two smaller things were corrected on the way. Cluster B had two default routes, one through the management gateway and one through 10.11.23.254 in the network of the bridges of datacenter B; the second was deleted.

bash
$ route delete -vserver DC1-B-XNAS001 -destination 0.0.0.0/0 -gateway 10.11.23.254

And the AutoSupport history showed the SMTP delivery in re-queued after twelve attempts; AutoSupport is the subject of Logging, monitoring and AutoSupport.

The log ends with aggr show on cluster B. How the ownership was redistributed and when metrocluster configure finally succeeded is not in it. Eight days later, on 29 September, a 9.1P8 image was installed on all four nodes, as the install dates in the upgrade notes show; whether the clusters were set up again in between I cannot tell from the Source material. The designed disk layout is in Storage design.

Administrators from Active Directory cannot log in

On 16 November 2017 the domain account admin01 could not be resolved on DC1-A-XNAS001. The notes of that day look like a data collection for a support case more than a diagnosis; they do not name a case.

First every cache of the security daemon was emptied on both nodes, 27 named caches per node, and the group-authentication cache was switched off so that every attempt would really go to the directory.

bash
$ diag secd cache clear -node DC1-A-ANAS001 -vserver DC1-A-XNAS001 -cache-name ldap-username-to-creds
$ security login ns-switch group-authentication cache timeout modify -timeout 0

A packet trace was started on the node for the two domain controllers, and then the lookup was run by hand.

bash
$ node run -node DC1-A-ANAS001 pktt start all -i 10.11.16.209 -i 10.11.16.210 -d /etc/crash
$ secd authentication show-ontap-admin-unix-creds -node DC1-A-ANAS001 -vserver DC1-A-XNAS001 -unix-user-name admin01
output 4 lines
  [  5002] TCP connection to ip 10.11.16.209, port 389 via interface 10.11.10.37 failed: Operation timed out.
  [  5004] Successfully connected to ip 10.11.16.210, port 389 using TCP
**[  5136] FAILURE: User 'admin01' not found in UNIX authorization source LDAP.
Error: command failed: Failed to get ONTAP admin UNIX credentials. Reason: "SecD Error: object not found".

Two different problems are in those four lines. The first domain controller did not answer on port 389 from the node management interface at all, which costs five seconds on every lookup. The second one answered, and did not return the user. getxxbyyy getpwbyname for the same user failed the same way.

The trace was stopped, the cache timeout set back to 10 minutes, and a full AutoSupport plus the service processor logs were sent.

bash
$ node run -node DC1-A-ANAS001 pktt stop all
$ system node autosupport invoke -type all -node *
$ system node autosupport invoke-splog -remote-node DC1-A-ANAS001

The trace files land in /etc/crash on the node and were fetched through the node's web interface. The note stops there, without the answer.

The continuation is in a second file, a page of trial and error on the LDAP client configuration: StartTLS on and off, session security none, sign and seal, port 389 and 636, each followed by diag secd connections test and marked FAIL, then ldapsearch and tcpdump from a Linux host against the same domain controller. The design that was finally built tunnels administrator logins through a CIFS SVM that is joined to the domain. That story, with the working configuration, is told in Access and directory integration.

Four errors in the event log

On 19 and 20 February 2018, the days around the upgrade to 9.1P11, four messages from the event logs of site 1 were followed up.

vldb.aggrBladeID.missing

output 1 line
2/19/2018 07:57:56  DC1-B-ANAS001  ERROR  vldb.aggrBladeID.missing: The volume 'DC1_S_VCVSM001_data' is located on the aggregate with UUID '<AGGREGATE_UUID>' whose owning dblade UUID '<NODE_UUID>' does not exist in the Volume Location Database.

Cluster B complained about a volume of an SVM that lives on cluster A. The volume location database of B had a record of the replicated volume that pointed to a node it did not know. The reference in the notes describes this state as a leftover of a forced MetroCluster switchover. Whether one of the tests of MetroCluster switchover and Tiebreaker preceded it the notes do not say.

The fix was to replicate the configuration of that SVM again, from the cluster that owns it.

bash
$ metrocluster vserver resync -cluster DC1-A-XNAS001 -vserver DC1-S-VCVSM001
$ metrocluster vserver show

The second command listed every SVM pair as healthy. A second check compares the volume records of the file system with the location database, at diagnostic privilege.

bash
$ set diag
$ debug vreport show
$ set admin
output 3 lines
This table is currently empty.

Info: WAFL and VLDB volume/aggregate records are consistent.

csm.sessionFailed

output 1 line
2/19/2018 00:09:35  DC1-A-ANAS001  ERROR  csm.sessionFailed: Cluster interconnect session (req=DC1-A-ANAS001:dblade, rsp=DC1-B-ANAS001:dblade, uniquifier=<SESSION_ID>) failed with record state ACTIVE and error CSM_CONNABORTED.

Sessions from a node of cluster A to both nodes of cluster B were aborted at the same second. The notes point to two bug reports of the vendor and give the repair in one line: upgrade from 9.1P8 to 9.1P11. That is the upgrade of ONTAP upgrade, done the next day. The notes do not record whether the message was seen again afterwards.

cf.hwassist.socBindFailed

output 2 lines
2/19/2018 08:01:56  DC1-A-ANAS001  ERROR  cf.hwassist.localMonitor: hw_assist: hw_assist functionality is inactive.
2/19/2018 08:01:56  DC1-A-ANAS001  ERROR  cf.hwassist.socBindFailed: hw_assist: bind failed to port 4444 on IP address <HWASSIST_IP>. Error 49

All four nodes logged this pair. Hardware-assisted takeover lets the service processor of a node tell the HA partner that the node has died, so that the partner does not have to wait for missed heartbeats. Each node listens for those alerts on UDP port 4444 of its node management address, and the partner has to be told that address. Here all four nodes tried to bind to the same address, which could not be the management address of each of them.

The node management addresses were confirmed in the node shell, where ifconfig -a marks them, and the state was read in both shells.

bash
$ system node run -node DC1-A-ANAS001
$ ifconfig -a
$ cf hw_assist status

At the cluster shell the same state has a column that says what to do.

bash
$ storage failover hwassist show
output 6 lines
DC1-A-ANAS001
                             Partner : DC1-A-ANAS002
                    Hwassist Enabled : true
                       Hwassist Port : 4444
                      Monitor Status : inactive
                     Inactive Reason : Fail to bind UDP socket.

The fix is one command per node: on each node, the management address of its partner.

bash
$ storage failover modify -hwassist-partner-ip 10.11.10.38 -node DC1-A-ANAS001
$ storage failover modify -hwassist-partner-ip 10.11.10.37 -node DC1-A-ANAS002

The same on cluster B with 10.11.10.40 for DC1-B-ANAS001 and 10.11.10.39 for DC1-B-ANAS002. Afterwards storage failover hwassist show printed Monitor Status : active and Keep Alive Status : healthy for all four nodes.

nvmem.battery.wrongCharge

output 1 line
2/20/2018 05:00:18  DC1-A-ANAS001  ALERT  nvmem.battery.wrongCharge: The NVMEM battery charger is charging the battery even though the battery is not requesting to be charged. To prevent data loss, the system will shut down in 24 hours.

An alert that announces a shutdown in 24 hours, at five in the morning of the day of the upgrade. It is a known false positive on the FAS8200 and its sibling models, documented in a bug report of the vendor: during the battery's learning cycle, which runs every 70 days and takes about 19 hours, the alert can fire and is followed a short time later by a message that all is normal again.

The catch is that the all-clear is a NOTICE, and the event log at admin privilege does not show it.

bash
$ set diagnostic
$ event show
output 2 lines
2/20/2018 05:00:28  DC1-A-ANAS001  NOTICE  nvmem.battery.normalCharge: The NVMEM battery charging status is normal.
2/20/2018 05:00:18  DC1-A-ANAS001  ALERT   nvmem.battery.wrongCharge: The NVMEM battery charger is charging the battery ...

Ten seconds apart. Nothing was changed.

LIFs of the MetroCluster partner out of sync

On 5 June 2019, at site 2, the MetroCluster check had one component that was not ok.

bash
$ metrocluster check run
$ metrocluster check show
output 7 lines
Component           Result
------------------- ---------
nodes               ok
lifs                warning
config-replication  ok
aggregates          ok
clusters            ok

metrocluster check lif show narrowed it down to one LIF, DC2-S-VCVSM005_nfs_lif1, check port-selection, on the stopped copy of the SVM on cluster B. The detailed view gave the reason.

output 5 lines
  Name of the Cluster: DC2-B-XNAS001
  Name of the Vserver: DC2-S-VCVSM005-mc
      Name of the Lif: DC2-S-VCVSM005_nfs_lif1
          Description: port-selection
Additional Information/Recovery Steps: Discovery of LIF with address 10.12.40.30 failed from destination cluster Ensure that the destination cluster has ports that have connectivity to the LIF on the source cluster.

In plain words: if cluster B had to start this SVM in a switchover, it had no port on which it could reach the network of that address. The SVM is the one for the KVM hosts of platform C, in VLAN 2020 and an IPspace of its own, see Network design.

The notes begin with what has to be true before any repair, and it is all outside the MetroCluster commands.

network interface show and network port show on cluster A showed the LIF up on port a0a-2020, and that VLAN port present and healthy on both nodes in the broadcast domain vlan-2020. Then the placement of the LIF on the partner was repaired and the SVM configuration replicated again; the notes run the resync a second time with the name of the -mc copy.

bash
$ metrocluster check lif repair-placement -vserver DC2-S-VCVSM005 -lif DC2-S-VCVSM005_nfs_lif1
$ metrocluster vserver resync -cluster DC2-A-XNAS001 -vserver DC2-S-VCVSM005

After that the detailed view had no recovery step left and metrocluster check lif show printed ok for port-selection. The whole sequence is a Config document: ONTAP: LIF resynchronisation procedure.

The notes do not say which of the prerequisites had been missing, or whether any was.

The commands that earned their place

QuestionCommandPrivilege
Is the MetroCluster healthy as a whole?metrocluster check run, then metrocluster check showadmin
What state are the two clusters in?metrocluster showadmin
Why did the last MetroCluster operation fail or warn?metrocluster operation showadmin
Which LIF would have no port after a switchover?metrocluster check lif show -instance -lif <lif>admin
Is the configuration of every SVM replicated?metrocluster vserver showadmin
Are the mirrors still resynchronising?storage aggregate show-resync-status -in-progress trueadmin
What happened, in order?event showadmin; diagnostic to see every message
Who owns which disk, and what is it used for?storage disk show -container-type <type>admin
Are bridges, switches and shelves seen?storage bridge show, storage switch show, storage shelf showadmin
Are the FC-VI links between the clusters up?metrocluster interconnect adapter showadmin
Can the HA partner be warned by the service processor?storage failover hwassist showadmin
Do volume records agree with the location database?debug vreport showdiagnostic
Why does a directory lookup of an administrator fail?secd authentication show-ontap-admin-unix-credselevated, run at the ::*> prompt
What is on the wire?node run -node <node> pktt start all -i <ip_address> -d /etc/crashnode shell
What does the node itself see?system node run -node <node> -command sysconfig -aadmin

The commands are as they were typed on ONTAP 9.1 to 9.3. Since 9.8 storage bridge is system bridge and storage switch is system switch fibre-channel.

Lessons

← solutionz