NetApp 13 - Troubleshooting
NetApp Solution · Previous: ONTAP upgrade
The last article collects the incidents that left notes behind: what the system said, what it turned out to be, what was checked and what fixed it. Some of them end with a fix, two end where the notes stop, and I say which is which.
Overview
| When | Where | Symptom | Outcome |
|---|---|---|---|
| September 2017 | Site 1, installation | metrocluster configure failed, disks shown as unknown | Not in the notes |
| November 2017 | DC1-A-XNAS001 | A domain account could not be resolved for an administrator login | Data collected; continued in Article 7 |
| February 2018 | DC1-B-XNAS001 | vldb.aggrBladeID.missing | SVM resynchronised |
| February 2018 | DC1-A-XNAS001 | csm.sessionFailed | Patch upgrade |
| February 2018 | All four nodes of site 1 | cf.hwassist.socBindFailed | Partner address set |
| February 2018 | DC1-A-ANAS001 | nvmem.battery.wrongCharge | False positive |
| June 2019 | Site 2 | metrocluster check: lifs warning | LIF placement repaired |
Installation day: configure fails, disks unknown
The largest file in the notes is a console log of 21 September 2017, when the MetroCluster of site 1 was being configured for the first time, on ONTAP 9.1P2. It is a log of looking, with tab completions and typing mistakes, and it has no conclusion. What can be read from it is this.
On cluster A the configuration had just failed.
$ metrocluster operation showoutput 5 lines
Operation: configure
State: failed
Start Time: 9/21/2017 09:24:35
End Time: 9/21/2017 09:29:30
Errors: Configure command timed out. Error: An internal error occurred. Retry the operation. For further assistance, contact technical support. Try the command again with "-refresh true".On cluster B metrocluster show printed Configuration: invalid with configuration-error for both clusters, metrocluster check show refused to run because MetroCluster was not configured, and metrocluster node show listed the two local nodes as ready to configure and nothing for cluster A.
The search then went through every layer between the clusters, and each one was healthy.
| Checked | Command | Result |
|---|---|---|
| Cluster peering | cluster peer show -instance | Available, authentication ok, all four intercluster addresses active |
| Intercluster and management reachability | network ping -node <node> -destination <ip_address> | Alive in both directions |
| FC-VI interconnect | metrocluster interconnect adapter show | On cluster B, both FC-VI adapters of each node up at 16G |
| Bridges | storage bridge show | All four ATTO bridges monitored, status ok |
| Shelves | storage shelf show | All twelve shelves of both datacenters visible, status normal |
| Aggregates | storage aggregate show | Root and data aggregates of all four nodes online, mirrored, normal |
What was not healthy was the picture of the disks. All 272 disks of both datacenters were visible from both clusters, as they must be in a fabric MetroCluster, but their ownership was lopsided.
| Seen from | Owned by the nodes of A | Owned by the nodes of B |
|---|---|---|
| Cluster A | 239 disks: 28 in aggregates, 211 spare | 32 disks, container type unknown |
| Cluster B | 239 disks, container type unknown | 32 disks: 26 in aggregates, 6 spare |
One more disk was broken. That a cluster shows the other cluster's disks as unknown is not the fault; it can see them and does not own them. What stands out is the split. The nodes of cluster A owned every working disk of the ten full shelves, including shelves 50, 51, 60, 61 and 70, which stand in datacenter B. The nodes of cluster B owned only the sixteen disks in each of the two partly filled shelves, 31 and 71, and one of their two data aggregates was smaller than the other three in the MetroCluster.
Two smaller things were corrected on the way. Cluster B had two default routes, one through the management gateway and one through 10.11.23.254 in the network of the bridges of datacenter B; the second was deleted.
$ route delete -vserver DC1-B-XNAS001 -destination 0.0.0.0/0 -gateway 10.11.23.254
And the AutoSupport history showed the SMTP delivery in re-queued after twelve attempts; AutoSupport is the subject of Logging, monitoring and AutoSupport.
The log ends with aggr show on cluster B. How the ownership was redistributed and when metrocluster configure finally succeeded is not in it. Eight days later, on 29 September, a 9.1P8 image was installed on all four nodes, as the install dates in the upgrade notes show; whether the clusters were set up again in between I cannot tell from the Source material. The designed disk layout is in Storage design.
Administrators from Active Directory cannot log in
On 16 November 2017 the domain account admin01 could not be resolved on DC1-A-XNAS001. The notes of that day look like a data collection for a support case more than a diagnosis; they do not name a case.
First every cache of the security daemon was emptied on both nodes, 27 named caches per node, and the group-authentication cache was switched off so that every attempt would really go to the directory.
$ diag secd cache clear -node DC1-A-ANAS001 -vserver DC1-A-XNAS001 -cache-name ldap-username-to-creds $ security login ns-switch group-authentication cache timeout modify -timeout 0
A packet trace was started on the node for the two domain controllers, and then the lookup was run by hand.
$ node run -node DC1-A-ANAS001 pktt start all -i 10.11.16.209 -i 10.11.16.210 -d /etc/crash $ secd authentication show-ontap-admin-unix-creds -node DC1-A-ANAS001 -vserver DC1-A-XNAS001 -unix-user-name admin01
output 4 lines
[ 5002] TCP connection to ip 10.11.16.209, port 389 via interface 10.11.10.37 failed: Operation timed out. [ 5004] Successfully connected to ip 10.11.16.210, port 389 using TCP **[ 5136] FAILURE: User 'admin01' not found in UNIX authorization source LDAP. Error: command failed: Failed to get ONTAP admin UNIX credentials. Reason: "SecD Error: object not found".
Two different problems are in those four lines. The first domain controller did not answer on port 389 from the node management interface at all, which costs five seconds on every lookup. The second one answered, and did not return the user. getxxbyyy getpwbyname for the same user failed the same way.
The trace was stopped, the cache timeout set back to 10 minutes, and a full AutoSupport plus the service processor logs were sent.
$ node run -node DC1-A-ANAS001 pktt stop all $ system node autosupport invoke -type all -node * $ system node autosupport invoke-splog -remote-node DC1-A-ANAS001
The trace files land in /etc/crash on the node and were fetched through the node's web interface. The note stops there, without the answer.
The continuation is in a second file, a page of trial and error on the LDAP client configuration: StartTLS on and off, session security none, sign and seal, port 389 and 636, each followed by diag secd connections test and marked FAIL, then ldapsearch and tcpdump from a Linux host against the same domain controller. The design that was finally built tunnels administrator logins through a CIFS SVM that is joined to the domain. That story, with the working configuration, is told in Access and directory integration.
Four errors in the event log
On 19 and 20 February 2018, the days around the upgrade to 9.1P11, four messages from the event logs of site 1 were followed up.
vldb.aggrBladeID.missing
output 1 line
2/19/2018 07:57:56 DC1-B-ANAS001 ERROR vldb.aggrBladeID.missing: The volume 'DC1_S_VCVSM001_data' is located on the aggregate with UUID '<AGGREGATE_UUID>' whose owning dblade UUID '<NODE_UUID>' does not exist in the Volume Location Database.
Cluster B complained about a volume of an SVM that lives on cluster A. The volume location database of B had a record of the replicated volume that pointed to a node it did not know. The reference in the notes describes this state as a leftover of a forced MetroCluster switchover. Whether one of the tests of MetroCluster switchover and Tiebreaker preceded it the notes do not say.
The fix was to replicate the configuration of that SVM again, from the cluster that owns it.
$ metrocluster vserver resync -cluster DC1-A-XNAS001 -vserver DC1-S-VCVSM001 $ metrocluster vserver show
The second command listed every SVM pair as healthy. A second check compares the volume records of the file system with the location database, at diagnostic privilege.
$ set diag $ debug vreport show $ set admin
output 3 lines
This table is currently empty. Info: WAFL and VLDB volume/aggregate records are consistent.
csm.sessionFailed
output 1 line
2/19/2018 00:09:35 DC1-A-ANAS001 ERROR csm.sessionFailed: Cluster interconnect session (req=DC1-A-ANAS001:dblade, rsp=DC1-B-ANAS001:dblade, uniquifier=<SESSION_ID>) failed with record state ACTIVE and error CSM_CONNABORTED.
Sessions from a node of cluster A to both nodes of cluster B were aborted at the same second. The notes point to two bug reports of the vendor and give the repair in one line: upgrade from 9.1P8 to 9.1P11. That is the upgrade of ONTAP upgrade, done the next day. The notes do not record whether the message was seen again afterwards.
cf.hwassist.socBindFailed
output 2 lines
2/19/2018 08:01:56 DC1-A-ANAS001 ERROR cf.hwassist.localMonitor: hw_assist: hw_assist functionality is inactive. 2/19/2018 08:01:56 DC1-A-ANAS001 ERROR cf.hwassist.socBindFailed: hw_assist: bind failed to port 4444 on IP address <HWASSIST_IP>. Error 49
All four nodes logged this pair. Hardware-assisted takeover lets the service processor of a node tell the HA partner that the node has died, so that the partner does not have to wait for missed heartbeats. Each node listens for those alerts on UDP port 4444 of its node management address, and the partner has to be told that address. Here all four nodes tried to bind to the same address, which could not be the management address of each of them.
The node management addresses were confirmed in the node shell, where ifconfig -a marks them, and the state was read in both shells.
$ system node run -node DC1-A-ANAS001 $ ifconfig -a $ cf hw_assist status
At the cluster shell the same state has a column that says what to do.
$ storage failover hwassist showoutput 6 lines
DC1-A-ANAS001
Partner : DC1-A-ANAS002
Hwassist Enabled : true
Hwassist Port : 4444
Monitor Status : inactive
Inactive Reason : Fail to bind UDP socket.The fix is one command per node: on each node, the management address of its partner.
$ storage failover modify -hwassist-partner-ip 10.11.10.38 -node DC1-A-ANAS001 $ storage failover modify -hwassist-partner-ip 10.11.10.37 -node DC1-A-ANAS002
The same on cluster B with 10.11.10.40 for DC1-B-ANAS001 and 10.11.10.39 for DC1-B-ANAS002. Afterwards storage failover hwassist show printed Monitor Status : active and Keep Alive Status : healthy for all four nodes.
nvmem.battery.wrongCharge
output 1 line
2/20/2018 05:00:18 DC1-A-ANAS001 ALERT nvmem.battery.wrongCharge: The NVMEM battery charger is charging the battery even though the battery is not requesting to be charged. To prevent data loss, the system will shut down in 24 hours.
An alert that announces a shutdown in 24 hours, at five in the morning of the day of the upgrade. It is a known false positive on the FAS8200 and its sibling models, documented in a bug report of the vendor: during the battery's learning cycle, which runs every 70 days and takes about 19 hours, the alert can fire and is followed a short time later by a message that all is normal again.
The catch is that the all-clear is a NOTICE, and the event log at admin privilege does not show it.
$ set diagnostic $ event show
output 2 lines
2/20/2018 05:00:28 DC1-A-ANAS001 NOTICE nvmem.battery.normalCharge: The NVMEM battery charging status is normal. 2/20/2018 05:00:18 DC1-A-ANAS001 ALERT nvmem.battery.wrongCharge: The NVMEM battery charger is charging the battery ...
Ten seconds apart. Nothing was changed.
LIFs of the MetroCluster partner out of sync
On 5 June 2019, at site 2, the MetroCluster check had one component that was not ok.
$ metrocluster check run $ metrocluster check show
output 7 lines
Component Result ------------------- --------- nodes ok lifs warning config-replication ok aggregates ok clusters ok
metrocluster check lif show narrowed it down to one LIF, DC2-S-VCVSM005_nfs_lif1, check port-selection, on the stopped copy of the SVM on cluster B. The detailed view gave the reason.
output 5 lines
Name of the Cluster: DC2-B-XNAS001
Name of the Vserver: DC2-S-VCVSM005-mc
Name of the Lif: DC2-S-VCVSM005_nfs_lif1
Description: port-selection
Additional Information/Recovery Steps: Discovery of LIF with address 10.12.40.30 failed from destination cluster Ensure that the destination cluster has ports that have connectivity to the LIF on the source cluster.In plain words: if cluster B had to start this SVM in a switchover, it had no port on which it could reach the network of that address. The SVM is the one for the KVM hosts of platform C, in VLAN 2020 and an IPspace of its own, see Network design.
The notes begin with what has to be true before any repair, and it is all outside the MetroCluster commands.
- On the network switches, in both datacenters: MTU 9000 and the right VLANs on every port a NetApp node is connected to.
- On the clusters: the same network interfaces, Ethernet ports, broadcast domains and IPspaces on both.
network interface show and network port show on cluster A showed the LIF up on port a0a-2020, and that VLAN port present and healthy on both nodes in the broadcast domain vlan-2020. Then the placement of the LIF on the partner was repaired and the SVM configuration replicated again; the notes run the resync a second time with the name of the -mc copy.
$ metrocluster check lif repair-placement -vserver DC2-S-VCVSM005 -lif DC2-S-VCVSM005_nfs_lif1 $ metrocluster vserver resync -cluster DC2-A-XNAS001 -vserver DC2-S-VCVSM005
After that the detailed view had no recovery step left and metrocluster check lif show printed ok for port-selection. The whole sequence is a Config document: ONTAP: LIF resynchronisation procedure.
The notes do not say which of the prerequisites had been missing, or whether any was.
The commands that earned their place
| Question | Command | Privilege |
|---|---|---|
| Is the MetroCluster healthy as a whole? | metrocluster check run, then metrocluster check show | admin |
| What state are the two clusters in? | metrocluster show | admin |
| Why did the last MetroCluster operation fail or warn? | metrocluster operation show | admin |
| Which LIF would have no port after a switchover? | metrocluster check lif show -instance -lif <lif> | admin |
| Is the configuration of every SVM replicated? | metrocluster vserver show | admin |
| Are the mirrors still resynchronising? | storage aggregate show-resync-status -in-progress true | admin |
| What happened, in order? | event show | admin; diagnostic to see every message |
| Who owns which disk, and what is it used for? | storage disk show -container-type <type> | admin |
| Are bridges, switches and shelves seen? | storage bridge show, storage switch show, storage shelf show | admin |
| Are the FC-VI links between the clusters up? | metrocluster interconnect adapter show | admin |
| Can the HA partner be warned by the service processor? | storage failover hwassist show | admin |
| Do volume records agree with the location database? | debug vreport show | diagnostic |
| Why does a directory lookup of an administrator fail? | secd authentication show-ontap-admin-unix-creds | elevated, run at the ::*> prompt |
| What is on the wire? | node run -node <node> pktt start all -i <ip_address> -d /etc/crash | node shell |
| What does the node itself see? | system node run -node <node> -command sysconfig -a | admin |
The commands are as they were typed on ONTAP 9.1 to 9.3. Since 9.8 storage bridge is system bridge and storage switch is system switch fibre-channel.
Lessons
- Read the event log before it is urgent. None of the four February messages is recorded with an outage. One of them, hardware-assisted takeover, was a protection that was inactive on all four nodes.
- An alert without its all-clear is half a message. Check at diagnostic privilege before reacting to a shutdown warning.
- An SVM needs its network on both clusters. The notes' own checklist says it: the same ports, broadcast domains and IPspaces on both clusters, and the same VLANs and MTU on the switches of both datacenters. Run
metrocluster check runafter every network change. - Finish the note. The installation log and the login notes both stop before the answer.