LINUXOR.SK ... open source notes ...

NetApp 09 - Logging, monitoring and AutoSupport

category: solutionz · date: 2019-12-31 · updated: 2026-10-02 · author: LALA

NetApp Solution · Previous: Encryption and certificates · Next: Unified Manager and API Services

A MetroCluster that nobody watches fails over quietly and then fails for real. This article covers everything the storage sends out about itself and everything that asks it questions: EMS events to the central syslog server, SNMPv3 for Sensu, the SNMP polling the cluster itself does on its switches and bridges, AutoSupport to the vendor, the generator that produced the as-built reports, and how a log bundle was sent to support.

Who talks to whom

The conceptual design drew the storage in the middle of the infrastructure services it depends on. Reduced to the logging, monitoring and management flows, it looks like this.

mermaid
flowchart LR
  adm["Administrators and operators"]
  subgraph mg["NetApp management"]
    ocum["Unified Manager"]
  end
  subgraph st["MetroCluster, site 1"]
    cl["Clusters A and B"]
    sw["Brocade FC switches"]
    br["ATTO bridges"]
  end
  subgraph inf["Infrastructure services"]
    sys["Central syslog"]
    sensu["Sensu"]
    mail["Mail server"]
    prx["HTTP proxy"]
  end
  sup["NetApp support"]
  adm -- "SSH, HTTPS" --> ocum
  adm -- "SSH, HTTPS" --> cl
  ocum -- "HTTPS" --> cl
  sensu -- "HTTPS, REST" --> ocum
  sensu -- "SNMPv3" --> cl
  cl -- "SNMP, health monitor" --> sw
  cl -- "SNMP, health monitor" --> br
  cl -- "syslog" --> sys
  sw -- "syslog" --> sys
  cl -- "AutoSupport, SMTP" --> mail
  cl -- "AutoSupport, HTTPS" --> prx
  prx -- "HTTPS" --> sup

The same flows as the firewall saw them at site 1. Addresses that the two design versions give differently (the mail server and the proxy) are left out.

FromToPortWhat
Clusters, FC switchesCentral syslog dc1-s-xcsys001.adm.example.net, 10.11.17.113syslogEMS events, Fabric OS log
Sensu dc1-a-vcsns001.adm.example.net, 10.11.16.129Cluster and node management LIFsUDP 161SNMPv3 queries
Cluster management LIF 10.11.10.33 and 10.11.10.34FC switches and ATTO bridges, out-of-bandUDP 161Health monitoring by the cluster itself
Nodesdc1-a-vcmsx001.adm.example.netTCP 25, 465AutoSupport mail
Nodesdc1-a-vcprx001.adm.example.netTCP 3128AutoSupport over HTTPS to the vendor
SensuAPI Services on DC1-A-VCOCM002TCP 8443REST, see Unified Manager and API Services

Site 2 has its own syslog server (dc2-s-xcsys001.adm.example.net, 10.12.17.113) and its own three Sensu servers. The notes hold the commands for site 1 only; site 2 was configured the same way against its own servers.

EMS events to the central syslog server

ONTAP writes its events to the Event Management System, EMS. The design asked for five things: include the severities INFORMATIONAL, NOTICE, ERROR, ALERT and EMERGENCY, exclude DEBUG, use TCP so that the server receives the messages, send without TLS "because central syslog server does not have implemented encryption", and address the server by IPv4.

ONTAP 9.1 had two ways to do this and the notes show both. The current one has three objects: a filter that selects events, a destination, and a notification that ties one to the other. The first configuration, on both clusters, used the built-in filter important-events.

bash
$ event notification destination create -name syslog -syslog 10.11.17.113
$ event notification create -filter-name important-events -destinations syslog

That filter passes only what ONTAP itself considers important, which is less than the design asked for. A filter of our own followed, with one include rule and one exclude rule. The notes give two ways to use it, a new notification or the existing one pointed at the new filter; the output below shows notification 2 with the new filter.

bash
$ event filter create -filter-name all-events
$ event filter rule add -filter-name all-events -type include -severity INFORMATIONAL,NOTICE,ERROR,ALERT,EMERGENCY
$ event filter rule add -filter-name all-events -type exclude -severity DEBUG
$ event notification modify -ID 2 -filter-name all-events
$ event notification show
output 4 lines
ID   Filter Name                     Destinations
---- ------------------------------  -----------------
1    default-trap-events             snmp-traphost
2    all-events                      syslog

The older way, which the notes head "old way (deprecated)", is a named destination and a route for every message name. It is in the notes too, and event destination show is where a surprise turned up: next to the factory destinations there was one named after the Unified Manager host, with that host as syslog target. It looks as if Unified Manager registered itself there when the cluster was added to it; the notes only show the entry.

bash
$ event destination show
$ event destination create -name example-syslog -syslog 10.11.17.113
$ event route add-destinations -messagename * -destination example-syslog

Two things are missing in the notes. The destination command has no protocol parameter, so the notes do not show where the TCP of the design was chosen; and cluster log-forwarding, the separate mechanism that forwards the audit log of management commands, does not appear at all. I cannot say from the Source material that command history ever reached the syslog server. The full set is a Config document: ONTAP: EMS events to syslog.

The Brocade switches log to the same server. On each switch the syslog server that was already configured was removed, the facility changed from LOG_LOCAL7 to LOG_LOCAL2 and the central server set.

bash
$ syslogadmin --remove -ip 10.13.13.207
$ syslogadmin --set -facility 2
$ syslogadmin --set -ip 10.11.17.113
$ syslogadmin --show -ip

SNMPv3 for Sensu

The central monitoring system was Sensu. The design gave it two roads into the storage: SNMP straight to the clusters for basic status, and REST through the management software, which is the subject of the next article.

On each cluster a local user sensu was created for the application snmp with the role readonly. The command, on cluster DC1-A-XNAS001, asks for the rest interactively.

bash
$ security login create -username sensu -application snmp -authentication-method usm -role readonly
PromptAnswer
Authoritative entity's EngineIDEnter, the local EngineID
Authentication protocol (none, md5, sha)sha
Authentication password<SNMP_AUTH_PASSWORD>
Privacy protocol (none, des)des
Privacy password<SNMP_PRIV_PASSWORD>

So the security level is authPriv, with SHA-1 and DES. DES was not a choice: the prompt of ONTAP 9.1 offers none or des. Beside DES the notes remark "the default is AES", which I read as the default of the monitoring side. No community and no trap host were configured; the as-built report of March 2018 shows the SNMP table of both clusters with empty communities, no trap hosts and traps disabled. Sensu polls, the cluster does not push. The notes say the queries should preferably come from the Sensu server's IPv6 address, with the IPv4 address as the alternative; which one was used in the end is not recorded.

The command set is a Config document: ONTAP: SNMPv3 user for Sensu.

What the cluster polls: switches and bridges

In a fabric-attached MetroCluster the cluster monitors its own plumbing. Each cluster polls the four FC switches and the four FibreBridges of the site over SNMP from its cluster management LIF, and raises health alerts when a switch or bridge stops answering. That is why the firewall tables of the design have UDP 161 from the cluster to the out-of-band addresses of switches and bridges.

bash
$ storage bridge show
$ storage switch show
output 6 lines
Bridge                   Symbolic Name Monitored Status  Vendor Model
------------------------ ------------- --------- ------- ------ -----------------
ATTO_10.11.15.47         bridgeA1      true      ok      Atto   FibreBridge 7500N
ATTO_10.11.15.49         bridgeA2      true      ok      Atto   FibreBridge 7500N
ATTO_10.11.23.47         bridgeB1      true      ok      Atto   FibreBridge 7500N
ATTO_10.11.23.49         bridgeB2      true      ok      Atto   FibreBridge 7500N

This polling is SNMPv1 with a community string, and it collides with hardening the switches. The notes hold two passes over the Brocade SNMP configuration, and they contradict each other.

PassSNMPv1Security level for GETSecurity level for SET
First, the hardeningDisabled with snmpconfig --disable snmpv12, authentication and privacy3, no access
Second, headed "SNMPv1 configuration"Left enabled, the six default community strings replaced0, no security3, no access

The first pass is what one wants on paper: v1 off, GET only with authPriv, SET forbidden, and snmpconfig --set snmpv3 to define the users (the dialogue of that command is not in the notes; the factory state was six users, snmpadmin1 to snmpadmin3 and snmpuser1 to snmpuser3, all with noAuth and noPriv). The second pass is what the cluster's health monitor needs: it speaks v1. The default communities were replaced on the switches, and both clusters were told the new read-only community for all four switches.

bash
$ snmpconfig --set secLevel -snmpget 0 -snmpset 3
$ snmpconfig --set snmpv1
$ configcommit

Then, on cluster DC1-A-XNAS001 and again on DC1-B-XNAS001, once per switch.

bash
$ storage switch modify -switch-name Brocade_10.11.15.45 -snmp-community <SNMP_COMMUNITY>
$ storage switch show -switch-name Brocade_10.11.15.45 -fields snmp-version,snmp-community

The output of the second command in the notes says SNMPv1, so the second pass is the one that stayed. The notes do not say in words that the first pass was rolled back because of the health monitor; that is my reading of the order and of the result. What remained of the hardening is that SET is forbidden and that nobody can read the switches with the factory community. The switch side is a Config document: Fabric OS: SNMP and syslog.

AutoSupport

AutoSupport is the node's habit of sending its state to the vendor: periodically, on events and on demand. It is configured per node, and the same command with thirty-odd parameters was run for all four nodes of site 1. The interesting parameters are these.

ParameterValueMeaning
-transporthttpsTo the vendor over HTTPS, not by mail
-proxy-urldc1-a-vcprx001.adm.example.net:3128The storage has no direct way out
-mail-hostsdc1-a-vcmsx001.adm.example.netFor copies to our own addresses
-fromnetapp@ad.example.net in the notes, netapp_mail@ad.example.net in the as-built reportChanged later to the mail account that existed
-to, -notetoops@example.netThree internal recipients
-supportdisable in the commandWith a note beside the output: must be changed to enable
-remove-private-datatrue in the commandSwitched to false "temporarily" for a support case
-ondemand-statedisableNo remote triggering by the vendor

The command in the notes is the cautious first version: delivery to the vendor off, private data removed, AutoSupport OnDemand off. OnDemand lets support ask a node to send a new or more detailed message; the node polls the vendor for such requests every 60 minutes, and with the state disabled it does not. Beside the output the notes carry two reminders: vendor support has to be changed to enable, and certificate validation has to be changed to false. The as-built report of March 2018 shows where the first one ended: support enabled, performance data enabled, private data not removed, throttle on, transport https through the proxy. The second is not visible in the report. The proxy had its own CA certificate, which the management servers had to trust too, and that would explain the wish to switch validation off; the notes do not say whether it was done.

Hiding private data did not survive the first support case. A line marked "temporarily enable the sending of sensitive data" sets -remove-private-data false on one node, and the report half a year later still shows it off, on all four nodes.

bash
$ system node autosupport modify -node DC1-A-ANAS001 -remove-private-data false

How it was tested

A test message was triggered by hand on 13 September 2017; autosupport show records it as the last subject sent, USER_TRIGGERED (TEST:TEST-20170913). A week later, during a troubleshooting session, the history showed what really worked.

bash
$ autosupport history show
output 11 lines
             Seq                                    Attempt  Percent  Last
Node         Num   Destination Status               Count    Complete Update
------------ ----- ----------- -------------------- -------- -------- --------
DC1-A-ANAS001M 195
                   smtp        re-queued            12       -        9/21/2017 10:03:54
                   http        sent-successful      1        100      9/21/2017 09:19:53
                   noteto      ignore               1        -        9/21/2017 09:18:19
               189
                   smtp        transmission-failed  15       -        9/20/2017 16:20:12
                   http        sent-successful      3        100      9/20/2017 15:29:22
                   noteto      transmission-failed  15       -        9/20/2017 16:20:12

HTTPS through the proxy delivered every time. Mail did not: fifteen attempts, the configured retry count, and then failed. The older design document has the likely reason in one sentence: the mail server accepted mail on port 465 with authentication, and "other Netapp components do not support SMTP authentication and due to this limitation will be not integrated with Email services". Only Unified Manager could log in to the mail server. The nodes kept their mail host setting anyway. I have no later history in the notes that would show mail from the nodes arriving.

The same mechanism was used later as a maintenance signal: before an ONTAP upgrade a message with MAINT=2h in the subject tells the vendor not to open cases for the next two hours. That is in ONTAP upgrade. The command set is a Config document: ONTAP: AutoSupport.

As-built documentation with NetAppDocs

The as-built reports quoted above were not written by hand. NetAppDocs is a PowerShell module from the vendor's tool chest that reads a cluster through its API and produces a Word document and an Excel workbook. It was run from a Windows workstation, in a PowerShell started as administrator, once per cluster for the cluster view and once per cluster for the SVM view.

bash
$ Set-ExecutionPolicy -ExecutionPolicy RemoteSigned -Scope CurrentUser
$ Import-Module NetAppDocs
$ $Credential = Get-Credential
$ Get-NtapClusterData -Name 'DC1-A-XNAS001.adm.example.net' -Credential $Credential | Format-NtapClusterData | Out-NtapDocument -WordFile 'C:\Netapp-doc\DC1-A-XNAS001.Docx' -ExcelFile 'C:\Netapp-doc\DC1-A-XNAS001.xlsx'
$ Get-NtapVserverData -ClusterName 'DC1-A-XNAS001.adm.example.net' -Credential $Credential | Format-NtapVserverData | Out-NtapDocument -WordFile 'C:\Netapp-doc\DC1-A-XNAS001_SVM.Docx' -ExcelFile 'C:\Netapp-doc\DC1-A-XNAS001_SVM.xlsx'

The same two pipelines were run for DC1-B-XNAS001. The credential was the cluster's local admin. The four documents, dated 20 March 2018, are the reports this Solution uses whenever it says "as built". They are the best cross-check there is, because they describe what the cluster was, not what somebody remembers typing. In this article alone they settled the AutoSupport sender, the support flag and the empty SNMP trap configuration.

Uploading a support bundle

When support wanted more than AutoSupport carries, the log directory of a node was packed by hand. This is from site 2, in the diagnostic privilege level, through the system shell of the node.

bash
$ set diag
$ systemshell -node DC2-A-ANAS002 -command sudo tar -cvzf /mroot/etc/crash/<CASE_NUMBER>.`hostname`_etc-logs.tar.gz --exclude stats /mroot/etc/log

The archive lands in the node's crash directory, named with the case number so that support can match it. From there it had to be fetched and uploaded from a workstation; the notes list the vendor's upload page, an alternative upload page and anonymous FTP as the three ways, and do not say which one was used.

Lessons

Reading it today

← solutionz