NetApp 09 - Logging, monitoring and AutoSupport
NetApp Solution · Previous: Encryption and certificates · Next: Unified Manager and API Services
A MetroCluster that nobody watches fails over quietly and then fails for real. This article covers everything the storage sends out about itself and everything that asks it questions: EMS events to the central syslog server, SNMPv3 for Sensu, the SNMP polling the cluster itself does on its switches and bridges, AutoSupport to the vendor, the generator that produced the as-built reports, and how a log bundle was sent to support.
Who talks to whom
The conceptual design drew the storage in the middle of the infrastructure services it depends on. Reduced to the logging, monitoring and management flows, it looks like this.
flowchart LR adm["Administrators and operators"] subgraph mg["NetApp management"] ocum["Unified Manager"] end subgraph st["MetroCluster, site 1"] cl["Clusters A and B"] sw["Brocade FC switches"] br["ATTO bridges"] end subgraph inf["Infrastructure services"] sys["Central syslog"] sensu["Sensu"] mail["Mail server"] prx["HTTP proxy"] end sup["NetApp support"] adm -- "SSH, HTTPS" --> ocum adm -- "SSH, HTTPS" --> cl ocum -- "HTTPS" --> cl sensu -- "HTTPS, REST" --> ocum sensu -- "SNMPv3" --> cl cl -- "SNMP, health monitor" --> sw cl -- "SNMP, health monitor" --> br cl -- "syslog" --> sys sw -- "syslog" --> sys cl -- "AutoSupport, SMTP" --> mail cl -- "AutoSupport, HTTPS" --> prx prx -- "HTTPS" --> sup
The same flows as the firewall saw them at site 1. Addresses that the two design versions give differently (the mail server and the proxy) are left out.
| From | To | Port | What |
|---|---|---|---|
| Clusters, FC switches | Central syslog dc1-s-xcsys001.adm.example.net, 10.11.17.113 | syslog | EMS events, Fabric OS log |
Sensu dc1-a-vcsns001.adm.example.net, 10.11.16.129 | Cluster and node management LIFs | UDP 161 | SNMPv3 queries |
Cluster management LIF 10.11.10.33 and 10.11.10.34 | FC switches and ATTO bridges, out-of-band | UDP 161 | Health monitoring by the cluster itself |
| Nodes | dc1-a-vcmsx001.adm.example.net | TCP 25, 465 | AutoSupport mail |
| Nodes | dc1-a-vcprx001.adm.example.net | TCP 3128 | AutoSupport over HTTPS to the vendor |
| Sensu | API Services on DC1-A-VCOCM002 | TCP 8443 | REST, see Unified Manager and API Services |
Site 2 has its own syslog server (dc2-s-xcsys001.adm.example.net, 10.12.17.113) and its own three Sensu servers. The notes hold the commands for site 1 only; site 2 was configured the same way against its own servers.
EMS events to the central syslog server
ONTAP writes its events to the Event Management System, EMS. The design asked for five things: include the severities INFORMATIONAL, NOTICE, ERROR, ALERT and EMERGENCY, exclude DEBUG, use TCP so that the server receives the messages, send without TLS "because central syslog server does not have implemented encryption", and address the server by IPv4.
ONTAP 9.1 had two ways to do this and the notes show both. The current one has three objects: a filter that selects events, a destination, and a notification that ties one to the other. The first configuration, on both clusters, used the built-in filter important-events.
$ event notification destination create -name syslog -syslog 10.11.17.113 $ event notification create -filter-name important-events -destinations syslog
That filter passes only what ONTAP itself considers important, which is less than the design asked for. A filter of our own followed, with one include rule and one exclude rule. The notes give two ways to use it, a new notification or the existing one pointed at the new filter; the output below shows notification 2 with the new filter.
$ event filter create -filter-name all-events $ event filter rule add -filter-name all-events -type include -severity INFORMATIONAL,NOTICE,ERROR,ALERT,EMERGENCY $ event filter rule add -filter-name all-events -type exclude -severity DEBUG $ event notification modify -ID 2 -filter-name all-events $ event notification show
output 4 lines
ID Filter Name Destinations ---- ------------------------------ ----------------- 1 default-trap-events snmp-traphost 2 all-events syslog
The older way, which the notes head "old way (deprecated)", is a named destination and a route for every message name. It is in the notes too, and event destination show is where a surprise turned up: next to the factory destinations there was one named after the Unified Manager host, with that host as syslog target. It looks as if Unified Manager registered itself there when the cluster was added to it; the notes only show the entry.
$ event destination show $ event destination create -name example-syslog -syslog 10.11.17.113 $ event route add-destinations -messagename * -destination example-syslog
Two things are missing in the notes. The destination command has no protocol parameter, so the notes do not show where the TCP of the design was chosen; and cluster log-forwarding, the separate mechanism that forwards the audit log of management commands, does not appear at all. I cannot say from the Source material that command history ever reached the syslog server. The full set is a Config document: ONTAP: EMS events to syslog.
The Brocade switches log to the same server. On each switch the syslog server that was already configured was removed, the facility changed from LOG_LOCAL7 to LOG_LOCAL2 and the central server set.
$ syslogadmin --remove -ip 10.13.13.207 $ syslogadmin --set -facility 2 $ syslogadmin --set -ip 10.11.17.113 $ syslogadmin --show -ip
SNMPv3 for Sensu
The central monitoring system was Sensu. The design gave it two roads into the storage: SNMP straight to the clusters for basic status, and REST through the management software, which is the subject of the next article.
On each cluster a local user sensu was created for the application snmp with the role readonly. The command, on cluster DC1-A-XNAS001, asks for the rest interactively.
$ security login create -username sensu -application snmp -authentication-method usm -role readonly
| Prompt | Answer |
|---|---|
| Authoritative entity's EngineID | Enter, the local EngineID |
| Authentication protocol (none, md5, sha) | sha |
| Authentication password | <SNMP_AUTH_PASSWORD> |
| Privacy protocol (none, des) | des |
| Privacy password | <SNMP_PRIV_PASSWORD> |
So the security level is authPriv, with SHA-1 and DES. DES was not a choice: the prompt of ONTAP 9.1 offers none or des. Beside DES the notes remark "the default is AES", which I read as the default of the monitoring side. No community and no trap host were configured; the as-built report of March 2018 shows the SNMP table of both clusters with empty communities, no trap hosts and traps disabled. Sensu polls, the cluster does not push. The notes say the queries should preferably come from the Sensu server's IPv6 address, with the IPv4 address as the alternative; which one was used in the end is not recorded.
The command set is a Config document: ONTAP: SNMPv3 user for Sensu.
What the cluster polls: switches and bridges
In a fabric-attached MetroCluster the cluster monitors its own plumbing. Each cluster polls the four FC switches and the four FibreBridges of the site over SNMP from its cluster management LIF, and raises health alerts when a switch or bridge stops answering. That is why the firewall tables of the design have UDP 161 from the cluster to the out-of-band addresses of switches and bridges.
$ storage bridge show $ storage switch show
output 6 lines
Bridge Symbolic Name Monitored Status Vendor Model ------------------------ ------------- --------- ------- ------ ----------------- ATTO_10.11.15.47 bridgeA1 true ok Atto FibreBridge 7500N ATTO_10.11.15.49 bridgeA2 true ok Atto FibreBridge 7500N ATTO_10.11.23.47 bridgeB1 true ok Atto FibreBridge 7500N ATTO_10.11.23.49 bridgeB2 true ok Atto FibreBridge 7500N
This polling is SNMPv1 with a community string, and it collides with hardening the switches. The notes hold two passes over the Brocade SNMP configuration, and they contradict each other.
| Pass | SNMPv1 | Security level for GET | Security level for SET |
|---|---|---|---|
| First, the hardening | Disabled with snmpconfig --disable snmpv1 | 2, authentication and privacy | 3, no access |
| Second, headed "SNMPv1 configuration" | Left enabled, the six default community strings replaced | 0, no security | 3, no access |
The first pass is what one wants on paper: v1 off, GET only with authPriv, SET forbidden, and snmpconfig --set snmpv3 to define the users (the dialogue of that command is not in the notes; the factory state was six users, snmpadmin1 to snmpadmin3 and snmpuser1 to snmpuser3, all with noAuth and noPriv). The second pass is what the cluster's health monitor needs: it speaks v1. The default communities were replaced on the switches, and both clusters were told the new read-only community for all four switches.
$ snmpconfig --set secLevel -snmpget 0 -snmpset 3 $ snmpconfig --set snmpv1 $ configcommit
Then, on cluster DC1-A-XNAS001 and again on DC1-B-XNAS001, once per switch.
$ storage switch modify -switch-name Brocade_10.11.15.45 -snmp-community <SNMP_COMMUNITY> $ storage switch show -switch-name Brocade_10.11.15.45 -fields snmp-version,snmp-community
The output of the second command in the notes says SNMPv1, so the second pass is the one that stayed. The notes do not say in words that the first pass was rolled back because of the health monitor; that is my reading of the order and of the result. What remained of the hardening is that SET is forbidden and that nobody can read the switches with the factory community. The switch side is a Config document: Fabric OS: SNMP and syslog.
AutoSupport
AutoSupport is the node's habit of sending its state to the vendor: periodically, on events and on demand. It is configured per node, and the same command with thirty-odd parameters was run for all four nodes of site 1. The interesting parameters are these.
| Parameter | Value | Meaning |
|---|---|---|
-transport | https | To the vendor over HTTPS, not by mail |
-proxy-url | dc1-a-vcprx001.adm.example.net:3128 | The storage has no direct way out |
-mail-hosts | dc1-a-vcmsx001.adm.example.net | For copies to our own addresses |
-from | netapp@ad.example.net in the notes, netapp_mail@ad.example.net in the as-built report | Changed later to the mail account that existed |
-to, -noteto | ops@example.net | Three internal recipients |
-support | disable in the command | With a note beside the output: must be changed to enable |
-remove-private-data | true in the command | Switched to false "temporarily" for a support case |
-ondemand-state | disable | No remote triggering by the vendor |
The command in the notes is the cautious first version: delivery to the vendor off, private data removed, AutoSupport OnDemand off. OnDemand lets support ask a node to send a new or more detailed message; the node polls the vendor for such requests every 60 minutes, and with the state disabled it does not. Beside the output the notes carry two reminders: vendor support has to be changed to enable, and certificate validation has to be changed to false. The as-built report of March 2018 shows where the first one ended: support enabled, performance data enabled, private data not removed, throttle on, transport https through the proxy. The second is not visible in the report. The proxy had its own CA certificate, which the management servers had to trust too, and that would explain the wish to switch validation off; the notes do not say whether it was done.
Hiding private data did not survive the first support case. A line marked "temporarily enable the sending of sensitive data" sets -remove-private-data false on one node, and the report half a year later still shows it off, on all four nodes.
$ system node autosupport modify -node DC1-A-ANAS001 -remove-private-data false
How it was tested
A test message was triggered by hand on 13 September 2017; autosupport show records it as the last subject sent, USER_TRIGGERED (TEST:TEST-20170913). A week later, during a troubleshooting session, the history showed what really worked.
$ autosupport history showoutput 11 lines
Seq Attempt Percent Last
Node Num Destination Status Count Complete Update
------------ ----- ----------- -------------------- -------- -------- --------
DC1-A-ANAS001M 195
smtp re-queued 12 - 9/21/2017 10:03:54
http sent-successful 1 100 9/21/2017 09:19:53
noteto ignore 1 - 9/21/2017 09:18:19
189
smtp transmission-failed 15 - 9/20/2017 16:20:12
http sent-successful 3 100 9/20/2017 15:29:22
noteto transmission-failed 15 - 9/20/2017 16:20:12HTTPS through the proxy delivered every time. Mail did not: fifteen attempts, the configured retry count, and then failed. The older design document has the likely reason in one sentence: the mail server accepted mail on port 465 with authentication, and "other Netapp components do not support SMTP authentication and due to this limitation will be not integrated with Email services". Only Unified Manager could log in to the mail server. The nodes kept their mail host setting anyway. I have no later history in the notes that would show mail from the nodes arriving.
The same mechanism was used later as a maintenance signal: before an ONTAP upgrade a message with MAINT=2h in the subject tells the vendor not to open cases for the next two hours. That is in ONTAP upgrade. The command set is a Config document: ONTAP: AutoSupport.
As-built documentation with NetAppDocs
The as-built reports quoted above were not written by hand. NetAppDocs is a PowerShell module from the vendor's tool chest that reads a cluster through its API and produces a Word document and an Excel workbook. It was run from a Windows workstation, in a PowerShell started as administrator, once per cluster for the cluster view and once per cluster for the SVM view.
$ Set-ExecutionPolicy -ExecutionPolicy RemoteSigned -Scope CurrentUser $ Import-Module NetAppDocs $ $Credential = Get-Credential $ Get-NtapClusterData -Name 'DC1-A-XNAS001.adm.example.net' -Credential $Credential | Format-NtapClusterData | Out-NtapDocument -WordFile 'C:\Netapp-doc\DC1-A-XNAS001.Docx' -ExcelFile 'C:\Netapp-doc\DC1-A-XNAS001.xlsx' $ Get-NtapVserverData -ClusterName 'DC1-A-XNAS001.adm.example.net' -Credential $Credential | Format-NtapVserverData | Out-NtapDocument -WordFile 'C:\Netapp-doc\DC1-A-XNAS001_SVM.Docx' -ExcelFile 'C:\Netapp-doc\DC1-A-XNAS001_SVM.xlsx'
The same two pipelines were run for DC1-B-XNAS001. The credential was the cluster's local admin. The four documents, dated 20 March 2018, are the reports this Solution uses whenever it says "as built". They are the best cross-check there is, because they describe what the cluster was, not what somebody remembers typing. In this article alone they settled the AutoSupport sender, the support flag and the empty SNMP trap configuration.
Uploading a support bundle
When support wanted more than AutoSupport carries, the log directory of a node was packed by hand. This is from site 2, in the diagnostic privilege level, through the system shell of the node.
$ set diag $ systemshell -node DC2-A-ANAS002 -command sudo tar -cvzf /mroot/etc/crash/<CASE_NUMBER>.`hostname`_etc-logs.tar.gz --exclude stats /mroot/etc/log
The archive lands in the node's crash directory, named with the case number so that support can match it. From there it had to be fetched and uploaded from a workstation; the notes list the vendor's upload page, an alternative upload page and anonymous FTP as the three ways, and do not say which one was used.
Lessons
- Hardening and health monitoring pulled in opposite directions. The cluster's own monitoring of the switches wanted SNMPv1 on devices we wanted to restrict to SNMPv3. The compromise was a private community and no SET.
- The mail path was never really there. AutoSupport reached the vendor over HTTPS from the first day; internal mail from the nodes failed against a mail server that demands authentication. It should have been either relayed without authentication from the storage addresses or removed from the node configuration.
- A setting marked "temporarily" was still there half a year later. Private data removal went off for one case and never came back on.
- The design promised TCP syslog and the notes cannot show it. In ONTAP 9.19.1 the same command has a
-syslog-transportparameter, and its default is unencrypted UDP; the command in the notes names no transport. I would verify on the syslog server what arrives, and add command-history forwarding, before signing that sentence again.
Reading it today
- Syslog can be encrypted. From ONTAP 9.12.1 the EMS syslog destination supports TLS, and
cluster log-forwardingdoes too. The deprecatedevent destinationandevent routecommands are absent from the command reference from 9.11.1 on. - SNMPv3 got better algorithms. ONTAP now offers SHA-256 and AES-128. Towards the FC switches the cluster uses SNMPv3 as well when they run Fabric OS 9.0.1 or later, and the bridges are managed in-band by default since 9.8, so the SNMPv1 compromise described above belongs to its time.
- AutoSupport mail can authenticate. The current command has parameters for SMTP encryption and for a mail-host user and password;
-notetois deprecated. - The hardware is at its end. The last ONTAP release for the FAS8200 is 9.16.1, and the Brocade 6505 has been out of support since April 2025.
- NetAppDocs still exists under the same name; its current version and which API it uses could not be confirmed.