NetApp 06 - SVMs and NFS
NetApp Solution · Previous: Storage design · Next: Access and directory integration
Aggregates and plexes are what the cluster has; a storage virtual machine (SVM) is what a client sees. This article goes through the six data SVMs of site 1 (DC1), what each of them serves, how NFS was set up for the KVM (RHV) clusters and for the Balabit SCB appliance, and the two things that did not work at the first attempt: the SCB backup over SMB and the first RHV storage domain.
Six SVMs, two directions of mirroring
In a MetroCluster every data SVM exists twice: active on the cluster that owns it (sync_source) and as a stopped copy with the suffix -mc on the partner cluster (sync_destination), which is started at switchover. Five of the six SVMs live on DC1-A-XNAS001 in datacenter A. One, DC1-S-VCVSM004, lives on DC1-B-XNAS001, and the reason is not data: each cluster needs its own SVM joined to Active Directory to authenticate administrators, which is the subject of Access and directory integration.
flowchart LR subgraph ca["Cluster DC1-A-XNAS001, datacenter A"] a0["Admin SVM DC1-A-XNAS001"] a1["DC1-S-VCVSM001 active"] a2["DC1-S-VCVSM002 active"] a3["DC1-S-VCVSM003 active, joined to AD"] a4["DC1-S-VCVSM004-mc replica"] a5["DC1-S-VCVSM005 active"] a6["DC1-S-VCVSM006 active"] end subgraph cb["Cluster DC1-B-XNAS001, datacenter B"] b0["Admin SVM DC1-B-XNAS001"] b1["DC1-S-VCVSM001-mc replica"] b2["DC1-S-VCVSM002-mc replica"] b3["DC1-S-VCVSM003-mc replica"] b4["DC1-S-VCVSM004 active, joined to AD"] b5["DC1-S-VCVSM005-mc replica"] b6["DC1-S-VCVSM006-mc replica"] end a1 -->|"mirror"| b1 a2 -->|"mirror"| b2 a3 -->|"mirror"| b3 b4 -->|"mirror"| a4 a5 -->|"mirror"| b5 a6 -->|"mirror"| b6
The first version of this drawing, from 2017, had only the SVMs 001 to 004; 005 and 006 were added when two more KVM platforms arrived.
| SVM | Owner cluster | Protocol | Serves |
|---|---|---|---|
DC1-S-VCVSM001 | DC1-A-XNAS001 | NFSv3 | Balabit SCB: system backup, backups and archives of audit trails |
DC1-S-VCVSM002 | DC1-A-XNAS001 | NFSv3, v4, v4.1 | KVM clusters of the management infrastructure |
DC1-S-VCVSM003 | DC1-A-XNAS001 | CIFS | AD tunnel for cluster A; planned backup and archive SVM |
DC1-S-VCVSM004 | DC1-B-XNAS001 | CIFS | AD tunnel for cluster B |
DC1-S-VCVSM005 | DC1-A-XNAS001 | NFSv3, v4, v4.1 | KVM clusters of platform C |
DC1-S-VCVSM006 | DC1-A-XNAS001 | NFSv3, v4, v4.1 | KVM clusters of platform B |
Each SVM has a local vsadmin account as the last resort and nothing else; nobody administered through the SVMs.
Creating an SVM
The notes keep one creation command, for the SCB SVM, run on DC1-A-XNAS001. The others were created the same way or in System Manager; the notes do not say which.
$ vserver create -vserver DC1-S-VCVSM001 -rootvolume DC1_S_VCVSM001_root -aggregate data_DC1_A_ANAS002 -rootvolume-security-style mixed -language C.UTF-8 -snapshot-policy default -is-repository false -foreground true -comment "VSM for Balabit" -ipspace DC1-S-VCVSM001_ipspace
output 2 lines
[Job 807] Job succeeded: Vserver creation completed
Three details in it did not survive. The aggregate was later renamed to DC1_A_ANAS002_data1, the naming used in Storage design. The root volume was created with security style mixed, because at that moment the SCB was still meant to write over SMB; the as-built report shows unix. And the SVM got an IPspace of its own, which is the reason it has its own default route in the table below.
Every SVM has a 1 GB root volume, which holds nothing but the junction points, and one data volume mounted under /<volume name>. No load-sharing mirror of the root volume was made. The as-built report of datacenter A, taken before SVMs 005 and 006 existed, gives the network side:
| SVM | LIF | Address | Home node and port | Default gateway |
|---|---|---|---|---|
DC1-S-VCVSM001 | DC1-S-VCVSM001_nfs_lif1 | 10.11.18.236/28 | DC1-A-ANAS002, a0a-1020 | 10.11.18.238 |
DC1-S-VCVSM002 | DC1-S-VCVSM002_nfs_lif1 | 10.11.10.94/27 | DC1-A-ANAS001, a0a-1026 | none |
DC1-S-VCVSM003 | DC1-S-VCVSM003_cifs_lif1 | 10.11.10.49/29 | DC1-A-ANAS001, a0a-1028 | 10.11.10.54 |
DC1-S-VCVSM004 | DC1-S-VCVSM004_cifs_lif1 | 10.11.10.57/29 | DC1-B-ANAS002, a0a-1027 | 10.11.10.62 |
One LIF per SVM serves data and SVM management at once. The KVM SVM has no route because its only clients sit in the same subnet. SVMs 005 and 006 answer on 10.11.40.30 and 10.11.89.94. The LIF of an NFS SVM sits on the node that owns its aggregate, so that NFS traffic does not cross the cluster interconnect.
flowchart TB scb["Balabit SCB cluster, 10.11.18.225/28"] kvm["KVM hypervisors, bond for storage traffic"] v1020["VLAN 1020"] v1026["VLAN 1026, 10.11.10.64/27"] subgraph node["Node of cluster DC1-A-XNAS001"] svm1["DC1-S-VCVSM001, NFS exports for SCB"] svm2["DC1-S-VCVSM002, NFS exports for KVM"] lif1["Data LIF 10.11.18.236/28"] lif2["Data LIF 10.11.10.94/27"] p1020["VLAN port a0a-1020"] p1026["VLAN port a0a-1026"] ifgrp["Interface group a0a, LACP over two 10GbE ports"] end scb --- v1020 kvm --- v1026 v1020 --- p1020 v1026 --- p1026 ifgrp --> p1020 ifgrp --> p1026 p1020 --> lif1 p1026 --> lif2 lif1 --> svm1 lif2 --> svm2
The same interface group carries the intercluster and management VLANs; those are in Network design.
NFS for the KVM clusters
The three KVM SVMs are built alike. NFS is enabled with versions 3, 4.0 and 4.1 (the as-built report also shows pNFS left at its default, enabled). There is one data volume per SVM, security style unix, and in it one qtree per RHV storage domain.
| SVM | Volume | Aggregate | Qtrees | Clients allowed |
|---|---|---|---|---|
DC1-S-VCVSM002 | DC1_S_VCVSM002_data | DC1_A_ANAS001_data1 | management_sd, iso_sd, data_sd | 10.11.10.64/27 |
DC1-S-VCVSM005 | DC1_S_VCVSM005_data | DC1_A_ANAS001_data1 | management_sd, iso_sd, data_sd | 10.11.40.0/27 |
DC1-S-VCVSM006 | DC1_S_VCVSM006_data | DC1_A_ANAS001_data1 | management_sd, iso_sd, data_sd | 10.11.89.80/28 |
management_sd holds the storage domain of the RHV manager, iso_sd the ISO domain and data_sd the virtual machine disks. A storage domain is addressed in RHV as the SVM's name and the junction path with the qtree, for example dc1-s-vcvsm002.adm.example.net:/DC1_S_VCVSM002_data/data_sd; RHV mounts it on every hypervisor by itself, so there is no fstab entry on the hosts and the notes hold no mount command.
No export policy was created. Volumes and qtrees all use the policy default of their SVM, and that policy holds a single rule:
| SVM | Rule | Client match | Protocols | RO | RW | Anonymous uid | Superuser |
|---|---|---|---|---|---|---|---|
DC1-S-VCVSM001 | 1 | 10.11.18.225 | nfs3 | sys | sys | 65534 | any |
DC1-S-VCVSM002 | 1 | 10.11.10.64/27 | nfs3, nfs4 | sys | sys | 65534 | any |
Superuser any means root on a hypervisor is root on the export. That is what "allow superuser access: yes" in the design stands for, and it is acceptable only because the storage VLAN contains nothing but the hypervisors and the LIF.
Why RHV-M could not mount it
The first attempt to attach a storage domain failed. The cause is described in a Red Hat knowledge base article with the telling title "Why does RHEV-M fail to mount NetApp NFS?" (solution 660143): RHV does all its storage I/O as user vdsm, group kvm, both with the numeric id 36, and a freshly created ONTAP volume belongs to root. The fix was applied on DC1-A-XNAS001 to the data volume, to the root volume of the SVM and to each qtree:
$ volume modify -vserver DC1-S-VCVSM002 -volume DC1_S_VCVSM002_data -user 36 -group 36 -unix-permissions 777 $ volume modify -vserver DC1-S-VCVSM002 -volume DC1_S_VCVSM002_root -user 36 -group 36 -unix-permissions 777 $ qtree modify -vserver DC1-S-VCVSM002 -volume DC1_S_VCVSM002_data -qtree data_sd -unix-permissions 777
Afterwards the volume shows the owner by name, because ONTAP's local UNIX users and groups resolve 36 to vdsm and kvm.
$ volume show -vserver DC1-S-VCVSM002 -volume DC1_S_VCVSM002_dataoutput 11 lines
Vserver Name: DC1-S-VCVSM002
Volume Name: DC1_S_VCVSM002_data
Aggregate Name: DC1_A_ANAS001_data1
Export Policy: default
User ID: vdsm
Group ID: kvm
Security Style: unix
UNIX Permissions: ---rwxrwxrwx
Junction Path: /DC1_S_VCVSM002_data
Junction Parent Volume: DC1_S_VCVSM002_root
Snapshot Policy: defaultOwner 36:36 on the volume and mode 777 on every qtree became a rule of the design ("because this is a Red Hat virtualization requirement"), and it had to be remembered whenever a qtree was added. Later two more qtrees appear in the notes, management-43_sd and data-43_sd, on both sites; the only command recorded for them is again the qtree modify … -unix-permissions 777. The notes do not say what they were for; the name suggests storage domains for a new RHV 4.3 environment. The full command set is a Config document: ONTAP: commands for the NFS SVMs.
Mode 777 is wider than documented: the oVirt administration guide asks for owner 36:36 and sets mode 0755 on the export directory. Nobody went back to narrow it.
The notes on Red Hat clients are otherwise only a reading list: the technical report on RHEL NFS clients with NetApp, a known NFSv4.1 problem of RHEL 7.1 (clients looping on NFS4ERR_BAD_STATEID) and a NetApp article on poor NFS performance with RHEL 7.3. No tuning was derived from them, and which NFS version the storage domains were finally mounted with is not recorded.
Two smaller facts from the as-built report: the snapshot policy of DC1_S_VCVSM002_data was set to none at some point (the listing above, from October 2017, still shows default), and the volume had grown from 2 TB to 10.1 TB. A second volume, DC1_S_VCVSM002_data2, is marked in the design as temporary, "to be removed after agreement with the virtualization team".
NFS for the Balabit SCB
The SCB appliance writes three kinds of data: its own configuration backup, backups of audit trails and archived audit trails, separately for the organisation and for three partners. Each got a qtree in the 100 GB volume DC1_S_VCVSM001_data on aggregate DC1_A_ANAS002_data1.
| Qtree | Holds |
|---|---|
scb_system_backup | SCB system (configuration) backup |
scb_org_backup, scb_org_archive | Audit trails of the organisation |
scb_partner1_backup, scb_partner1_archive | Audit trails of partner 1 |
scb_partner2_backup, scb_partner2_archive | Audit trails of partner 2 |
scb_partner3_backup, scb_partner3_archive | Audit trails of partner 3 |
The junction path is /DC1_S_VCVSM001_data/<qtree>, the only client is the SCB cluster at 10.11.18.225, and the SVM speaks NFSv3 only. It also has ndmp among its allowed protocols; nothing in the notes uses it.
The SMB detour
The first plan, in October 2017, was SMB. The SVM was joined to the domain ad.example.net (with SMB1 towards the domain controllers switched off and SMB2 on, following a community thread titled "CIFS not joining AD domain"), and a user scb_system_backup was prepared for the appliance. The SCB then refused:
output 4 lines
scb/backup[11339]: ERROR (root@localhost) Backup error, cannot initialize backup method; Error mounting remote server on SMB kernel: CIFS VFS: Error connecting to socket. Aborting operation. kernel: CIFS VFS: cifs_mount failed w/return code = -115 kernel: CIFS: Unknown mount option "sec"
The notes stop there, with two references on local users and workgroup mode in clustered ONTAP, and continue with NFS. Why the SMB mount failed was never written down; the protocol was changed instead.
NFSv3 then needed one setting on the server. It was found through NetApp's "Top 10 NFS issues" article and switches off the check that mount requests come from a reserved source port:
$ vserver nfs modify -vserver DC1-S-VCVSM001 -mount-rootonly disabledWhen a mount was refused, two commands on the cluster told why: one shows the effective permissions on a path, the other evaluates the export rules for a given client.
$ vserver security file-directory show -vserver DC1-S-VCVSM001 -path /DC1_S_VCVSM001_data/scb_system_backup $ vserver export-policy check-access -vserver DC1-S-VCVSM001 -volume DC1_S_VCVSM001_data -client-ip 10.11.18.225 -authentication-method sys -protocol nfs3 -access-type read
The backup SVM that was never finished
DC1-S-VCVSM003 was designed as the target for backups and archives of the management infrastructure. Two drawings show the idea. In the first, an archive server has one leg in the KVM storage VLAN, one in an archive LAN and one in a backup VLAN that was never numbered.
flowchart LR subgraph s2["DC1-S-VCVSM002, SVM for KVM virtual machines"] q1["management-sd"] q2["iso-sd"] q3["data-sd, holds OS and data disks of the VMs"] q4["export-sd"] end subgraph s3["DC1-S-VCVSM003, SVM for backup"] k1["kvm-backups"] k2["kvm-exports"] end hv["KVM hypervisors, bond0"] st["VLAN 1026, KVM storage traffic, 10.11.10.64/27"] arch["Archive server"] alan["Archive LAN, VLAN 1004"] bk["Backup traffic, VLAN not assigned"] vm["KVM VM traffic networks"] rt["Router"] s2 --- st hv --- st hv --- vm arch --- st arch --- alan arch --- bk s3 --- bk alan --- rt rt --- vm
The second drawing adds a further SVM that exports one volume per virtual machine straight into the guests over a routed "VM data NFS" network, so that the operating system disk stays in data-sd and application data bypasses the hypervisor.
flowchart LR subgraph sx["Additional SVM for VM data"] d1["Volume data-VM1, NFS export"] d2["Volume data-VM2, NFS export"] end subgraph s2["DC1-S-VCVSM002, SVM for VM operating systems"] q3["data-sd"] end subgraph s3["DC1-S-VCVSM003, SVM for backup"] k1["kvm-backups"] k2["kvm-exports"] k3["data-VM1-backup"] end vm1["VM1"] vm2["VM2"] dn["VM data NFS network, VLAN not assigned"] st["VLAN 1026, KVM storage traffic"] arch["Archive server"] bk["Backup traffic, VLAN not assigned"] vm1 -->|"OS disk"| st vm2 -->|"OS disk"| st st --- q3 vm1 -->|"data"| dn vm2 -->|"data"| dn dn --- d1 dn --- d2 arch --- st arch --- bk bk --- s3
None of it was built. The drawing calls the additional SVM VCVSM004, a name that was then used for the AD tunnel of cluster B; there is no export-sd qtree; and the design document, in its final version of November 2019, still says of the backup SVM: "This section will be completed when backup/archive solution will be implemented." What exists is the SVM itself, a 100 GB NTFS-style volume DC1_S_VCVSM003_data, and its real job, the domain tunnel.
A note file on home directories belongs to the same pile of intentions. It is a list of links on CIFS home directories in ONTAP 9 and on automounting NFS home directories defined in AD, and one line, yum install cifs-utils, on a Linux host. It led nowhere.
Site 2
Site 2 (DC2) has the same six SVMs with the prefix DC2-S-, owned the same way (DC2-S-VCVSM004 on cluster B). The differences that can be read from the Source material:
- The volumes of the KVM and tunnel SVMs were named without underscores:
DC2SVCVSM002_data,DC2SVCVSM005_root. Only the SCB SVM kept the site 1 pattern (DC2_S_VCVSM001_data). DC2-S-VCVSM003has a root volume only; the placeholder data volume for backups was not created there.- The CIFS SVMs of site 2 were joined to the domain controllers of site 1, because site 2 had no directory service of its own when it was built.
- The NFS export section for site 2 in the design document is a copy of the site 1 section with only the SCB names changed. Client networks, LIF addresses and export rules of the site 2 KVM SVMs are therefore not documented, and they are not repeated here.
Lessons
- The export rules sit in the policy
default. It works, and it means that any new volume in the SVM is exported to the same clients the moment it is mounted in the namespace. A named policy per consumer costs one command more. - The uid 36 requirement is the kind of thing that is found once and forgotten until the next qtree. It belongs in the procedure for adding a storage domain, not in a troubleshooting note.
- An SVM that is "for backup, later" and in the meantime authenticates every administrator is two purposes in one object. When the backup solution comes, it should get an SVM of its own.